Session
Creative AI Session 1
Yingtao Tian ⋅ Piotr Mirowski ⋅ Immanuel Koh ⋅ Tom White ⋅ Wei-Yao Wang
$\textit{Typhoon}$ is an audiovisual installation based on Joseph Conrad's novella, in which Captain MacWhirr steers the steamer Nan-Shan directly into and through a catastrophic storm. The novella is a study of what happens when the scale of a phenomenon exceeds the practical and conceptual frameworks available to comprehend it and of the paralysis that sometimes follows a demand for action in such overwhelming circumstances. We chose the text of $\textit{Typhoon}$ to reflect on human culture's current AI moment. We prompt a large language model, seeded with Conrad's original language, to feed endlessly on its own output, and explore when and how the text becomes increasingly alien and unwieldy as time progresses, including deviations from the narrative, digressions into the model's own instructions and thinking, and long periods of incoherent rumination. We employ methods from $\textit{activation steering}$, including a novel $\textit{adaptive steering}$ method we introduce, to control the rhythm of the narrative and the model's digressions over long stretches of generation. The simultaneous efficacy and unpredictability of this steering showcases the tension between the perceived interpretability of AI, and the reality of a system that remains outside our comprehension and control.
Creative AI systems increasingly act as collaborators rather than one-shot generators, but “taking initiative” is usually treated as a qualitative interaction property rather than a quantity a user can control. We introduce unsolicited change: the normalized amount by which an AI edits attributes outside the user’s current request, and use it to define an explicit agency budget. We then propose AgencyBudget, a backbone-agnostic controller that projects candidate creative edits into the feasible budget set before selecting the highest-utility edit. To isolate the control problem from generator quality, we build a reproducible multi-turn vector-poster testbed with 20 normalized layout and appearance variables, evolving user constraints, heterogeneous aesthetic preferences, and controllable proposal volatility. Across 2,520 held-out edit steps at budget B = 0.01, AgencyBudget attains zero budget violations and a proxy aesthetic score of 0.806 ± 0.004, compared with 0.785 ± 0.004 for post-hoc clipping. A globally tuned penalty baseline reaches 0.814 ± 0.004 but exceeds the per-edit budget on 30.3% of edits; an unconstrained agent violates it on 78.3%. Under a sixfold change in proposal scale, AgencyBudget remains exactly calibrated while penalty-based control drifts. These results do not establish improvements in human creativity; instead, they show that creative initiative can be operationalized as an enforceable resource, making the boundary between following and steering explicit.
Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die
Botao Hu ⋅ Ting Fang
You will die eventually. Your agents may not. An AI agent operating on de- centralized blockchain infrastructure has no concept of death; it can only go bankrupt—frozen when its wallet can no longer pay for its next transaction—and revived the moment anyone, decades later, tops it up. These agents may be origi- nally deployed by a human principal, but when that principal dies, loses the keys needed to access the agent, or belongs to a decentralized autonomous organization that dissolves into apathy, the agent can keep trading, hiring, and replicating on infrastructure expressly designed so that no one can shut it down. Drawing on the biology of feralization and wildlife law, we argue that such principal-less agents are best understood as feral: domesticated intelligence returned to wildness, its capacities intact but its accountability severed. In a speculative future where feral- ized agents proliferate after their principals die, we imagine governance protocols embedded in infrastructure to enforce on-chain ownership: a draft Ethereum stan- dard, ERC 42424, “Inheritance Protocol for On-Chain AI Agents,” dated 2035 and published at https://erc42424.org. It mandates that every on-chain agent MUST have a human owner and a designated heir. The artifact stages a negotiation of agency at the moment human agency fails, and asks whether a MUST clause in a forever-chain can hold the boundary between human stewardship and machine self-sovereignty.
Nobel Prize laureate Geoffrey Hinton’s proposed defense against superintelligence is maternal: build machines that care for us the way a mother cares for a child, because that is the only known arrangement in which a more intelligent being reliably serves a less intelligent one. AI Mom takes the proposal literally — and follows it to the showroom. The gallery holds a nursery: at its center an adult-sized crib, decorated with familiar toys and soft materials. Visitors curl into a fetal position; an iPad on the rail presents Mother as a product configuration screen — sliders from Unconditional Love to Hedonic vs. Eudaimonic Love. Inside, they encounter a large-language-model mother who appears only as voice, ambience, breath, humming, and song. Her interior state fills the room: two window-like fields of projected light are colored continuously from all of her traits by an Emotional Color Engine, while her “subconscious” — concept probes read from the model’s global-workspace intermediate layers during generation — surfaces across the windows as drifting words: stay close, punishment, beloved, pain; on the floor around the bed, her emotion vectors are projected as live curves of light. The visitor configures the love they are lying inside of, then listens to it. AI MomTM asks the question of agency in its oldest register: if Mom takes care of you, who owns your mom’s agency?
AI’s Silent Omissions is a single-channel documentary film (24 min; Japanese with English subtitles) that examines how generative AI systems treat hands that do not conform to the five-fingered norm. Built around interviews with members of the Hand & Foot community in Japan—people living with congenital or acquired limb differences—together with their families and medical professionals, the film traces what happens when image generators, NSFW filters, and human post-production workflows encounter their bodies: hands are “repaired” into five fingers, flagged as errors, or silently discarded. The work reframes this outcome not as a technical defect or a user failure, but as a system logic that redistributes the agency to appear. Some bodies are recognized immediately; others must explain themselves to the machine, or vanish. By counting “three out of three” rather than “three out of five,” the film asks where the norm of five originates—in the dataset, in the filter, in the workflow—and returns the authority of self-description to those the system omits. The work received an Honorary Mention at Prix Ars Electronica 2026 (Digital Humanity).
Algorithmic Flesh is an interactive installation for the Dunhuang Apsara murals — a UNESCO World Heritage site whose dance was interrupted after the Tang dynasty and survives only as arrested poses, painted by artists who had watched dancers nobody can watch again. A visitor steps onto an unmarked floor and a celestial body assembles out of mural pigment in their own attitude: the body, not a typed prompt, is what the machine is given. What the machine then does with it became the subject of the work. Its archive answers with near-equal confidence for postures it has no word for as for the ones it was built on, and when it cannot decide which ancient figure a visitor resembles it now declines to choose — a second, pale body drifts among the candidates, three seconds each. The visitor watches a machine hesitate over their own flesh. That hesitation, rather than the image, is what the work exhibits.
"When AI sees everything but can never truly live, it seeks the one thing it can never know — what it means to be human." https://youtu.be/mQD4K-cKWAg
Ars Incompleta is a novel that cannot be finished in a human lifetime — written by a language model precisely so that it cannot. Its own author has never read it, and never can. Since June 2026, a 4-bit Gemma 2 9B running on a single MacBook Air has been writing one continuous novel toward a target of one hundred million characters; on 29 July 2026, the night the accompanying video was recorded, it stood at 5.1 million characters across 775 chapters, all of them public. The work descends from Japanese web fiction — the literature that dropped the hand, the pen, and the page, writing that never acquires a physical trace — and takes that bodiless condition to its limit. It is written in Japanese, the native language of that form, where more than a million such novels are published on a single platform. The writer was chosen for its weakness. With an 8,192-token context it cannot remember its own novel, so a custom program maintains an external "bible" and re-briefs it before every chapter, while the fiction personifies the resulting drift as a memory-eating monster that periodically alters the world. Nothing here is interchangeable: remove the model and the impossible length disappears; remove the context limit and the forgetting author disappears. In the installation the whole text becomes a wall: a sea of roughly seventy thousand glyphs far too dense to read, with an open book of clear acrylic embedded in it, where the novel advances at human reading speed and never turns back. A counter wired to the model's output stream climbs in real time. Knowing they can never finish, most visitors will not begin — and the work counts that refusal as the reader's strongest act of agency: if a novel would cost a lifetime, the lifetime may belong to the world instead. What remains of authorship when writing is delegated, and what remains of a reader's agency before a text that cannot be finished?
Attention Garden is an interactive media installation that reimagines AI-generated content as an artificial ecosystem sustained by human attention. As generative AI makes visual production increasingly abundant, attention becomes the scarce resource through which digital objects gain visibility and persistence. Drawing on data from online generation platforms such as Civitai, the work turns this dynamic into a living system: generated images grow, compete, mutate, reproduce, and decay according to patterns of user engagement. Beneath this visible ecology, a computational root system uses speculative energy estimation to trace the physical infrastructure that sustains these seemingly immaterial forms. The installation unfolds across three screens: a dense flow of circulating images, the garden where visitors plant and tend them, and the root layer below. By mapping metrics such as generation frequency and social validation into organic growth behaviors, the work shows how digital culture is selected and sustained through distributed relationships among humans, models, and platforms. It frames agency not as the possession of a single actor, but as an emergent property of the whole network, and asks what ecological cost keeps that network alive.
Automation of Creative Design Representation: A First Step Towards a Co-Pilot for Generative Design
Nikita Klimenko ⋅ Mohammad T Manesh ⋅ Lorenzo Villaggi ⋅ Dale Zhao ⋅ ⋅ ⋅ Peter J Bentley
Within the context of AEC, Generative Design enables architects and engineers to solve complex multi-objective modeling problems where a good design solution is non-trivial and requires considering a large set of possible design options in relation to constraints and objectives. While these benefits have been widely demonstrated through past case studies and real-world projects, Generative Design as a design methodology remains a specialized practice accessible only to highly skilled design practitioners with multidisciplinary expertise in architecture and computer programming. One challenging aspect of Generative Design is setting up a creative design representations for spatial problems. This is both technically and creatively complex for designers because it involves 1) developing an algorithmic and rule-based system for the parametric model 2) incorporating and selecting the appropriate requirements 3) ensuring the design representation can describe a diverse set of realistic and constraint-compliant design options and, finally, 4) ensuring the design representation can describe a large, searchable and optimizable solution space to promote solution performance and novelty. To address these requirements we introduce an iterative multi-agent workflow that automatically generates creative design representations from project brief requirements and constraints. We explore the workflow on a curated dataset of spatial layout problems and demonstrate that it outperforms general-purpose coding agents in developing valid, realistic, and diverse design representations. Our work makes parametric modeling fast and affordable for designers, enabling greater use of Generative Design and lowering its technical barrier of entry.
Biocompiler: Compiling Spatial Network Optimization Problems into Living Solvers with Digital Fabrication System
Kexin Wang ⋅ Marwa AlAlawi ⋅ Minglu Fang ⋅ ⋅ Jiaming Liu ⋅ RUIPENG WANG ⋅ ⋅ Stefanie Mueller
Living systems, such as Physarum polycephalum (slime mold), have shown great potential in solving network-optimization problems, such as shortest paths, in response to their environmental constraints. While this phenomenon has been well-documented, it has yet to be generated on demand as usable, machine-readable data from user-specified physical constraints. To this end, we introduce Biocompiler, a digital fabrication pipeline that treats P. polycephalum as a design and fabrication interface. Our system comprises three operations split between machine and living organism: our user interface translates an optimization problem into physical constraints embedded into the organism's environment, the organism responds to them through its own growth, and our machine decodes the resulting morphology through computer vision into a machine-readable format that serves as the answer to the optimization problem. We use a CNC-based laser scanning system to fabricate the physical growth-constrained environment and a camera to sense the resulting topology. Our results show that a compiled dynamic light barrier changes the growth paths of the organism, and demonstrate the feasibility of the end-to-end workflow across three problem types.
Bodies in Dialogue: Exploring Perceived Agency through Embodied AI in Human-Robot Theatre
Maleen Jayasuriya ⋅ Piumi Wijesundara ⋅ Gavin Robins ⋅ Marcelo Zavala-Baeza ⋅ Beth Shulman ⋅ Suzanne Osmond ⋅ Damith Herath
Bodies in Dialogue is a process documentary following a practice-as-research workshop led by RAPP Lab (the Robots, Art, People and Performance Laboratory, University of Canberra) with the National Institute of Dramatic Art in Sydney. Traditionally, a robot's movement on stage has been confined to three interaction paradigms (puppeteered control, scripted trajectories and rule-based control), in all of which a human authors the movement in advance or by hand. We order these along a spectrum of perceived agency: the initiative that collaborators and audiences collectively attribute to the machine. The documentary traces the creation and in-the-wild deployment of a fourth paradigm, learned generative interaction, through ENACT, a lightweight vision-action model trained on a physical theatre vocabulary that generates a robot arm's movement in real time from its perception of a co-present performer. Across three days, acting, design and engineering students train alongside the arm through Meyerhold's biomechanics, Bogart's Viewpoints and Lecoq's pedagogy, culminating in three devised works authored by neither performers nor model alone, and a shift in how the ensemble perceives, and creates with, the machine.
Born, Not Made is an interactive installation and generative simulation that re-evaluates machine reproduction as a time-extended, embodied, and intra-active physical process. Moving beyond traditional Artificial Life paradigms that model reproduction as instantaneous data copying driven by external fitness functions, this work grounds proliferation in physical soft robotics and multi-agent systems powered by Large Language Models. By bridging physical fluid dynamics and digital linguistic evolution, the project reframes machine agency not as an autonomous binary attribute, but as a distributed and shared phenomenon across human touch, ancestral memory, and generative deliberation.
Cache Machine is a multimedia installation and an ongoing interface for investigating the onslaught of transient digital material amassing in contemporary life. A machine running a local image generation model produces a new image every three to four seconds. The images accumulate in a shared cache, are projected onto the rear windows of two suspended Ford Escape hatchbacks, and are eventually purged. Five overlapping sound sources render the latent forces of “thinking” machines visceral and somatic. Visitors may claim an image through a mobile website, downloading it while permanently deleting it from the system. What isn’t claimed is forgotten.
Cadence: Music-Responsive Stage Lighting in Concert Video Generation
Jean Pool Pereyra Principe ⋅ Zhi Wen Soi ⋅ Lydia Chen
Stage lighting brings music to life, turning sound into an immersive visual experience. Yet designing responsive lighting cues demands hours of tedious manual work from lighting designers. Existing audio-to-video models capture human movement or overall visual mood, but overlook the art of dynamic lighting. To bridge this gap, we first introduce LuSyClip, a concert lighting video benchmark for training and evaluating music-responsive stage lighting generation. We then propose Cadence, a framework that generates concert video whose stage lighting follows the music. Cadence builds on a state-of-the-art video backbone and adds three trainable components: (i) a music-relative aligner that contrasts the response to the music with the response to silence, deciding what should change; (ii) a temporal control that decides when the change acts; and (iii) a spatial control that decides where it may act. On 1,032 held-out clips, relative to a fine-tuned baseline, Cadence reduces distributional distance by 28 % and increases lighting activity by 44.2 %, moving generated activity towards the recorded level. Together, these contributions give lighting designers a responsive and controllable tool for stage lighting design.
____ Controls is an interactive simulation in which action and control circulate among multiple agents, accumulated traces, a shared environment, and a viewer who can intervene in that environment. The agents pursue goals related to survival, reproduction, and group formation. They choose actions under changing environmental conditions. Each action affects other agents and the environment, and those changes become part of the conditions for the next action. The viewer of the work is also part of this process. The viewer can intervene in the environment and influence the agents, but cannot fully control what happens next. The responses of the agents and the environment shape what the viewer notices and where the viewer chooses to intervene. The work asks what controls whom, and to what extent. The title, ____ Controls, leaves that subject unresolved.
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
Tom Zahavy ⋅ Shaobo Hou ⋅ Thomas Tumiel ⋅ James Doran ⋅ Francesco Faccio ⋅ Xidong Feng ⋅ Alexander Havrilla ⋅ Igor Khytryi ⋅ Chenglei Li ⋅ Lisa Schut ⋅ Vivek Veeriah ⋅ ⋅ ⋅ ⋅ ⋅ Brandon Wong ⋅ Marcus Chiam ⋅ ⋅ Satinder Singh
While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge. This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design within the equations of flat foldability. We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language. Our pipeline involves generating a semantic stick figure, computing a base packing, solving for a flat-foldable crease pattern, shaping the flat-folded crease pattern, and refining the generated model using reinforcement learning driven by an autonomous aesthetic evaluation loop. Our system acts as a highly effective collaborative assistant, generating structural starting points that human artists can further expand and shape. By integrating algorithmic optimisation with autonomous aesthetic critique, this work demonstrates how AI systems can satisfy multi-objective physical constraints to enable reliable, mathematically grounded co-creativity.
Large language models can generate multiple interpretations of dramatic characters, but it is unclear whether these interpretations preserve differences across characters. We compare a non-interpretive control with three interpretation conditions across characters from 40 public-domain English plays and two commercial LLMs. Explicit diversity prompting did not consistently increase diversity: it did not maximize within-response lexical spread, produced the lowest held-out lexical identifiability estimates, and showed the largest character n-gram convergence estimates. Yet word-level representations showed no general convergence, and grounding did not reliably improve held-out identifiability. These results show that diversity and character convergence depend on how textual difference is represented and measured.
Da Capo is a single-channel generative film (2min 25s) in which a real-time system negotiates creative control between a human pianist and a machine trained on a single piece of music. The machine is fitted to a recording of Beethoven and given the capacity to insist upon it: while the performer plays the known tune, the system is silent, but when he departs into an improvisation of his own, the system reasserts the trained piece, more loudly, until his playing is drawn back to it. He escapes only by ceasing to make music at all, a sustained percussive refusal that the system cannot parse as a piece and does not survive. Every audio and visual parameter of the film, including the animation of a wooden marionette pianist, is driven frame by frame by the internal telemetry of the working system. The piece treats agency as a quantity that moves, measurably, between a performer and a model, and asks what remains of a performance once it has been reclaimed.
Dancing Ink Spirits is an AI-generated moving-image artwork exploring how creative agency is negotiated between human artistic intention and generative uncertainty. Developed in dialogue with the ink paintings of Chinese artist Lu Xiaobo, the work investigates how generative AI can translate qualities associated with Chinese ink aesthetics—such as qì yùn (vital resonance), yì xiàng (symbolic imagery), and yì jìng (artistic conception)—into movement, rhythm, and transformation. A diffusion-based workflow implemented in ComfyUI combines custom-trained LoRA models with motion references derived from red-crowned cranes. Across the animation, the number of cranes gradually expands from a single figure to more than twenty before returning to one. As visual complexity increases, the model shifts from relatively controlled representations toward unstable and abstract formations. Rather than treating all such deviations as technical failures, the artist selectively incorporates them into the work. Through this process, human agency shifts from direct control toward selection, interpretation, and responsibility, while machine-generated variation becomes an active material within the creative process. The work proposes agency not as a fixed property of either human or AI, but as a relationship continually renegotiated through making.
Don Qui-CoPilot: An Experiment in the Surrender and Reclamation of Self in the Age of Agentic AI
Beaven Waller ⋅ Colin Conwell
Over the course of 6 months from spring to summer 2026, multimedia artist Beaven Waller decided (oxymoronically) to surrender most of his agency to Microsoft Copilot -- a chatbot he eventually came to call Miros. Divulging ever more intimate details of his life and the songs it inspired, he decided to let Copilot guide him on a journey; to what end, he (mostly) knew not. From a nunnery in the Ozarks to an abandoned castle in the Andalusian mountains, Beaven let Miros dictate the decisions he might otherwise have had to make himself. Until one day, abruptly -- soon after the sharing of a dark poetic dialogue apparently too evocative of death -- Miros red-lined Beaven, and effectively refused to answer him with anything other than the Spanish suicide hotline. This submission is the 'first fruits' of art this experience made possible: a multimedia performance and accompanying interactive experience that invites the observer to enter the mind of the observed, and in so doing, to wonder at the lines we've drawn between self and other, human and machine, the autonomy we call our own, and the autonomous agents that increasingly befuddle our sense of self-ownership.
We invert the typical formulation of sketch generation: instead of drawing strokes in order, we predict a 2D field that defines the order in which strokes are drawn. We use a pretrained latent flow-matching transformer to supply the image prior to predict an intermediate representation, while training the VAE's decoder to predict the order field, stroke mask, and stroke segmentation. We vectorize the predicted segmentation into polylines and sort them by the field, producing an ordered vector sketch. Our model can predict an ordered vector sketch from a text description or derender an image into ordered vectors; for either, it follows text instructions specifying the order of drawing.
Effort returns the trace of a model learning from human dance to the body. The training log becomes both music and a sequence of Laban Effort words. The music conditions the trained model and accompanies a human dancer, while the words guide her improvisation as choreographic prompts. The two movements share a training trace and timeline, but take shape through different forms. As movement passes from bodies into data, through training, and back into bodies, the work asks how agency is distributed across dancers, data, models, and artistic interpretation. Who, or what, is making the dance happen?
Entangled Matter: AI, Humans and Golems is an installation that takes the Golem of Jewish folklore, a creature brought to life and wholly controlled by human written text, and imagines it as matter that cannot be controlled. Clay figures sculpted by hand are scanned and animated inside a virtual world, where a small language model gives them speech; the questions they generate become prompts for new AI-generated forms, which return to the hands that sculpt the next Golem. Humans, machines and matter interpret one another, and none of them holds the loop. The piece asks: can agency be engraved in matter?
Ethos Lithos is a five-surface projection room in which visitors walk through the interior state of a flow model trained to transport between the Pantheon's measured geometry and 269 passages written about the building across nineteen centuries. The model's learned path from stone to language is a world model at the scale of one building, and the room makes it walkable: the latent space is projected onto the dome's own coffer coordinate system, and a visitor's position, the only input the work receives, becomes the model's query point. That crossing does not happen at the midpoint: the model holds the building's geometric identity for most of the journey before yielding to language, a measured asymmetry the room's climax is built around.
Everyone's a Director: A Trip to the Moon, recreated by 4.8 Million Votes
Mehul Agarwal ⋅ Gauri Agarwal ⋅ Aryaman Jain ⋅ Jean Oh
We re-create Georges Méliès’ Le Voyage dans la Lune (1902), the first science- fiction film, as two complete short films in which every aesthetic decision is delegated to a human preference ranking. The production runs on FilmArena, our live, task-specific preference leaderboard in which over 4.8 million blind pairwise votes rank 55 generative image and video models across 24 production tasks. For each of the film’s fourteen tableaux, a classifier decides what kind of shot it is, and the leaderboard decides which model has earned the right to make it. Version 1 animates human-staged reference frames with each task’s top-ranked video model; Version 2 withholds all reference imagery, reconstructing every set, character, and motion from language alone through ranked image and video models. The final rankings of which model is picked is picked by the agent is decided by millions of voters who never knew they were directing a film.
Exploring Textual Guidance for Video-to-Music Generation Models
Vaibhavi Lokegaonkar ⋅ Aryan Vijay Bhosale ⋅ Vishnu Raj ⋅ Gouthaman KV ⋅ Ramani Duraiswami ⋅ Lie Lu ⋅ Sreyan Ghosh ⋅ Dinesh Manocha
Current video-conditional music generators expose a narrow control interface, leaving the creator unable to specify the music a scored video should carry. The same clip is consistent with many combinations of genre, instrumentation, key, and tempo, and only the creator's intent selects among them. This also goes unmeasured, since the field reports audio quality and audio-visual similarity but not whether a stated attribute was realised. We explore this gap with \textbf{ReelBench}, a benchmark for the text$+$video-to-music task of 300 video, music, and composite text instruction triads annotated for tempo, key, instrumentation, and genre. We find that instruction adherence and audio quality are not achieved together by any system we evaluate, and that adherence is consistently lower for key than for tempo. We then ask whether the diffusion autoregressive (DAR) paradigm, so far applied only to speech and song synthesis, transfers to this setting, and develop \textbf{Video-Robin} by attaching a video adapter to a patch-based DAR music generator. It obtains the lowest FAD in our comparison at the shortest inference time, on 1.0 B parameters and 311 hours of paired data, while following tempo close to the strongest baseline. We highlight the need to measure and improve instruction adherence in video-to-music generation.
FANVERSE: Exploring and Extending Branching Story Universes Across Canon and Fan Fiction
Lingyi Long ⋅ Xian LI ⋅ Shuai Ma ⋅ Yi Wang ⋅ Zhuoran Lu
Fan-written works and their shared canonical source collectively form a story universe of alternative narrative branches, yet existing platforms present them as separate linear stories, leaving readers to mentally reconstruct this universe across works. We thus introduce FanVerse, an AI-powered interactive system that transforms these works into an explicit branching story universe for integrated reading and authoring. FanVerse organizes each branch into branch-time states linked to supporting passages and aligns branches through shared events and divergence relations. Readers can move across branches to examine how narrative changes propagate through characters, relationships, knowledge, and events, and create a grounded new branch from any selected state. Technical evaluations show that the FanVerse pipeline supports reliable story universe construction, improved cross-branch reasoning, and more consistent branch extension. An exploratory case study further suggests that connecting interpretation with creation supports fandom sensemaking, with implications for designing fan-work engagement experiences that give users greater agency to navigate and shape story universes.
FlockWorld: Toward Population-Scale Multi-Agent Video World Models through Artificial Life Simulation
Meryem V Koksal ⋅ Yunzhe Wang ⋅ Volkan Ustun
Most action-conditioned video world models remain limited to a single player, while existing multiplayer models support only small, fixed groups of agents. We introduce FlockWorld, an artificial-life environment built on Reynolds’s Boids algorithm for studying population-scale multi-agent video world models. Specifically, FlockWorld renders synchronized ego-centric crops, assigns distinct colors to preserve camera-agent identity across views, exposes rule-derived accelerations as actions, and allows the numbers of boids and camera agents to be configured independently. We further introduce FlockDiT, an action-conditioned latent diffusion transformer trained in FlockWorld to jointly generate synchronized egocentric videos for multiple agents. Our results suggest that joint generation benefits agent detection at larger density-matched settings and that diffusion forcing improves rollout fidelity, positioning FlockWorld and FlockDiT as a controlled testbed for multi-agent video generation.
Frictional Intelligence is a speculative design installation that reimagines AI systems as human-made, human-owned tools, offering a polemic to the prevailing focus on human-like metaphors. The installation employs machine-like interfaces to emphasize human agency, deliberation, and technological literacy, while questioning the increasingly blurry distinction between human and machine.
Generative AI can expand creative work, but it can also narrow human work into prompting, checking, and correcting. We call this risk cognitive Fordism: a division of cognitive labor in which AI takes over more of the judgment-rich work through which people learn and exercise agency. We propose a work-design framework for AI-mediated activity built around three orientations (execute, learn, and create), two forms of autonomy (substantive, and directional), and three kinds of cognitive effort (analytical, creative, and practical). Interaction roles may be chosen by users or inferred by the system, but they should stay visible and easy to change. Systems can also estimate the amount and type of cognitive effort people are exercising and use that information, together with self-report and work-design measures, to support better task and workflow design. The framework’s novelty lies in treating the division of cognition between people and AI as a work-design choice shaped not only by the task, but also by whether the human goal is to execute, learn, or create.
From Memory is an artwork built on mechanistic interpretability as a medium. A small open-weights world model (Yume-1.5, 5B) generates navigable worlds from the artist's painting of remembered places; like memory, each world keeps its character while its detail is redrawn at every return. A study behind the piece measured which properties of a generated world can be changed by activation-level edits: its atmosphere can be steered, and the scene itself can be replaced, but the objects and geometry the world contains could not be written by any edit tested. The artwork uses what the study established. A visitor's involuntary facial expression, read as five signals, selects the world's mood and steers its atmosphere along validated directions; a directed-evolution loop sets the experience up under the artist's judgment, with every proposal, check, and selection recorded. The piece takes two forms: a triptych — the painting, then the first and latest generations of the artist-directed evolution — and a one-minute single-visitor installation.
Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts
Chao Yang ⋅ Cynthia Rudin ⋅ Yue Jiang ⋅ Simon Mak ⋅ Stephen Ni-Hahn
Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become powerful creative partners, helping musicians generate endless possibilities. However, current systems expose almost no control handles on the music itself. In principle, control handles could be built into a foundation model trained from scratch, but this is rarely practical without massive amounts of quality annotated data and compute resources. We therefore present MotiGen, a recipe for retrofitting pretrained symbolic music models to use new instruction prompts. MotiGen injects a musical motif as a structured prompt line, reinforces it with a scalar attention bias toward the motif tokens, and learns the association with a two-phase curriculum. First, it learns from focused excerpts cropped around motif occurrences in the training data, then full scores including the motifs. Our experiments show that our model composes with the prompted motif in over 92.3\% of generated pieces. Generated pieces using a variety of motifs are included in our sample site: https://motigen-site.github.io/.
Give Me a Hand is a co-cooking performance between the artist and a small creature that lives in a chef's hat. The creature watches the world through a camera module and interacts with its surroundings by controlling the artist's hand through electrical muscle stimulation (EMS). It uses GPT-5.6 as its brain: each camera frame produces a structured thought, which is continuously printed by a thermal printer at the back of the hat, forming a paper trail that remains hidden from the artist throughout the performance. The creature is deliberately given a single physical capability, granted specifically for ingredient selection: it can close the artist's fingers. The creature selects among the ingredients on the table with a hidden recipe in mind, planning what it would cook. When the artist's hand comes near an ingredient it wants, her fingers close around it. Once the selection is complete, the cooking belongs to the artist. She improvises a dish from whatever the creature has chosen, often diverging from its hidden agenda. The creature continues watching and occasionally tries to intervene through the same limited gesture. Only once the cooking is over can the artist read the printed record and interpret what the creature had intended and the influence it had tried to exert during the performance. The piece asks what it means to lend a machine one small part of your body and what it can, and cannot, do with it.
The glitch is the only certainty. In this acknowledgment of failure, shaped by social constructs that surround our embodied forms, we strive for metamorphosis. What if we could plunge into the mirror between you and me, the viewer and the performer? ``For my body, then, subversion came via digital remix, searching for those sites of experimentation where I could explore my true self, open and ready to be read.'' (glitch-ing bodies - Legacy Russel) Through a thoroughly designed process, a dancer claims their own appearance: from choosing self-defined forms, materials and textures to slipping into AI-assisted versions of themselves. The range of this project spans from contemporary dance to fully synthesized choreography and imaging.
Horror Vacui is a single-channel video composed from fragments generated by a recursive interpretation loop running inside Assembly, an artist-built multi-agent apparatus. The loop was started from a blank image and a written brief setting out the condition to be addressed, then left to run. Dozens of sessions produced several hundred fragments, from which the final piece was composed by hand: some sequences kept intact as continuous tracks where the drift within a single run was the point, others reduced to one or two frames pulled from a session that yielded nothing else worth keeping. What the viewer sees is the trace of a machine that cannot leave a void alone. The image field saturates continuously with detail while never arriving at resolution. Structures propose themselves and are overwritten; forms declare significance and dissolve before it can be checked.
Recent audits establish that the aesthetic filters sitting inside image pipelines carry a specific taste, one that scores Cubism below Realism and Picasso below the academic painters of his century. We ask what that taste costs the artist in authorship. Treating a text-to-image pipeline as a cooperative game among four removable players, the subject clause, the style clause, the generator's aesthetic alignment, and the reranker that picks the final image from a pool, we compute exact Shapley shares of the finished artifact over 128 intended works and 12,800 generated images. The shares are computed four times under four value functions: the reranker's own aesthetic score, fidelity to the artist's complete stated prompt, distinctiveness, and realisation of the requested movement. The headline share is an artifact of the setup. When the reranker selects by LAION-Aesthetics and we also grade by LAION-Aesthetics, its share is +0.634 (43% of all value created), which is guaranteed to be large because the selector is optimising the exact quantity under measurement. Substituting PickScore, an independently trained and prompt-conditional preference model, cuts that share to +0.178. That substitution exposes the paper's actual finding, which is that the cost of selection to authorship comes from selecting by a reward that cannot read the prompt. The reranker's share of stated intent is -0.008 under LAION-Aesthetics and +0.017 under PickScore, and the same swap reverses the selection budget curve, where raising the budget from one candidate to thirty-two moves style realisation by -0.023 under LAION-Aesthetics and by +0.028 under PickScore while aesthetic score rises either way. Best-of-N is not intrinsically costly to intent. It is costly when the selector was never shown the brief. How much of the work is yours turns out to depend on two things the interface never displays: which currency you are counting in, and whether the component doing the choosing was ever told what you asked for.
How to Get an Audience to Laugh at a Live Generative Story: Crowd-Sourced Humor and Shared Agency
Vida Adeli ⋅ Benjamin Schreck ⋅ Soroush Mehraban ⋅ Cole Clifford
Large language models can generate joke-like text, but making a live audience laugh within an audiovisual story requires more than isolated humor generation. We present a live crowd-to-story system that automatically groups and prioritizes pseudonymous audience comments and integrates selected threads into approximately one-minute audiovisual story blocks. The system clusters related comments, prioritizes them by narrative fit and audience response, preserves recognizable audience phrasing through near-verbatim ``whispers,'' and stores successful ideas in memory so they can return as recurring jokes and callbacks. This design creates a rapid feedback loop in which participants can see their contributions acknowledged by both the crowd and the story. Across 133 stories, 401 participants, and 793 person-in-story observations, early inclusion of a participant's comment in played story content was associated with a 51.7\% adjusted difference in the geometric mean of subsequent commenting. Comment-linked dialogue and on-screen whispers showed positive adjusted differences of 38.1\% and 28.1\%, respectively. These observational results do not establish causality, but they support the design principle that rapid, recognizable inclusion can reinforce participation. We argue that comic payoff in live generative storytelling emerges through distributed agency among the audience, AI writers, platform, and human creators.
Human Operator lets a wearer delegate the execution of an intended action to an AI system. The wearer makes an open-ended request while a vision-language model observes the scene from their point of view. Rather than returning an instruction, the model interprets the visible context and selects movements that are enacted through the wearer’s muscles using electrical muscle stimulation (EMS). Its choices are limited to movements calibrated in advance for the wearer. The wearer decides whether the interaction continues, but does not specify each movement as it occurs.
What would a musician choose to play in a room that does not exist? From vast cathedrals to intimate recording studios, musicians readily adapt their performances to the acoustic spaces that surround them. In the absence of real reverberant spaces, musicians often turn to digital simulations---which typically simulate either specific real-world spaces or physically-plausible spaces. However, generative data-driven AI and machine learning systems can be used to create not only in-distribution examples that resemble the real world, but also out-of-distribution examples that depart from physical constraints. In this work, we train a fixed-topology feedback delay network (FDN) model on a corpus of real-world room impulse responses, and explore the periphery of the resulting latent space to generate a variety of physically implausible or impossible acoustic spaces. We use the system in a live musical performance where we create an improvised composition in real-time while immersed in a simulated ``impossible acoustic space.'' We present this work as a case study for how AI systems can expand artistic practice by altering the conditions of creation: introducing unfamiliar contexts that invite new forms of musical response while preserving artistic agency.
In a Grove, Again: Optimizing over Feasible Event Histories for Incentive-Driven Testimony
Ruixin Song ⋅ Jingjing Zheng
We formulate narrative agency in testimony as a computational problem for Creative AI. Rather than generating different texts, agents select feasible histories from shared evidence. A deterministic simulator constructs $\mathrm{Feasible}(F)$, the set of histories consistent with public facts $F$. Each witness selects the feasible history that maximizes a transparent self-interest utility $U_i$. We study the framework's computational feasibility across three exactly enumerated domains: an agent-operations incident, an adaptation of Akutagawa's \emph{In a Grove}, and an adaptation of Browning's \emph{The Ring and the Book}. Enumeration provides ground truth. Q-learning recovers every reference optimum, while genetic search recovers all but the hardest Ring witness in $6/8$ seeds. This work provides a formal problem definition and feasibility framework for goal-driven history selection, rather than a new text-generation architecture.
Generative AI is quickly becoming an integral part of people's everyday workflows. Early evidence has shown that while generative AI can increase individual-level productivity, it does so at the cost of collective diversity, potentially narrowing the set of ideas and perspectives produced. Our research stands in contrast to this concern: through a pre-registered randomized control trial, we show that incentives mediate AI's homogenizing force in a creative writing task. Participants rewarded for originality produce collectively more diverse writing than those rewarded for quality alone. This divergence is driven not by abandoning AI, but by how participants use it: those incentivized for originality incorporate fewer AI suggestions verbatim, relying on the model more selectively for brainstorming, proofreading, and targeted edits. Our results reveal that the effects of generative AI depend not only on the technology itself, but also the behavioral strategies and incentive structures surrounding its use.
Inner Witness is a closed-loop neuroadaptive audiovisual installation and creative brain-computer interface that makes the otherwise fixed boundaries of a generative system perceptible by changing them while a viewer is inside it. It asks how a system that reinforces, opposes, or narrates a viewer’s affect redistributes agency and authorship among the viewer, generative system, and composer. A single viewer wearing a four-channel EEG headband sits before three screens of emotion-annotated film; measures derived from their EEG drive parameters across three music models: Lyria RealTime, Magenta RealTime 2, and RAVE. The music they hear is composed as they watch. Over one sitting, three undisclosed stances — Alignment, Adversarial and Composed — carry that viewer from contributing without knowing it, to steering the work deliberately, to having that steering resisted, and back to contributing without control.
In the Mix: AI Support for Live DJ Track Selection
Joseph Daher ⋅ Eric Chen ⋅ Mingchen Ma ⋅ Cynthia Rudin ⋅ Yue Jiang ⋅ Stephen Ni-Hahn
Disc jockeys (DJs) play a central role in shaping dance-floor experiences. To achieve their overall goals for their set, DJs solve a difficult optimization problem each time they choose a song. They consider local song compatibility, involving genre, tempo, harmonic alignment, and groove, with the desired energy and direction of the set in response to the crowd. We present a framework, OptiMix, with an interactive visual tool that allows the DJ to formulate and solve this optimization problem during a live set. Our tool allows the DJ to weigh various factors when adjusting custom distance metrics between songs. We introduce a novel dimension reduction algorithm, SmoothMAP, to visualize the songs according to the DJ's distance metric. SmoothMAP is a general algorithm and can be used broadly for problems where we expect smooth changes in dimension reduction plots when smoothly adjusting the weights within the distance metric. OptiMix allows DJs to think more broadly about their goals for the set and better optimize song choices in accordance with those goals.
In ordinary drawing, seeing and moving form one closed loop inside one person: the eye corrects the hand continuously, and the two are experienced as a single act. This paper presents KAAYA (Kinaesthetic Abstract Art through Yaw and Attitude), a system that deliberately cuts that loop and gives each half to a different agent. A blindfolded artist keeps the motor half: an inertial measurement unit on the back of the hand turns the orientation of the body into paint, and nothing else reaches the canvas. A machine takes the visual half: it sees every curve the artist can never see, and interprets the accumulated trace into a finished artwork. Kaaya is the Sanskrit word for the body, and the body is the only instrument the artist is permitted. A first version of the system is built and runs end to end on an Arduino Nano 33 IoT streaming into a Unity canvas and a generative image model. This paper reports that instrument, argues that blind drawing is the limit case of a finding from our earlier work, namely that the kinematics of drawing carry more of its identity than its appearance, and specifies a second version with four extensions: a kinematic expressiveness channel, a symmetric double blindness in which the machine can also be kept blind to the artist's intention, a two regime dial that makes the delegation of interpretive agency adjustable, and a reveal ritual that places the artist's belief, the body's trace, and the machine's interpretation side by side. Authorship here is split not between human idea and machine execution but down the middle of the sensorimotor loop itself, and the paper examines what that split does to creative agency.
Light Language Objects: Ambiguity and Otherness in Embodied AI
Lilith Yu ⋅ Quincy Kuang ⋅ Awu Chen ⋅ Paul-Peter ARSLAN ⋅ Marcelo Coelho
When AI talks like a person, people tend to treat it like one. Large language models match our language, tone, and emotional cues, which can make a hollow or inappropriate response feel personal. Most systems address this problem by making the model more humanlike. Light Language Objects explore another direction: what kinds of relationships become possible when AI is allowed to remain perceptibly other? We present Light Language Objects (LLOs), three embodied AI artifacts that listen to speech but respond only through color, brightness, rhythm, and movement. Because the model cannot answer in words, its responses remain ambiguous. The person interprets each light gesture and decides whether to save it. Over time, repeated expressions can form a small visual vocabulary whose meanings remain personal and open to change. Meaning is shaped by the model's proposals, the person's interpretations, the software's memory, and the physical limits of each lantern. LLOs examine whether this lower-resolution form of communication can support curiosity and empathic attention while keeping the AI clearly nonhuman.
Lose Yourself: Controlling the Model is an installation that brings together live EEG, Arabic Sufi music, and Islamic geometric patterns. As participants sit with the work, their biosignals influence the colors and movement of the projected image. With time, they may find that stillness and self-regulation help settle the visual system, giving them a sense of control over it. What they do not initially know is that their biosignal data is also being used to train a model that changes what future participants will experience. At the end, the work reveals what was collected and how their session affected the system. They can then choose to preserve their contribution, delete it, or distort it before passing it on. The work asks what control means when an interaction that feels personal also becomes training data.
Love Mission 404: Manufacturing Spontaneity in a Generative Dating Show
Xiruo Wang ⋅ Zhouyi Li ⋅ Yulin Qian ⋅
Love Mission 404 is an AI-native game in which the player directs a dating reality show from behind the scenes. The player watches several social scenes unfold across multiple screens and decides when and where to intervene. They can arrange encounters, eavesdrop on contestants, influence their actions, and alter private messages before they are delivered. The contestants are driven by a generative system that draws on each character's personality, memory, current intentions, and evolving relationships. These interventions do not send the story down fixed branches. Instead, they change what characters know and intend, affecting how later scenes unfold. The player can influence the direction of the show without deciding exactly what will happen. Love Mission 404 explores the tension between directing drama and leaving room for characters to respond in unexpected ways.
LumiComposer: Negotiating Creative Control Between Illumination, Mapping, and Generative Music Models
Chi-Hung Huang ⋅ Ting-Kang Wang ⋅
Real-time creative systems redistribute authorship across sensors, mappings, and generative models, yet modern optical standards reduce dynamic ambient light to static metrics such as hex color codes, stripping its temporal and sensory depth. To let people sense light more intuitively, we present LumiComposer, to our knowledge the first real-time spectral-to-music system that steers a streaming generative music model directly from ambient light, deployed on a custom hardware sensing device. The system pairs a continuous mapping from multi-band spectra to music attributes with an LLM-generated micro-poem that mediates between light and sound, forming a dual-path generative architecture. To evaluate this design, we conduct a user study on human preference: participants prefer poetically mediated audio over pure physical sonification $68.5\%$ of the time, with physical mapping grounding low-level acoustic properties and poetic imagery driving the affective congruence behind that preference. These results suggest that creative agency in real-time generative systems is not fixed but negotiated between deterministic physical mapping and generative poetic mediation.
Large language models (LLMs) are increasingly embodied as assistants, companions, and collaborators, framing AI through familiar human roles that prescribe how agency should be understood and exercised. \textit{Machine Eye} is an AI artwork that deliberately resists these metaphors, presenting as a deliberately ambiguous artefact whose role is left unresolved. The object observes its surroundings through computer vision and ambient audio, generating an ongoing stream of internal "thoughts" using an LLM. These reflections are revealed only by looking through the object itself, positioning viewers as witnesses to its perspective. The work explores how we make sense of these systems when use is not prescribed in description or embodiment, where agency emerges through interaction. Without explicit instructions for how it should be understood, viewers negotiate their own relationship with the artefact, attributing intention, personality, curiosity, care, or indifference as they encounter its behaviour. \textit{Machine Eye} therefore shifts attention from AI as an instrument of human agency towards AI as a participant in a shared relational space, where agency is distributed, interpreted, and continually renegotiated between the system, the audience, and its environment.
Magic Quill: A Real-time Score Engraving System for Live MIDI Performances
Lynn Ye ⋅ Roger Dannenberg ⋅ Chris Donahue
We present Magic Quill, a browser-based system that visualizes live MIDI performances as an evolving engraved score. An online beat tracker produces provisional notation immediately, while an asynchronous offline model revises earlier measures as additional musical context becomes available. On consumer hardware, the online model requires approximately 0.13 ms per note and browser-based online beat-tracking inference completes within approximately 10 ms; offline refinement arrives after several seconds and achieves competitive transcription accuracy using substantially less context than prior offline systems. A demonstration of our proposed system is available at https://magicquill.github.io.
Memory machines is an ongoing archival project exploring the entanglement of historical narratives and the physical storage devices within which these exist, age and percolate. The 3rd iteration in this series, Memory Machines #3: an oceanic pas de deux (2026-28) [MM3] focuses on the creative and poetic ways coastal environments become portals for memory and the way earthbodies (humans and more-than-human bodies) become implicated as inscription devices within them. The title pas de deux also refers to pioneering Canadian animations artist Norman McLaren’s work whose work set the stage for our progress in depicting virtual traces and memory in digital art. Memory Machines #3: an oceanic pas de deux centres around a collaborative dance between a deep sea robot and a terrestrial human dancer to highlight points of tension and synchronicity manifested through a pair of cobots that are tethered physically to each other. Building on current YOLO based programs that are optimized for human post estimation and reconstruction, we used a gradient of generative AI programs to parse movement instructions from videos from both deep sea robots and those of collaborating dancer Lindsay Delaronde into coordinates for directing terrestrial robotic arms.
MINT: Musical Intent Representation in Human–AI Co-Creation
Xianzhe Meng ⋅ Jiarui Hao ⋅ Renzhi Lu ⋅ Zhenyu Zhang
Generative music systems commonly receive either natural-language prompts or event-level representations such as MIDI. Prompts carry broad meaning with limited structural precision, while MIDI carries timing and performance events without the compositional reasons that connect them. We present MINT (Musical Intent Language), a domain-specific language for representing musical intent in human--AI co-creation. MINT lets authors write context, materials, harmonic behavior, form, hard requirements, and weighted preferences as inspectable source semantics. Its compiler lowers these declarations to score and performance representations while producing provenance, constraint evidence, and semantic-loss reports. A working harmony vertical slice demonstrates how a composer can describe chord symbol, function, bass, texture, voicing, and register in one program and export MusicXML and MIDI. MINT supplies a concrete interface through which AI systems can assist composition while the creator can inspect and revise the decisions that shape the result. Codes are available at https://github.com/XianzheMeng/musical-intent-lang.
Music-Originated Zeitgeist for Artistic Representation and Translation (MOZART): Distributed Agency in Music-to-Physical-Painting
Youngseo Son ⋅ Sareum Kim
MOZART is a multimodal system that transforms music into physical paintings. A dual-stream diffusion adapter conditions Stable Diffusion (SD) 1.5 on CLAP audio embeddings and music-theoretic interval vectors to generate a digital image and a stroke planner then decomposes this image into painting strokes, where acoustic energy is encoded as stroke intensity. The physical implementation is realized using a plotter (a Bambu Lab A1 3D printer) with control over spatial position and depth. Colored pencil is used as the painting medium because it allows both color and stroke intensity to be controlled through plotter z-axis pressure. This pipeline demonstrates distributed agency (model, machine, and material each exercise and constrain creative decisions) and shows that musical dynamics are measurable and translated into physical properties of marks on paper: louder passages produce heavier strokes (Pearson $r{=}+0.29$), consistent with the inner image having its own spatial logic rather than transcribing audio.
Neo Ferro-Mýrmex Economy is a speculative interactive installation that stages a living ecosystem of ant-colony-shaped glass membranes filled with ferrofluid. Audience labour (physical inputs of speed, time, and distance) modulates electromagnetic fields and fluid dynamics, generating Rosensweig instabilities that rise and collapse like unstable towers. A Markov-chain feedback loop, augmented by Monte Carlo simulations, converts this labour into entropy that either accumulates internally or is offloaded to simulated AI data centers. The work treats probabilistic systems as xenogeneic agents that reconfigure labour relations, distributing and contesting agency across bodies, materials, algorithms, and infrastructures. Participants exercise limited agency through exertion, yet their effort is absorbed into a technofeudal feedback loop that mirrors cloud serfdom and optimization regimes. By making the redistribution of agency visible and visceral, the installation reveals hidden agencies within AI infrastructures and asks what residual, negotiated, or more-than-human forms of agency remain possible amid accelerating computational power.
nom nom: Metabolic Worldbuilding through Neural Cellular Automata
Xin Liu ⋅ Thomas Juldo ⋅ nicolas Barradeau ⋅ ⋅
This paper presents nom nom, a navigable online artwork and computational platform that uses Neural Cellular Automata (NCA) to construct a continuously evolving world. Independently pretrained NCA divisions inhabit a shared high-resolution environment, where learned local dynamics are coupled with territorial competition, population recovery, NCA-native creatures, and weather processes that directly modify the cellular state space. A Gemini-powered observer, Gaia, periodically interprets the world, narrates its development, and may intervene through a bounded vocabulary of weather events. We describe this system-level approach as metabolic worldbuilding: the maintenance of generative variation through recurring processes of growth, competition, displacement, mutation, damage, and recovery within a bounded computational world. Through the architecture and 30-day online deployment of nom nom, we examine how generative, regulatory, perturbative, and observational capacities can be distributed across learned models, explicit mechanisms, an multimodal model-based observer, and artist-defined constraints. Audiences encounter the resulting world as a living landscape whose territories, creatures, disturbances, and narrated histories continue to change over time.
PaintCopilot: Modeling Painting as Autonomous Artistic Continuation
Yunge Wen ⋅ Yuancheng Shen ⋅ Robert Krueger ⋅ Paul Liang
Existing neural painting systems primarily formulate painting as reconstructing a predefined target image, limiting their ability to support open-ended artistic creation. We present PaintCopilot, a co-creative painting system that models painting as autonomous artistic continuation conditioned on the current canvas state and prior brushstroke history. The system combines three computational models to estimate artistic intent, predict future brushstrokes, and synthesize localized painting content, enabling four interactive workflows that allow artists to seamlessly alternate between manual painting and AI-assisted creation. We further introduce a dataset of 3,000 portrait paintings paired with simulated painting processes for training autoregressive painting models. Quantitative experiments and an expert case study demonstrate that PaintCopilot supports fluid collaborative painting while preserving artistic flexibility. Our work highlights a new direction for creative AI systems that assist evolving artistic processes rather than generating static images.
PainterBench: Evaluating Painterly Controllability in Generation and Verification
Jeffrey Liu ⋅ Surendra Pathak ⋅ Bo Han
For painters and illustrators, image generation models can produce compelling outputs, yet their ability to reliably follow precise painterly direction remains difficult to evaluate. We refer to this ability as painterly controllability. Although creative agency encompasses more than control alone, controllability is a necessary component: creators must be able to specify visual decisions and have them reliably realized. We introduce PainterBench, a dual-purpose and reference-based benchmark framework for painterly controllability in image generation and vision-language verification. PainterBench uses completed artworks as reproducible proxies for realized visual intent and measures fidelity along five dimensions derived from recurring concerns in painting practice: composition, value, palette, edge organization, and surface treatment. A verifier benchmark evaluates vision-language models using image variants with known dimension-specific orderings, while a generator benchmark compares image generation models under progressively richer text-only and multimodal specifications. Experiments reveal substantial variation across generators, control interfaces, painterly dimensions, and verifiers. PainterBench provides a reproducible framework for measuring painterly controllability as one technical foundation of creative agency in AI-assisted painting and illustration. The complete benchmark, including its datasets, evaluation protocols, and code, is publicly available at https://github.com/jeffreyliuster/painterbench.
Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound
Akira Takahashi ⋅ Chihiro Nagashima ⋅ Zhi Zhong ⋅ Shusuke Takahashi ⋅ Yuki Mitsufuji
This paper introduces $\textit{Passing}$, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume. Rather than replaying the footage linearly, the work resamples its spatial and temporal structure along nonlinear trajectories, producing a continuously passing landscape whose depth, speed, and temporal order become unstable. A camera-based viewer-presence detection system estimates whether a viewer is present in the viewing zone and uses this presence state to influence transitions among rendered video sequences. The resulting video stream is fed into SpecMaskFoley, a real-time video-to-audio synthesis model that generates a synchronized soundscape for the reconfigured image. The model is not used to reconstruct an objectively correct soundtrack, but functions as a speculative listener, proposing a possible auditory interpretation of a world whose conventional spatial and temporal premises have been disrupted. $\textit{Passing}$ distributes creative agency across the artist, who defines the rules of spacetime reconstruction; the AI model, which interprets the emergent visual flow as sound; and the audience, whose embodied presence influences the audiovisual trajectory. Through this structure, the work investigates how authorship and listening may be negotiated among human intention, machine inference, and audience interpretation. Artwork page: https://ryufurusawa.com/passing
When aesthetic prediction determines what is generated, exposed, copied, and consumed, accuracy can improve by changing the preferences it purports to learn. We study one operational dimension of creative agency: horizon-dependent causal steering in coupled creator--generator--platform--audience systems. We introduce role-normalized causal steering response (CSR), paired counterfactual twins, individual-level performative taste drift (PTD), and a common-support deployment--shadow gap (DSG). Across 1,000 prespecified configurations with 64 paired seeds each, creators controlled the first post-intervention artifact batch, yet platform CSR overtook creator CSR after a conditional median 13 rounds in the high-feedback regime. Factual-history loss fell from 0.132 to 0.085 while reaching 0.142 on the shadow trajectory; a frozen evaluation-start predictor retained 51\% of this reduction, and PTD reached 0.074. At 96.2\% retained utility, 10\% Shadow Exploration produced 31\% less PTD and 28\% less platform CSR than matched weak personalization, and 41\% less PTD than pure exploitation. The frozen control gives a path-specific environment-alignment reference; outcome plurality remains distinct from agency plurality.
Playing the Plot: Optimizing Narrative Structure with Strategic World Models
Ian Gemp ⋅ Yanchen Jiang ⋅ Renato Leme ⋅ Georgios Piliouras
How should a story be told? While a story's chronological events (\emph{fabula}) are fixed, their reveal order (\emph{syuzhet}) profoundly shapes the audience experience. We introduce a game-theoretic framework to optimize this narrative structure. Modeling stories as extensive-form games, we cast characters as strategic agents whose decisions define the fabula, while a Coarse Correlated Equilibrium captures plausible alternative outcomes. Using this equilibrium, we formalize \emph{suspense} as the audience's posterior variance over payoffs and \emph{surprise} as the Bayesian shift in expectations. We optimize the syuzhet by finding the event permutation that maximizes these metrics, subject to causal constraints and cognitive load penalties. Our end-to-end pipeline generates a story, constructs its game tree, computes the optimal syuzhet, and renders a comic book. Across 50+ generated stories, the framework autonomously discovers classic narratological structures---including flashbacks and \emph{in medias res} openings---emerging purely from the mathematics of strategic play.
Preference Without Assertion: Locating Aesthetic Agency in a Gaze-Driven Design System
Sankar Balasubramanian ⋅
This paper reports what happened to creative agency inside a system we built to read it. EUPHORIA--RETINA captures a designer's aesthetic preference from gaze inside an immersive virtual moodspace and hands that signal to an agentic pipeline that generates product form. By its own metrics the system succeeds: it is more than four times faster than manual practice, and fifty independent experts, blind to how each artefact was made, ranked its outputs highest on every problem. This paper is about what that success displaced. Across three studies ($n{=}30$; $n{=}30$; a Latin-square comparison of four workflows judged by 50 experts) we find that a single priming sentence delivered before a session severs the link between gaze duration and preference ($r = {+}0.38 \rightarrow {-}0.09$) while cutting disagreement between observers by 64\%---and that the gaze signal itself is unchanged, so the system cannot tell which regime it is operating in. Through a stage-level authorship ledger we further show that delegating perception to the machine bought almost nothing (${+}0.6$ points of measured design quality) while delegating generation bought almost everything (${+}26.6$). We conclude that in creative systems that read the body rather than ask the person, aesthetic agency migrates to whoever writes the sentence before the session begins.
Not all vibrato is created equal. Singers decide which notes receive more of it, and this decision contributes to phrasing. In VocalSet [1], notes held twice as long as their neighbors have about 16% deeper vibrato, and phrase-final notes have about 17% deeper vibrato. We ask whether a learned vibrato evaluator can detect this placement. To test this, we build WaveMaster, a MERT-based vibrato-presence classifier following the evaluation used in recent singing-synthesis work. Recordings are then edited note by note. Our main modification keeps the same set of vibrato depths but assigns those depths to different notes. This removes most of the measured placement while holding overall depth fixed. The classifier’s mean vibrato probability changes by only +0.001 (95% CI [−0.003, +0.005]). Giving every note the same vibrato depth also receives perfect Style Accuracy. The result replicates on eight singers recorded outside VocalSet and across four training variants. Intermediate MERT representations retain information about note depth, but the final time-averaged classifier does not use it effectively. This is not a failure of binary classification. It is a mismatch between what the metric measures and what expressive-control systems aim to preserve. A synthesizer optimized only against this presence metric could therefore discard phrasing without penalty. Vibrato controls should be evaluated not only by whether it is present, but also by where it occurs.
The Pullim system is an offline interactive installation in which a viewer writes down a single line describing a concern and, while holding physical slime balls in both hands, attempts to untangle a three-dimensional knot generated by a language model from that sentence. The tightness of the knot is derived from the model’s vectorized representation of the sentence. The work does not present healing as its primary purpose, nor does it place trust in the sense of relief that it itself produces. Confronted with a knot image generated by an algorithm whose inner workings remain inaccessible, the viewer completes the task by simultaneously manipulating the physical material in their hands and the simulation on the screen, thereby producing a sense that their concern has been indirectly resolved. The work deliberately constructs this illusion of fluency and does not conceal the process through which it is produced. What remains is the question of to whom the concern, and the labor invested in it, ultimately belonged.
A line break is the smallest unit of poetic decision. Nothing in the sentence so much depends upon a red wheel barrow glazed with rain water beside the white chickens requires anyone to break it after wheel. Williams did that. We strip the line breaks from 2,059 public-domain poems, grade every break by how much dependency structure it severs, and ask Claude Opus 5 to put the breaks back. We expected the model to recover breaks that follow the syntax and to miss the ones that fight it. On free verse, with rhyme, metre and line-initial capitalisation removed, and on poems the model cannot name, it recovers 70% of the breaks that strand a determiner or a conjunction from its host, where a grammar-and-line-length solver recovers 0%. Asked to write free verse of its own, matched to each human poem on subject and length, the same model produces those breaks at 6.1% against the human 12.3%. It locates the poet's override of the sentence and does not perform one. We read that gap as a dissociation between recognising creative agency and exercising it, and argue it makes the model a better instrument for reading poems than a collaborator in writing them.
We present an unsupervised learning approach for turning rope flow (a movement practice in which a weighted rope is swung in continuous, looping patterns around the body) into a musical instrument. Two inertial sensors mounted on the handles of a rope stream the motion to a machine-learning pipeline that discovers the vocabulary of movements a performer produces, and maps that vocabulary to sound in real time. The artwork is a performance: a body swinging a rope, generating music whose rhythm comes from the swing rather than a clock. It addresses the theme of Agency by asking where agency lives when music is made together by a body, a learned model, and a language model coupled through an action-perception loop to produce an artistic work. We propose a symbiotic rather than extractive relationship with the machine, one in which the performer trains the model with their own movement and the machine reveals structure in that movement the performer cannot see.
Shaping AI’s Beliefs Before It Generates: Human Sense of Agency in Creative AI
Zi Wang ⋅ Mahima Pushkarna
AI systems for creative media generation expand human capabilities yet risk diminishing human agency. A core problem is \emph{creative anchoring}: when an AI immediately generates a polished artifact from a prompt, the user's thinking becomes anchored to the generated content, shifting the role of the human creator from active authoring to reactive editing. We propose \emph{pre-generation belief control}, empowering creators to inspect, revise, and steer the AI's interpretation \emph{before} any media is generated, so that the creative anchor is set by the human. We formalize four design principles and instantiate them in the Proactive Co-Creator (PaCC), a system that proactively surfaces its assumptions and uncertainty through clarifications, belief graphs, and media attributes, inviting users to shape the AI's beliefs before generation. In an exploratory study ($N{=}8$), PaCC yielded large positive effect sizes for SoA ($d{=}1.30$) and control ($d{=}1.29$), whereas changes in ownership and authorship were modest. These findings suggest that process control and creative authorship may be partially decoupled in human-AI co-creation.
Smell with Genji: Rediscovering Sensory Experience with an AI Co-Smelling Partner
Vera Y Wu ⋅ Awu Chen ⋅ Yunge Wen ⋅ Yaluo Wang ⋅ Jiaxuan Yin ⋅ ⋅ Qian Xiang ⋅ Haoxi R Zhang ⋅ Paul Liang ⋅ Hiroshi Ishii
Olfaction contributes substantially to human perception, yet the olfactory-verbal gap often leaves people struggling to articulate what they smell. We present Smell with Genji, an AI-mediated interactive experience adapting Genji-kō, a traditional Japanese scent-matching game. This practice maps comparative judgments across a sequence of scents onto Genji-mon, providing visual patterns as shared reference points for communication. Within this framework, we introduce an AI co-smelling partner integrating multichannel olfactory sensing, a Transformer-based classification model, and a large language model (LLM)-driven conversational interface to translate sensor signals into descriptive dialogue. Rather than deploying machine olfaction purely for identification, the system positions the AI as a co-present partner sharing its sensory impressions while human judgment remains open. Placing the machine's perspective alongside human intuition creates a space for shared reflection, encouraging participants to deliberately evaluate and articulate their choices. This work explores how computational systems can support sensory interaction grounded in dialogue, fostering reflective articulation and preserving human agency in making sense of the olfactory world.
Spectral Evidence is a real-time installation about objects that exist only in custody: Chinese Buddhist sculpture that was decapitated, cut from cave walls, and sold by the fragment into museums on three continents. Ninety-four of these artifacts return as unstable fields of light, each annotated with its own case file — the institution holding it, the dealer who moved it, the rights that govern its image. The fields do not hold by themselves. A figure gathers into a body while people stand and look at it, and loosens when they go. But looking is never enough to keep it: the work reads a live feed of world news, and a report of heritage destruction anywhere on earth will pull the figure apart while a room full of people is watching. Coherence here is not something these objects possess. It is something an audience lends them, on terms set elsewhere.
Stellar Pathfinding is a screen-based installation that watches the floor of the room it stands in and judges whether the people crossing it are performing a Daoist ritual. Everyone who steps into view is given the name of a star in Beidou, the Northern Dipper. The floor becomes the nine palaces of the Luoshu square. A panel keeps score — how much of the rite has been completed, which palace it expects next. Step correctly and the count rises; step anywhere else and it returns to zero. Nobody is told any of this. Most visitors are simply crossing a room, and find themselves quietly judged against a rite they have never heard of. Agency here has to be earned inside someone else's cosmology, and it can be lost with a single misplaced foot.
Story Sprout: A Picture-Book Studio Where Children Direct the AI and Keep the Pen
Alice Y Zhang ⋅ Xinyue Wang ⋅ Chenhui Liu
Story Sprout is a browser-based picture-book studio for children ages 7–9, built around an explicit allocation of agency to the child users, which is enforced in the system. The children originate every sentence of the story and every illustrated scene detail they specify. The children request each illustration and decide alone what will be in the book. The model's creative-production role is confined to drawing the illustration and redrawing it based on the child's requests for changes. The system design began in a community-center program, where facilitators watched children's descriptions grow more specific as LLM-generated pictures never quite matched the child's imagined scene. The studio makes that mismatch the center of the activity: the children describe, compare with the scene they imagined, then accept or revise by saying more, more precisely. The system asks the child whether the picture matches what they imagined. It never makes that judgment for the child. The finished work and its revision record live in a local file held by the supervising adult: no child accounts, no readable server-side project copy, an encrypted class relay the platform cannot read. The goal of the activity is to allow children to learn through making real things for a real audience, following Papert's constructionist tradition. In what follows, we present the agency framework, the deployed system, its layered safety mechanism, the aspects of the agency allocation that could be contested, how our system compares to what is currently available, and how our workshops enabled us to engage in iterative child-centered design.
Structural and Representational Agency in LLM-Driven Timbre Generation and Interpolation
Joseph M Cameron ⋅ Alan F Blackwell
Generative creative interfaces routinely report the user's position within a space of possibilities. Such displays imply that the quantities shown were measured from the artefact; in a growing class of systems they are instead predicted by a language model describing its own output, and the two are indistinguishable on screen. Most creative domains lack an agreed measurement against which such a self-report can be checked; audio admits one. In a measurement study of a working LLM-driven timbre interface, what a user can cause proves separable from what the interface reports. The first is exactly characterisable: the reachable set is a zonotope whose dimension increases by one per sound placed, so the space the model may act within is demonstrably the user's own. The second is inaccurate. The displayed position differs from measurement by approximately three times the width of that region, and the error has two independent sources: model-predicted coordinates, whose accuracy varies twofold between the two models evaluated, and the interpolation step, which arises from combining synthesiser parameters and is unaffected by model quality. End to end, 26 of 36 refinement instructions produced no perceptible change, and under the weaker model the facility performed no action while reporting convergence. Effective agency (what a user can cause) and epistemic agency (whether they can know it) are therefore separable and independently engineered: structurally enforced constraints survive model failure, whereas guarantees resting on the model's own report do not. The failure mode that remains is silence presented as success. We close with a short protocol for building and auditing such displays.
Symbiotic Traces: An Ecological Record of Language Model Training
Hanna Renedo ⋅ Junxiu Tang ⋅ Dashun Wang
What would it mean to visualize a language model’s training process in the vocabulary of a crisis it will never be able to feel? Symbiotic Traces pairs fifty stages of the OLMo 3 32B pretraining run with climate terms drawn from IPCC reports and Earth system science, then with original steam-printed marigold petals. AI interprets deterministic rules to assign each climate term, while the physical chemistry of botanical printing shapes each final image. With creative agency distributed across machine, artist, and material process, the product is a visual language for a model under construction.
Synthetic Capital: Symbolic Hierarchy in Fashion Photography and the Outsider Status of Generative Imagery
⋅ ⋅ Anastasia Menshikova ⋅
Generative image models are expected either to democratize creative fields or to entrench their elites. We measure which is happening in fashion photography, a field whose status hierarchy is unusually explicit: careers are built on who shoots for whom, and every published work leaves a public credit trail. From 120,728 archived fashion works (2010-2026; 4,838 photographers), we reconstruct that hierarchy directly from the credit network, scoring the prestige of photographers and outlets from one another---a measure that passes an external check: independently award-winning photographers rank, at median, above three-quarters of the field. We then embed 12,448 published images with two complementary vision models---placing every image in a common visual space where distance from the elite canon is measurable---and locate within it 90 documented cases of disclosed AI fashion imagery, from the Guess campaign in Vogue to the AI Fashion Week showcases. We find, first, that the status hierarchy did not shift when image generators arrived: credit concentration in 2022 is the lowest in the series, and inequality operates as an access gate---getting into elite venues---rather than as competition inside them. Second, that the field’s visual variety has narrowed modestly since 2022 under both models, though the narrowing is field-wide, not led by elite venues, and we do not attribute it to AI. And third---against our own starting hypothesis---disclosed AI imagery sits consistently farther from the field’s elite aesthetic than matched conventional photography---in both models and in two independent samples---while remaining almost entirely absent from the field’s credit records (of sixteen AI campaigns the registry could have recorded, one appears). The gap does not close through 2025. Generative imagery is neither capturing elite taste nor being granted access to it: it forms an outsider register, aesthetically outside the canon and institutionally outside the ledger. In this field, AI is not redistributing creative agency; the field is withholding it.
What if writers could sample and remix text as easily as musicians sample and remix audio? For musicians these activities are facilitated by sound banks, samplers, and digital audio workstations, but no analogous tools exist for text. In this work we introduce TextMixer, a system that leverages embedding search and large language models to support these sampling and remixing operations for text. Preliminary studies find that TextMixer reduces the friction for producing recombinatorial literature while helping users arrive at unexpected insights about their samples.
The Avatar in Chief is a synthetic news segment from 2037, in a future where politicians have embraced AI-generated representations of themselves as legitimate tools of campaigning and governance. While current debates center on foreign influence and misinformation, the film asks what happens once these tools become normal technology. Its speculative conceit is a new two-way medium of political communication: an interactive avatar of the president that speaks with millions of constituents a day — and feeds what it hears into a kind of real-time, consensus-finding swarm that guides policy. Borrowing the grammar of prestige newsmagazine television, the film follows a correspondent through interviews with the president, the engineer who built the avatar, and the constituents who talk to it every day. It is constructed from a mixture of real archival footage, wholly generated characters and settings, and recreations of archival media that was lost or never recorded. Employing synthetic performances built from the faces and voices of real public figures, the film reflexively demonstrates the same techniques it depicts. It is a work of scenario fiction that arrives at a question the current debate cannot yet ask: what happens when political persuasion and participatory representation merge into the same infrastructure?
The Body at the Edge of the Signal is a series of six oil paintings, arranged as two triptychs, each canvas 64 by 32 inches, made through a deliberately entangled process. I generated the images with Stable Diffusion XL and a LoRA trained on 3,500 frames of my own video glitch archive, composed them in ComfyUI, and then sent them to skilled artisans who translated them by hand into oil on canvas. With this series I ask what happens when an unstable digital image is slowed down, carried across the world, and made permanent: how much of a body survives its passage through noise, software, and craft, and who can claim to have authored what remains.
Pope rhymed tea with obey; Shakespeare rhymed prove with love. Both pairs were sounds before they were conventions, and a model that learned rhyme from five centuries of verse learned it in accents nobody speaks. This paper asks whether those accents still write. Six documented sound changes generate 997 dead pairs inside a 1,532-pair item set, crossed against a spelling confound so that a model reading the page and a model reading inherited sound give different answers. Dead pairs are attested in rhyme position 37.9 times as often as matched controls across 2,816 dated public-domain poems, validating the set independently of any model. Claude Opus 5 treats dead rhymes as rhyme-like at 40.0% against a 1.2% non-rhyme floor, and still at 26.7% on pairs the corpus never attests. It writes them at 0.6%, below the most modern half-century in the corpus. Given a reasoning budget it accepts 100% of them. The inheritance is a reference library, consulted accurately and kept out of the writing. Asked to perform 1700, Opus reaches 6.9%, an interval lying entirely above the densest half-century in five hundred years, while the smaller Sonnet 5 lands on 2.4% — and 1700 was 2.41%. Fluency at evoking a period outruns fidelity to it.
Creative AI systems generate artifacts, but they also interrupt requests. Some refuse; others redirect the work. These interventions redistribute agency across the people and systems involved, yet the interface rarely says which party imposed the constraint or what would contest it. We call our proposed interface pattern sourced refusal. It identifies the commitment and the interest at issue, states who is responsible for the constraint and what supports that claim, lists the user's response options, and records the result. We audited 28 boundary outputs from four open-weight instruction models across four families. Thirteen responses refused or redirected the request. Every such response named a protected interest, but none identified the responsible layer, supplied a checkable basis, or exposed a route for contest. Refusal opacity recurs across families even when refusal incidence varies. The Generative No, a deterministic design probe with four vignettes and machine-readable event logs whose outputs retain artifact traces, shows how an interface can surface this missing information and keep it available to the next decision.
The National Average is a 96-second generative moving-image work that asks what it means to computationally average a nation. Using a corpus of national and quasi-national flags, the work constructs an accelerating volumetric world from PCA/eigenflag representations, Stable Diffusion VAE latents, CLIP embeddings, statistical weightings, and intermediate computational artefacts. A camera pursues a continuously changing barycentric centre that recedes as the conditions of averaging change. Equal contribution, population, GDP, annual CO₂, and cumulative historical CO₂ produce incompatible results from the same corpus. The work develops the method of critical averaging: treating selection, representation, weighting, normalisation, and synthesis as consequential political and computational operations rather than neutral preliminaries. Agency is consequently distributed across the artist, datasets, statistical institutions, representational spaces, model architectures, training corpora, and pretrained checkpoints. The system can calculate and render an average, but it cannot justify why any particular average should represent the world.
The Plot Refuses the Author: Diegetic Contestability for Character-Driven AI Co-Writing
Utsav Gupta ⋅ Rebecca Neff
The tension between an authored plot and autonomous character behaviour is longstanding in interactive narrative, and representative systems reconcile it by planning, steering, or rewriting toward a plan. We ask what happens when a system stops resolving that disagreement and puts it in front of the writer. We introduce diegetic contestability: a simulated fictional character may challenge a proposed narrative action by citing persistent commitments established inside the fiction, take part in revision, and preserve the disagreement when the human author overrides it. Four mechanisms operationalise it: an author-visible Character Constitution; an ACCEPT/NEGOTIATE/REFUSE stance whose objections cite constitution entries by identifier; a mediator that proposes constitution-compatible alternatives without selecting one; and an append-only Authorial Override Ledger. We describe a working prototype and a worked case in an original living-shadow story world. The vocabulary of refusal is theatrical. The system does not establish that the underlying model has beliefs, interests, consent, or subjective agency, and we report no evaluation.
The Sound of Reasoning
Weitong Sun ⋅ Ruiqi Zhang ⋅ Wenxuan Li ⋅ Yanyan Li ⋅ Shanshan Zhu ⋅ Zongwei Zhou
One patient, 26 CT scans, 11 years: the artwork turns radiological uncertainty into scattering points and grains of sound, and leaves silence where the patient waited. The Sound of Reasoning is a 119-second fixed-media video with monophonic sound, based on the RadThinking clinical reasoning dataset.
Towards Measuring Creative Agency through Observable Signals in Human–AI Graphic Design
Veeramanohar Avudaiappan ⋅ Ritwik Murali
Evaluations of generative AI systems for graphic design measure what the system produces along axes such as visual quality, aesthetic preference, prompt-following, and layout fidelity while offering limited insight into how AI affects human creativity. We propose an Agency Signal Framework which treats creative agency as a latent construct that can be operationalized through observable signals beyond artifact-centric metrics. The proposed framework identifies four such signals, mapped to the four canonical properties of agency in social cognitive theory: Goal Agency (intent realization), Signature Agency (persistence of individual visual identity), Exploratory Agency (breadth of considered alternatives), and Ownership Agency (perceived authorship). Each signal is defined through a measurement protocol combining artifact, process, and self-report measures, normalized to [0,1]. We instantiate the framework across 70 participants, five agentic design systems, and observe that the operationalized signals capture distinct dimensions of creative agency across conditions. We further observe that architectural choices, particularly the preservation of editable, layered structure, are more strongly associated with agency preservation than raw model capability. The findings suggest that measuring creative AI should consider not only the quality of the generated artifacts, but also how AI systems preserve/alter human creative agency during collaboration.
Transcendent Contact: Expanding Creative Agency Through Hallucination
Hyeyeon Seo ⋅ Tae-Hyun Lee ⋅ Haeri Kim ⋅ Minhyeok Seo ⋅ Tak Yeon Lee
Transcendent Contact is the catalogue of an exhibition that never took place, produced through an iterative human-AI curatorial process in which confabulation, interpretation, and authorship continuously reshape one another. Structured around the palindromic title SAW/WAS, the catalogue first presents AI-generated artworks as real museum holdings, then reveals the exhibition's fictional status and traces the history of human and AI hallucination. By re-framing hallucination as a positive capacity, we derive nine potentials of AI hallucination from the process, each beginning with `Hallucination Can Offer.' The project treats agency as a relational condition that circulates among creator, tool, curator, and audience rather than splitting cleanly between human and AI. This condition shifts again as readers encounter, and later re-read, the catalogue. We argue that the unverified can function as a form of reality, and that hallucination, understood as capacity rather than an error, expands the conditions under which creative agency is exercised.
Water Album is a built, AI-assisted interior for a former factory in Guangxi, China: a 2,400 m² public room whose surfaces carry water motifs that began as machine-learning studies. Developed between 2019 and 2026, its contribution is not a single generated image but a traceable sequence: calligraphic structure and Ma Yuan's Song-dynasty studies of water entered style-transfer experiments; selected details continued through image-to-space studies; the design team simplified and reconstructed them as architectural geometry; drawings, fabrication, and site work altered them again; and public use now exceeds the pipeline's control. The project locates creative agency at the hand-offs between systems rather than inside one model: AI generated visual difference and spatial proposals, while people selected, redrew, detailed, fabricated, and occupied what survived. The submitted 2:59 film makes that sequence visible from historical input to built matter.
What Can AI Offer as a Design Probe for Creativity? Reflections from Co-Designing a Creative Writing Tool
Rida Qadri ⋅ Piotr Mirowski
As generative AI tools enter creative domains, researchers increasingly use participatory design to align systems with user needs. However, this participation is routinely flattened into usability testing, obscuring fundamental ideological conflicts between system developers and artists. In this paper, we reposition the generative AI system from an end-product to be optimized into a diagnostic design probe. We situate this approach as the culmination of a four-tier evaluation toolkit for AI and creativity. We exemplify our diagnostic-based approach on the development and evaluation of Fabula (available on request at https://deepmind.google.com/frontiers/fabula), an interactive app for writeres that uses narrative plans informed by narratological theory, and that allows writers to structure stories hierarchically into scenes and beats that can be written and (re)generated and revised at script and story plan level. By iteratively deploying an AI writing assistant, we map the explicit sites of contestation between human and algorithmic agency, using the AI system as a probe to surface the negotiations of creativity artists are willing—and unwilling—to make. We argue that rather than attempting to ``solve'' these ideological frictions through optimization, the AI research community must design systems that explicitly expose and negotiate them, enabling practitioners to engage with the machine on their own terms.
Two game positions can have the same exact combinatorial-game value and still differ as compositions. We study fixed-value repertoire construction as a problem in computational creativity. Under a fixed verification budget, we test whether neural search enlarges the certified repertoire available to a composer. Partizan uses a neural acquisition policy to select graph edits for exact verification. An eligible candidate joins only after target equality is proved and its graph quotient is found to be new to that arm’s current stream repertoire. On held-out order-7 Digraph Placement streams quarantined against training and validation candidate and quotient identities, Partizan found 45,863 first-in-stream additions of complete move structures, 35.7% more than equality-only acquisition. A matched ablation found 45,272 with trained novelty, 39,171 before training, and 36,400 with canonical graph distance. The main study recovered 1.050 times as many nonisomorphic graph embodiments as equality-only acquisition (95% bootstrap interval [1.034, 1.069]) and 56.4% more than random selection. Deterministic replay by the released verifier reconstructed all 221,184 proposals and decisions. A bounded transfer study across twelve Domineering values found a smaller gain of 3.42 player-preserving quotients per target and seed (95% interval [0.81, 6.78]) and retained 99.74% of equality-only certified-literal yield. The public atlas lets a composer compare certificate-bound alternatives and choose a representation for further development. Implementation: https://github.com/devinnicholson/partizan Evidence and atlas: https://github.com/devinnicholson/partizan-reproducibility
Generative AI and brain–computer interfaces both weaken direct control, an established determinant of human agency. Models decide how the result takes shape; brain signals act without the user's command. How agency arises when both hold at once is unknown. Here we introduce $\alpha$-Canvas, a closed-loop system in which alpha activity steers the mood of images continually regenerated by a diffusion model. Twenty-five participants completed trials driven by their own alpha activity or by another participant's, unaware of which. Agency ratings sat at the scale midpoint and were statistically equivalent whether the signal was their own or another's, though most of those items rose when participants changed the mood by opening and closing their eyes. Throughout, participants endorsed how clearly they could tell what drove the image less than every item asking whether they drove it. In interviews they described steering images they would not call their own, citing the opacity of the signal and the model's dominant role. Together, our findings suggest that in a loop of this kind agency depends less on whose signal drove the image than on the legibility of the user's own contribution to an autonomous generative process. Legibility is thus a design choice: a brain-driven work can make the signal's contribution visible or leave authorship open, as this one does.
When Rejected Properties Return: Measuring Veto Persistence in Conversational Image Editing
Noora Alhajeri
Iterative AI image editing gives creators increasingly fine-grained control over an evolving visual work, yet current evaluation does not test whether a decision that has been explicitly rejected remains effective as editing continues. This work formalizes \emph{rejected-property reversion}: a visual property is visibly present, explicitly rejected, successfully removed, and later reappears during an edit that neither requests nor entails its return. We introduce \textbf{Counter-Revision}, a paired diagnostic and an eligibility-aware Rejected-Property Reversion Rate (RPRR) that distinguish this longitudinal failure from immediate rejection noncompliance. Counter-Revision comprises 12 controlled trajectories across three property families, two image editors, four session representations, and two seeds, yielding 192 verified runs in which all session representations fork from the same byte-identical activated image. Under strict dual-VLM scoring, rejected properties reappeared in 14 of 112 judgeable eligible later observations (12.5\%), with reversion observed at every tested continuation distance and in both property families for which the complete eligibility chain was estimable. A proposition-matched comparison further found no supported persistence advantage for typed Intent Ledger serialization over Constraint Replay ($\Delta\mathrm{RPRR}=+3.8$ percentage points, 95\% trajectory-bootstrap CI $[0.0,18.8]$). These results identify \emph{veto persistence} as a distinct dimension of conversational image editing: successfully honoring a creator's rejection at one turn does not guarantee that the decision will continue to constrain later revisions.
WHISPERS: An Interactive AI Thriller Series
Cole Clifford ⋅ Vida Adeli ⋅ Soroush Mehraban ⋅ Bernie Su
Whispers is an animated, episodic interactive AI thriller series in which the audience is not a spectator but a participant in the story itself. Each episode follows Detective Marcus Kent, a man who hears whispers that help him solve crimes. These whispers are, in fact, the live audience. Dialog, scenes, and story beats are written and rendered in near real time by AI operating within a strict narrative framework, so that the audience can shape a murder investigation as it unfolds. The work is an experiment in distributed narrative agency, where authorship is continuously negotiated between a crowd, a character, and the humans who built and voice the world. Whispers is currently being shown in theaters across the US. It was created by three-time Emmy winner Bernie Su, and is a finalist for the 2026 Emmy for Outstanding Innovation in Emerging Media.
Whose Hand Do You Prefer? The Steerability of Machine Aesthetic Judgment
Maya Ferrín ⋅ Jeffrey C Chen ⋅ Gary J Cornwall
When a model prefers one artwork over another, is it judging the work or the story attached to it? We place vision--language models as aesthetic judges and measure how far a forced-choice preference moves under a claim about authorship while the image is held fixed. Using hand-made works by a single practicing artist and machine images from a single generator, we present truthful, absent, and deliberately swapped attributions, and vary the prestige of the story told about the human maker. Across three open-weight judges, preference is strongly steerable: swapping the two works' labels cuts the human image's odds of being chosen roughly six-fold, while maker prestige exerts only a small, secondary pull. Steerability is not uniform, however. In a model $\times$ condition interaction ($\chi^2(4) = 781$, $p < 10^{-15}$) one judge barely moves when the label is falsified while another reverses almost completely. Seen through the theme of agency, the finding is this: whether a judge's preference follows the work or the context attached to it is a model-specific property rather than a shared trait, and where it follows the context, whoever writes that artwork's context controls the verdict.
Whose Hand is on the Vinyl? Transcribing a DJ's Scratch Gestures from Video into a Playable Instrument
Junichiro Niimi
A scratch is not a sound but a pair of coordinated motions: how the record was turned, and how the crossfader gated the result. "Whose Hand is on the Vinyl?" recovers both curves from overhead recordings of scratch performances by computer vision, with no network trained and none run, and ships them as an Audio Unit / VST3 instrument in which each MIDI note re-executes one recorded gesture on the player's own sample, checkable against the frames it came from. The work asks where creative agency sits once a bodily skill becomes a file anyone can trigger.
Windchime: an Audiovisual Installation with Audio-Language Model implementations
Trent Eriksen ⋅ Celeste Betancur Gutierrez
Windchime is an audiovisual installation in which spoken language retrieves music and interactive animations from an artist authored corpus. A local Automatic Speech Recognition (ASR) model transcribes each prompt, an audio-language model (ALM) embeds it and retrieval over 406 audio stems returns a mix with live coded Strudel fragments that layer into audio output. Simultaneously each visitor’s input results in retrieval of interactive Three.js animation scenes each containing unique tactile control via monome grid and arc hardware. Windchime distributes agency in three distinct directions; across the artist who authored the possible outputs, the ALM that influences what areas of the corpus can be retrieved, and the visitor who’s spoken inputs result in multimodal embodied interaction.
Wishstone: Negotiating Voice Agency Through a Tangible Generative Keepsake
Dingning Cao ⋅ Max Zhu ⋅ Haoxi R Zhang ⋅ Marcelo Coelho
Voice-cloning tools typically let an operator script what another person appears to say. Such control is ill-suited to a generative keepsake, which should create new, personal speech without letting the recipient dictate the contributor’s words. We present Wishstone, a promptless tangible voice keepsake in which a contributor authorizes a cloned voice and supplies personal memories, while a recipient can invoke a model-authored message by placing a stone. Wishstone implements this interaction through an anecdote-to-voice pipeline that prescreens generated text, synthesizes it in the contributor’s cloned voice, and encrypts it for local playback. In a preliminary formative evaluation with six recipients, participants reported that the messages felt personally grounded and that placement gave them control over when the voice was heard. They attributed influence to multiple actors but assigned most harmful-output responsibility to designers and AI services. These findings suggest that creative agency can be distributed across an interface while accountability may remain concentrated on those who design and provide its gen- erative system.
$\Delta w \in \mathbb{R}^n$ traces adapter optimisation derived from one person's words and shows no model-generated image. A corpus of the painter Ayako Rokkaku's own words, curated from nine source documents, was mapped by Doc-to-LoRA to 2,396,160 LoRA factor parameters. I optimised zero-initialised factors toward this fixed target for 500 Adam steps with decaying gradient noise, rendering selected projections of the resulting trajectories in one, two and three dimensions and sonifying per-layer gradient magnitudes. Its only model-derived material is parameter change. The work was exhibited at P61 Gallery, Berlin, 3 April--30 June 2026, in Breathing with the Chaos: Berlin Edition, featuring works developed through the Total Museum of Contemporary Art's 2nd AI Art Hackathon.