Paper deep dive
Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation
Zoe De Simone, Angie Boggust, Fredo Durand, Ashia Wilson, Arvind Satyanarayan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/18/2026, 1:31:24 AM
Summary
Creo is a multi-stage text-to-image (T2I) system designed to align generative AI workflows with human creative ideation. By decomposing image generation into five distinct stagesâviewpoint, composition, color, lighting, and styleâCreo allows users to progressively refine visual ideas using intermediate, sketch-like representations. This approach mitigates common T2I issues like premature anchoring, unintended editing drift, and lack of user control, ultimately fostering greater user agency and output diversity.
Entities (4)
Relation Signals (3)
Creo â implements â Multi-stage generation
confidence 100% ¡ Creo is a multi-stage T2I system that scaffolds image generation by progressing from rough sketches to high-resolution outputs
Creo â improves â User Agency
confidence 95% ¡ Creo improved usersâ sense of control, authorship, and ability to iteratively refine ideas.
Creo â mitigates â Anchoring Effect
confidence 90% ¡ these findings suggest that multi-stage generation... is a key design principle for... mitigating well-known behavioral biases such as anchoring effects
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate outputs that make implicit visual decisions on behalf of the user, often introduce fine-grained details that can anchor users prematurely and limit their ability to keep options open early on, and cause unintended changes during editing that are difficult to correct and reduce users' sense of control. To address these concerns, we present Creo, a multi-stage T2I system that scaffolds image generation by progressing from rough sketches to high-resolution outputs, exposing intermediary abstractions where users can make incremental changes. Sketch-like abstractions invite user editing and allow users to keep design options open when ideas are still forming due to their provisional nature. Each stage in Creo can be modified with manual changes and AI-assisted operations, enabling fine-grained, step-wise control through a locking mechanism that preserves prior decisions so subsequent edits affect only specified regions or attributes. Users remain in the loop, making and verifying decisions across stages, while the system applies diffs instead of regenerating full images, reducing drift as fidelity increases. A comparative study with a one-shot baseline shows that participants felt stronger ownership over Creo outputs, as they were able to trace their decisions in building up the image. Furthermore, embedding-based analysis indicates that Creo outputs are less homogeneous than one-shot results. These findings suggest that multi-stage generation, combined with intermediate control and decision locking, is a key design principle for improving controllability, user agency, creativity, and output diversity in generative systems.
Tags
Links
- Source: https://arxiv.org/abs/2604.13956v1
- Canonical: https://arxiv.org/abs/2604.13956v1
Trouble viewing inline? Open PDF directly â
Full Text
90,769 characters extracted from source content.
Expand or collapse full text
Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation Zoe De Simone MIT CSAILCambridgeMAUSA zoed@mit.edu , Angie Boggust MIT CSAILCambridgeMAUSA , Fredo Durand MIT CSAILCambridgeMAUSA , Ashia Wilson MIT CSAILCambridgeMAUSA and Arvind Satyanarayan MIT CSAILCambridgeMAUSA Abstract. Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I generate outputs which make implicit visual decisions on behalf of the user, often introduce fine-grained details which can anchor the user prematurely, limiting usersâ ability to keep options open early on, and cause unintended changes during editing that are difficult to correct and reduce usersâ sense of control. To address these concerns, we present Creo, a multi-stage T2I system that scaffolds image generation by progressing from rough sketches to high-resolution outputs, exposing intermediary abstractions where users can make incremental changes. Sketch-like abstractions invite user editing and allow to keep design options open when ideas are still forming, due to their provisional nature. Each stage in Creo can be modified with manual changes and AI-assisted operations, enabling fine-grained and stepwise control through a locking mechanism that preserves prior decisions so subsequent edits affect only specified regions or attributes. Users remain in the loop, making and verifying decisions across stages, while the system applies diffs rather than regenerating full images, reducing drift as fidelity increases. A comparative study with a one-shot baseline shows that participants felt stronger ownership over Creo outputs, as they were able to trace their decisions in building up the image. Furthermore, embedding-based analysis indicates that Creo outputs are less homogeneous than one-shot results. These findings suggest that multi-stage generation, combined with intermediate control and decision locking, is a key design principle for improving controllability, user agency and creativity, and output diversity in generative systems. human-AI collaboration; generative AI; text-to-image models; creativity support â copyright: noneâ conference: ; ; â ccs: Computing methodologies Artificial intelligenceâ ccs: Human-centered computing Interaction design Figure 1. Creo is a multi-stage image generation workflow. Unlike current text-to-image systems, Creo starts from a rough sketch, allowing users to progressively make visual decisions, like viewpoint, composition, color, lighting, and style. 1. Introduction Text-to-image (T2I) systems enable users to rapidly generate high-fidelity images from minimal prompts, lowering the barrier to image creation (Ramesh et al., 2022; Saharia et al., 2022; Rombach et al., 2022). However, most systems generate fully rendered images in a single step, making implicit visual decisions about composition, lighting, and style before the userâs intent has fully formed. These decisions can prematurely anchor users to details they have not yet considered and make edits feel like repairing model-made choices rather than developing oneâs own ideas. This interaction paradigm contrasts with how creative ideas typically develop. Creative work rarely begins with a fully specified vision (Bigelow et al., 2023). Instead, creators progressively refine structure and appearance over time. Many visual workflows explicitly support this progression. For example, artists in animation and film begin with storyboards and rough pre-visualizations before committing to lighting and rendering (Thomas and Johnston, 1995; Okun and Zwerman, 2010; Parent, 2012), while comics artists move from thumbnails to penciling, inking, and coloring (McCloud, 1994; Eisner, 2008; Cohn, 2013). These staged representations allow creators to focus on particular decisions while postponing others, enabling broad exploration without prematurely committing to fine-grained details. In contrast, single-shot T2I systems collapse these decisions into a single step, forcing early commitment and limiting exploratory control. Together, this mismatch suggests a need for generative systems that support progressive ideation and intent formation, allowing users to develop and refine visual decisions over time rather than specifying outcomes upfront. Prior work on controllable generation has expanded what aspects of an image users can specifyâfor example through structural guidance such as edges, poses, or layouts (e.g., ControlNet (Zhang et al., 2023)). However, these approaches often still introduce highly detailed imagery early in the process. In contrast, we argue that improving controllability requires structuring when and how visual decisions are introduced during image creation. To address this, we present Creo, a T2I system that scaffolds image generation as a multi-stage process, progressing from rough sketches to high-resolution outputs. Rather than generating a fully rendered image immediately, Creo decomposes image creation into a sequence of stages over which users retain control. At each stage, users interact with sketch-like intermediate representations that are intentionally under-specified and unfinished, inviting user input and keeping design options open while ideas are still forming. As these representations evolve, they become progressively more detailed, allowing users to first establish high-level aspects such as viewpoint and composition, and then refine lower-level properties such as color, lighting, and style. Creo provides direct manipulation and AI-assisted tools that enable fine-grained control over specific aspects of the image. Users can edit intermediate representations by drawing, erasing, or masking regions and describing localized changes, allowing targeted modifications without affecting unrelated elements. Through this staged interaction, users remain actively involved across the generation process, making and verifying decisions as the image progressively develops rather than reacting to fully specified outputs. To support stable iteration, the system includes a locking mechanism that preserves prior decisions so that subsequent edits apply only to selected regions or attributes. Instead of regenerating entire images, Creo applies incremental, diff-based updates, reducing unintended changes and drift as fidelity increases. We evaluate Creo through a comparative study against a conventional one-shot generation workflow. Participants reported a stronger sense of ownership over images created with Creo, as they were able to trace their decisions in building up the image. Participants engaged with the full range of stages and both direct and AI-assisted tools, despite this not being required, indicating that intermediate stages were not superfluous but provided distinct and necessary support for different aspects of the workflow. Additionally, embedding-based analysis shows that outputs from Creo are less homogeneous than those produced by the baseline, suggesting broader exploration of the design space. While baseline systems were often rated higher in immediate visual quality, Creo improved usersâ sense of control, authorship, and ability to iteratively refine ideas. Together, these findings suggest that structuring generation as a multi-stage processâcombined with editable intermediate representations and mechanisms for preserving decisionsâis a key design principle for improving controllability, user agency, and output diversity in generative systems. More broadly, this work highlights the importance of supporting progressive intent formation, where systems help users develop ideas over time rather than requiring them to specify outcomes upfront. By aligning generative workflows with the iterative nature of creative thinking, such systems can mitigate well-known behavioral biases such as anchoring effects and premature commitment (Green, 1989; Jansson and Smith, 1991; Wadinambiarachchi et al., 2024; Doshi and Hauser, 2024), enabling more exploratory and user-driven ideation. 2. Background and Design Goals Prior research across cognitive science, creativity support tools, and generative image systems points to a mismatch between how people form visual ideas and how current T2I workflows generate images. In creative practice, ideas often develop gradually through intermediate representations that help creators refine one aspect of an image at a time. For example, artists will isolate and test different color palettes on small regions of an image prior to choosing a direction for the entire image. In contrast, many T2I systems produce fully rendered outputs early, collapse multiple visual decisions into a single generation step, and rely heavily on prompt-based interaction to control visual decisions. 2.1. Human Requirements of Visual Ideation Early visual ideation is intentionally underspecified. In creative practices, early visual ideas are often deliberately underspecified. Cognitive science research shows that people imagine scenes without committing to attributes such as color, texture, or background until later in the process (Bigelow et al., 2023). This ambiguity is beneficial because it allows multiple candidate interpretations of a scene to remain viable until later design decisions constrain them and helps preserve flexibility while ideas are still forming. Sketching makes this process visible. Drawing externalizes evolving thought (Fan et al., 2023), and early sketches are treated as exploratory (Purcell and Gero, 1998). Classic accounts describe sketching as a reflective process of seeing, interpreting, and revising (Tversky, 2002; Goldschmidt, 2014; SchĂśn, 1983). Line drawings preserve structure while suppressing detail, making them well suited for reasoning about form without committing to appearance (Hertzmann, 2020; Livingstone and Hubel, 2002). Together, these findings suggest that ideation benefits from representations that preserve ambiguity. Intermediate representations support dimension-specific reasoning. Creative workflows typically progress through representations that isolate specific types of decisions. For example, artists establish composition before refining lighting or color (Dow, 1914; Loomis, 1947), and workflows in animation, film, and comics similarly move from structure to appearance (Thomas and Johnston, 1995; Okun and Zwerman, 2010; Parent, 2012; McCloud, 1994; Eisner, 2008; Cohn, 2013). This ordering reflects dependencies between decisions, as later properties rely on a stable structural foundation (Purcell and Gero, 1998; Verstijnen et al., 1998; Brun et al., 2016). Intermediate representations also support attention and reasoning by limiting what is being decided at a given time (Goodwin, ). Creativity support tools have long leveraged similar strategies, enabling users to begin with coarse sketches or low-fidelity wire frames and progressively refine them (Igarashi et al., 1999; Bae et al., 2008; Swearngin et al., 2018; OâDonovan et al., 2015; Iarussi et al., 2013; Fernquist et al., 2011). Recent systems such as DrawMyPhoto (Williford et al., 2019) and ArtKrit (Ma et al., 2025) externalize dimensions such as composition, value, and color to support focused reasoning. Drawing assistance tools guide contour structure or proportion without requiring fully rendered outputs (Iarussi et al., 2013; Fernquist et al., 2011). Other tools provide abstractions for working with color and tone, including histogram views (Chevalier et al., 2012) and direct-manipulation palette systems (Shugrina et al., 2019, 2017). Across these systems, intermediate representations act as cognitive scaffolds that support incremental, focused reasoning. This work suggests that visual ideation is inherently staged: creators progressively and independently refine image dimensions. 2.2. Challenges in T2I Model Workflows Early specificity shifts ideation toward reactive editing and constrains exploration. Most T2I systems generate fully specified, high-fidelity images from the outset. Even minimal prompts can produce detailed commitments to lighting, texture, background, and style. While visually compelling, these outputs introduce many decisions that users may not have intended or even considered (Hopkins et al., 2025), often before their intent has stabilized. As a result, interaction becomes reactive. Users typically iterate by modifying prompts and regenerating full images, but each iteration still presents a complete, highly detailed output. Rather than developing ideas through intermediate representations, users adapt their thinking in response to model-produced artifacts. Early exposure to polished outputs can also create a false sense of completion, where users accept compelling images despite misalignment with higher-level goals such as composition or viewpoint. Correcting these discrepancies often requires localized edits like masking, inpainting, or prompt adjustments. This shifts the creative process away from concept formation toward repairing model-generated artifacts. Prior work describes this dynamic as a âgulf of envisioning,â where users must specify details they cannot yet imagine and interpret unpredictable changes in generated outputs (Agrawala, 2023; Subramonyam et al., 2024). Early specificity also shapes exploration through well-established cognitive mechanisms. Design fixation shows that exposure to example solutions can anchor designers on particular features, even when those features are incidental (Jansson and Smith, 1991). Generative model outputs can play a similar role: users may begin ideation by reacting to generated images rather than constructing ideas incrementally. Recent work suggests that this can reduce divergent thinking (Wadinambiarachchi et al., 2024) and lead to more homogeneous outputs across users (Doshi and Hauser, 2024). Current interaction patterns amplify these effects. Interfaces often present grids of polished outputs or support prompt refinement (Feng et al., 2023; Brade et al., 2023), encouraging comparison between finished artifacts rather than early exploration. Hybrid systems that combine sketches with high-fidelity generation may similarly introduce detailed imagery before structural decisions are fully developed (Zhang et al., 2023; Sarukkai et al., 2024). Together, these findings reveal a mismatch between the timing of model-generated detail and the development of user intent. When highly specific visual commitments are introduced too early, they can anchor exploration, shift interaction toward reactive editing, and reduce diversity of outcomes. This suggests that generative systems should align the level of visual specificity they produce with the userâs stage of ideation. Visual decisions are entangled in current interaction paradigms. In most T2I systems, decisions such as viewpoint, composition, color, lighting, and style are resolved simultaneously within a single generation step. Because these attributes are entangled, modifying one aspect often unintentionally affects others. For example, changing lighting conditions may alter color, mood, or object appearance, requiring repeated regeneration. Prompt-based interaction further complicates this process. Users must translate evolving visual ideas into explicit language, even when intent is still tacit (SchĂśn, 1983; Tversky, 2002). Recent work on controllable generation expands the range of attributes users can specify, for example through structural guidance or localized editing (Zhang et al., 2023; Sarukkai et al., 2024; Voynov et al., 2023; Tang et al., 2025; Lee et al., 2025; Alzayer et al., 2025; Zhang et al., 2025; Shi et al., 2026). However, these approaches primarily address what can be controlled, while less often addressing when different visual decisions should be introduced. 2.3. Design Goals These observations point to a consistent mismatch: human ideation proceeds through staged, partially specified representations, while current T2I systems introduce detailed, entangled decisions early. From this mismatch we distill three design goals (DG1 â-DG3), which inform the design of Creo described in Section 3. DG1: Introduce visual detail progressively. Generative systems should align the level of detail they produce with the userâs stage of ideation, enabling early exploration before committing to fine-grained appearance, and without prematurely committing a user to details they have not yet reasoned about. DG2: Decompose image creation into separable decisions. Inspired by artistic practices, generative systems should structure image creation around distinct visual dimensions (e.g., composition, lighting). This allows users to incrementally build up their image as intent forms, refining one aspect without unintentionally altering others. Such separation enables localized reasoning and editing while preserving previously made decisions. DG3: Support interaction through representations accessible to both humans and models that invite editing. Generative interfaces should employ intermediate visual representations that are both intuitive for users to manipulate and meaningful for models to condition on. Rather than requiring users to edit fully rendered images or specify pixel-level changes through prompts, systems should expose abstractions that allow users to control one aspect of the image at a time and are easily accessible and editable. 3. Creo Figure 2. Creo decomposes image generation into multiple stages. From a prompt, it generates (1) multiple viewpoints, after which the illustrator (2) refines composition, (3) color, (4) lighting, and (5) style in any order. Stages can be completed in any order. To bridge the disconnect between human visual ideation and current image generation workflows, we introduce Creo (Figure 2). Creo is a text-to-image system that supports progressive ideation (DG1) by structuring image generation as a sequence of creative decisions (DG2). Rather than generating a fully specified image in a single step, Creo begins with sketches that users progressively refine into high-fidelity images (DG3). 3.1. Multi-Stage Image Generation We designed Creo to give users fine-grained control over image generation without requiring them to specify every visual property at the outset. To do so, Creo decomposes text-to-image generation into five stages: viewpoint, composition, color, lighting, and style. In each stage, users make the design decisions that are relevant to that part of the creative process, such as sketching new objects during composition or choosing a color palette. This design facilitates progressive ideation (DG2), allowing users to build up the final image in a way that mirrors the creative process, starting with low-level structural decisions and then refining higher-level aesthetic details. Importantly, Creoâs stages are independent, allowing users to explore one aspect of their imageâs design at a time. While standard text-to-image systems always output a fully rendered image (e.g., adding color even when none is specified), Creo better matches the level of detailed described in the prompt. For instance, if a user has not entered the color stage, then the image will remain in grayscale while they make decisions about lighting, keeping options like color and style open for later refinement. Creoâs stages are also additive, enabling users to progressively build a high-fidelity image. As users move through the stages, their decisions compound, and the image evolves from a black-and-white sketch to a fully rendered scene. Users work with sketches in viewpoint and composition, add hues in color, define depth in lighting, and develop an aesthetic in style. As a result, Creo allows users to work through their creative process, yielding an image that fully results from a userâs intent rather than an AI modelâs assumptions. By structuring image creation into stages, Creo isolates user edits. A central challenge in text-to-image editing is that modifying one aspect of an image often unintentionally affects others. Instead of modifying a fully rendered image and resolving unintended changes, users verify their composition, color, and lighting choices step by step. This staged process helps prevent errors from compounding as the image becomes more detailed. By decomposing image generation into stages, Creo aligns with established creative processes. The five stages are inspired by documented workflows in illustration, studio art, and architectural rendering (Szuc, 2020; Artists & Illustrators, 2021; Nicolaides, 1941; Ching, 2014), where, practitioners often establish composition before committing to color, lighting, and style. While individual creative processes vary based on the creator and their task, we distill these recurring practices into a compact set of five stages. As a result, Creoâs stages act as a minimal grammar of visual ideation, allowing users to control their creative process while leveraging the efficiency of image generation models. 3.2. Staged Representations and Tools To support creative processes, Creo defines stages that allow users to progressively introduce visual details, like lighting and color. However, progressive ideation is not only about matching the visual detail in the image to the userâs intent, but allowing users to meaningfully edit these underspecified images (DG1). To do so, Creo maps each stage (Section 3.1) to an intermediate representation. Intermediate representations reflect the types of visual design decisions a user makes in that stage. Instead of editing a rendered image, users work on representations that isolate specific aspects of the image and remove dependencies on others. For instance, the composition stage operates over black-and-white sketches, while lighting adds shading to the sketch, and only style renders a high-fidelity image. This allows users to reason directly about the decision at hand without being constrained by details that they have not yet specified. In particular, Creo uses sketch-based representations to capture structure without committing to appearance. Sketches reduce unrequested detail not by hiding it, but by not introducing it in the first place. This matches usersâ intent by removing dependencies on details that have not yet been decided, and supports progressive ideation through changes in representation rather than incremental refinement of a single image. For example, in a sketch, repositioning an object involves only adjusting simple shapes, rather than reconciling perspective, shadows, or occlusion. Moreover, since sketches are perceived as provisional they encourage users to modify and explore the representation rather than treat the image as final. Each stage and its representation (e.g., black-and-white sketch) is mapped to a set of interaction techniques that operate on that representation (e.g., lasso and move). This pairing allows users to act directly on the aspect of the image they are currently refining (DG3). While we initially explored a unified set of tools across all stages, this made interaction cumbersome as tools designed for detailed images (e.g., lighting) are difficult to apply to simplified representations and sketch-based tools do not translate well to appearance-level editing (e.g. what is the effect of brushing on lighting or how should drawing alter a high-resolution image). Instead, each stage provides tools that match its representation. For instance, in a sketch-based stage like composition, users work with tools such as drawing, erasing, and transforming shapes to define structure. As the image becomes more detailed, Creo introduces tools for applying color through brushes or regions, and for adjusting lighting through directional and intensity controls. By scoping tools to each stage, users can act directly on the aspect of the image they are currently refining, without needing to manage unrelated details. This reduces interference between decisions and keeps interaction aligned with the current level of abstraction. Within each stage, Creo combines direct manipulation and AI-assisted tools. We combine these two types of interactions to balance between giving users precise editing control and allowing them to delegate tedious edits to an AI model. As a result, users can provide partial input, such as rough sketches, color strokes, or masks, and rely on the system to complete or refine the result. For example, a user may block in approximate color regions and allow the system to propagate them into object boundaries or specify a lighting direction and let the model infer shading. This allows users to control the level of effort they invest, focusing on decisions they care about while delegating others to an AI model. To limit unintended changes, Creo supports localized edits through masking. Users can select specific regions of the image and apply changes only within those areas. Model-generated updates are composited back into the existing image, leaving the rest unchanged. This allows users to refine individual elementsâsuch as adjusting the color of an object without affecting the surrounding scene. Figure 3. Creo supports non-linear creative workflows by combining edits in earlier stages (e.g., adding a hat to the composition) with existing upstream decisions (e.g., dog color). 3.3. Interaction Workflows Although Creo is organized into stages, creative workflows are not strictly linear. Users often revisit earlier decisions after seeing later results, requiring the system to support revision without restarting. To support non-linear workflows, Creo provides a persistent preview of the final image alongside the decisions in each stage (Figure 3). Users can move between stages in any order and revise earlier decisions (Figure 9), such as coloring an image and then returning to composition to add a new object or change the layout. This keeps different stages conceptually separate, while always revealing the current composed result. In informal pilot studies, we found that using Creoâs set of unordered stages, as opposed to a timeline of sequential edits (Willis et al., 2021), helped users understand the impact of their edits without overloading the workspace. To ensure users can revisit stages without losing their existing progress, Creo implements re-propagation. When a user edits an earlier stage, those changes are propagated forward to downstream stages while preserving existing decisions where possible. For example, in Figure 3), when a user modifies a dogâs fur type and adds accessories after specifying color and lighting, they will see the updated composition rendered with the previously defined palette and illumination. By scoping changes to a specific aspect of the image and referencing previous intermediary representations as locked decisions, the system reduces unintended interactions between decisions and maintains consistency across stages. Re-propagation is implemented by ensuring that every generation step is conditioned on both the current stage as well as a the userâs decision state. The decision state includes the current composed image, decisions in earlier stages, and stage-specific instructions defining which attributes are editable (Figure 9). Rather than regenerating images from scratch, Creo incrementally updates the current image while locking the userâs previous decisions. This allows users to compare alternatives, revert edits, and branch from earlier decisions without losing progress. 3.4. Design Rationale and Tradeoffs Designing Creo required navigating several tradeoffs. Below we describe the key tensions that shaped our design. 3.4.1. Progressive Detail vs. Editable Representations. A central tension is how to progressively introduce visual detail (DG1) while ensuring users can meaningfully manipulate incomplete images (DG3). A natural approach is to vary the level of detail within a single image, for example generating a high-resolution output while simplifying underspecified regions. For example, given the prompt âa living room with an orange sofa near a window,â where the sofa is well specified but the surrounding layout remains ambiguous, a system might fully render the sofa while blurring the background. In practice, however, this does not support progressive ideation. Even when regions are blurred, users must still interpret and edit a fully rendered scene. For example, repositioning the sofa requires reasoning about perspective, shadows, and surrounding objects even in visually de-emphasized areas. As a result, users often avoid direct manipulation of fully rendered images and instead regenerate outputs or adjust prompts, reinforcing a reactive workflow. This contrast is clear when compared to Creoâs sketch-based interaction. Editing a sketch of the same scene decreases the cognitive and interaction cost. Repositioning the sofa only requires adjusting simple outlines without needing to account for lighting or texture. Moreover, while rendered images appear âfinishedâ and discourage editing, sketches appear provisional and invite modification. This reveals a limitation of applying DG1 in isolation. Progressively revealing detail is insufficient if interaction remains tied to pixel-level representations that are challenging to edit. As long as users operate on rendered images, they must reason about appearance rather than structure. 3.4.2. Speed vs. Structured Control. Another tension lies between rapid iteration and decomposed control. One-shot generation enables fast exploration by producing complete images from a single prompt, but entangles multiple decisions and makes targeted edits difficult. Decomposing generation into stages introduces structure and control, but increases interaction overhead. This raised the question of how many stages to include and what decisions each should expose. We initially explored designs that mapped sketches directly to fully rendered outputs, similarly to existing systems such as ControlNet (Zhang et al., 2023) and Block and Detail (Sarukkai et al., 2024). While this provided control over the composition of the image, it left many other aspects of the image such as color, lighting, and style up to the model, and we found it anchored details prematurely and limited control over the aesthetic aspects of the image. Instead, we grounded Creoâs stage decomposition in studio art practices (e.g., thumbnailing, refinement, color studies, shading) (Loomis, 1947; McCloud, 1994; Eisner, 2008; Dow, 1914) and extended this progression with a final rendering stage. Informal feedback indicated that users consistently valued control over composition, color, and lighting, even if their priorities differed We therefore adopted this set of stages as a balance between expressive control and interaction complexity. 3.4.3. Linear vs. Non-Linear Workflows. Creative workflows are non-linear, and users frequently revisit earlier decisions after observing downstream results. Thus, while staged decomposition introduces clarity, forcing a sequence of stages would be misaligned with practice. To support non-linear workflows, we explored a timeline-based approach (Willis et al., 2021), but found that organizing interaction temporally made it difficult to distinguish between types of visual decisions and how edits would affect the image. Instead, Creo adopts a semantic organization (Photoshop, 2026) where each stage corresponds to a visual decision (e.g., color) but users can move between stages in any order. However, stage-based design introduces a cognitive challenge: revisiting earlier stages can feel like âgoing back in time,â requiring users to recall downstream decisions. We explored showing multiple stages simultaneously, but found this visually cluttered. Instead, we provide a persistent preview of the final image, allowing users to understand the impact of edits without overloading the workspace. 3.4.4. User Control vs. Automation. A final tension concerned balancing how much manual control to give users via direct manipulation tools vs. how to leverage AI modelâs abilities to alleviate busy work. We address this by supporting multiple levels of control within each stage. Users can directly manipulate representations (e.g., drawing structure, applying color, specifying lighting direction) or provide partial input (e.g., masks, palettes, or high-level prompts) and rely on the model to complete the result. For example, users may sketch approximate color regions and use AI-assisted tools to refine them, or specify a lighting âvibeâ and allow the system to infer detailed illumination. This flexibility gives users agency to allocate their effort, focusing on decisions they care about while delegating othersâand accommodates different creative styles. 3.4.5. General Purpose vs. Sketch-Specific Models. Creo relies on T2I models to translate user edits into an updated image. Since many stages rely on sketch-based representations, we initially explored using specialized sketch-generation models (Vinker et al., 2025). However, we found that prompt-based adaptations of general-purpose models were sufficient to produce sketch-like outputs, allowing us to use a single model family across all Creo stages. 4. User Study: One-Shot vs Progressive Generation Evaluation To evaluate whether structuring image generation into progressive stages influences creative ideation, we conducted a controlled user study comparing Creoâs staged workflow with a conventional one-shot text-to-image interface. Our goal is to understand how these workflows shape (1) breadth of idea exploration, (2) usersâ ability to make targeted changes without disrupting prior decisions, (3) when users commit to a design direction, and (4) their sense of authorship and whether they feel they are actively shaping the image or reacting to model-generated outputs. 4.1. Study Overview We employed a within-subjects, counterbalanced design with two conditions: Progressive Staging (Creo). Participants used Creoâs staged workflow, which decomposes image generation into sequential decision stages including viewpoint, composition, color, lighting, and rendering style. Participants could move freely between stages and revisit earlier decisions. One-Shot Baseline. Participants used a multi-turn prompt-centric text-to-image interface (ChatGPT) that supported iterative prompt refinement, regeneration, and masked edits. To reduce concept carryover between conditions and ensure participants were starting each task fresh conceptually, participants completed two related design prompts: design your ideal living room and design your ideal kitchen. Each participant completed both prompts, using a different tool for each prompt. Tool order and prompt order were counterbalanced across participants. 4.1.1. Participants: We recruited 15 participants through posts in online (Reddit) and university communities focused on illustration, art, and design. Participants had prior experience in creative workflows across domains including graphic design and marketing (n=6), interior design and architecture (n=5), illustration and concept art (n=2), animation (n=1), and fine arts (n=1). Participants represented a mix of early-career practitioners (1â3 years, n = 4), to mid-level experience (4â7 years, n = 6), to highly experienced practitioners (8+ years, n = 5). Participants had a range of prior exposure to generative AI image tools: 4 reported using them regularly, 6 occasionally, and 5 had limited or no active use. Sessions lasted approximately 60 minutes and were conducted remotely via video conferencing with screen sharing enabled. Participants received a gift card as compensation for their time. 4.1.2. Study Protocol Sessions began with a brief discussion of participantsâ creative workflows, use of T2I generative AI tools in their work. Participants then completed both conditions in counterbalanced order. Before each condition, the moderator introduced the interface. Participants were asked to think aloud while working for approximately 14 minutes per condition. After each condition, participants reflected on their experience, followed by a final composition discussion between the two workflows. Additional details on materials, task instructions, counterbalancing, and moderation are provided in Appendix D.6. 4.2. Analysis Overview We conducted a mixed-methods analysis combining structured interaction logs with qualitative analysis of participantsâ verbal reflections. Our analysis is organized around three research questions corresponding to exploration, control, and workflow structure. From screen recordings, we reconstructed interaction sequences and segmented them into discrete actions (e.g., construct, evaluate, generate, repair). Each action was annotated with its intent (on-intent, pivot, drift) and whether it was user-driven or model-led, and whether it constituted a direction change (a major shift in the overall concept or design trajectory). Actions were grouped into iterations to capture branching exploration. These annotation categories were derived iteratively from our research questions, aiming to operationalize constructs such as exploration, control, and workflow structure. We refined the schema through multiple passes over a subset of sessions, adjusting definitions to ensure they were both interpretable from observable behavior and consistently applicable across participants. To scale annotation, we implemented an LLM-based coding pipeline that labeled each action a user took according to this schema. We validated this process by manually cross-checking a subset of annotated sessions against the original recordings, ensuring consistency between automated labels and observed user behavior. Each annotation is used to derive session-level metrics (e.g., direction changes, drift, repair actions), which we define briefly below and describe in full in Appendix D.6. RQ1: Exploration and anchoring. To characterize ideation breadth and commitment, we measure the number of direction changes per session, individually coded from the study recordings, and the number of distinct iterations (i.e., separate exploration branches). We additionally compute the proportion of actions labeled as drift (unintended deviations introduced by the model) and pivot (intentional redirection by the user). As a proxy for anchoring, we compute the L2 distance between unit-norm CLIP embeddings of the first and final images in each session. This summarizes how much endpoint appearance (in CLIP space) drifts from the initial frame: larger distance means the last image lies farther from the first in the modelâs representation; smaller distance means the session remains closer to its starting point, which we treat as stronger anchoring in embedding space. RQ2: Control and predictability. To examine interaction control, we measure the proportion of construct actions (direct content specification or modification) and evaluate actions (inspection of outputs). We approximate perceived agency as the proportion of user-driven actions (actions where users explicitly specify or modify content). We quantify revision burden as the ratio of repair actions (corrections of unintended changes), and measure unintended side effects via the rate of invariant violations, defined as edits that unintentionally alter aspects outside their intended scope (e.g., changing layout during a color edit). RQ3: Workflow structure. To understand how participants appropriate staged interaction, we analyze iteration patterns and stage transitions in Creo, including revisiting earlier stages and skipping intermediate stages. We also summarize stage usage as the proportion of actions occurring within each stage and the proportion of participants who engaged with each stage. Qualitative analysis. We analyzed think-aloud protocols and post-task interviews using thematic coding to capture participantsâ experiences of authorship, control, predictability, and commitment. These qualitative themes are used to interpret the behavioral patterns observed in each condition, linking quantitative measures (e.g., constructive vs. evaluation-heavy behavior, direction changes, repair actions) to participantsâ reported experiences of constructing, steering, or reacting to model outputs. All metrics are computed at the session level and aggregated across participants, and are based on action counts and proportions rather than time-based metrics. Full annotation and coding details are provided in Appendix D.6. 4.3. User Study Results Our mixed-methods analysis showed that Creo and the one-shot baseline (GPT) supported different modes of ideation. GPT positions users as reacting to finished outputs, while Creo enabled users to progressively construct images through editable intermediate states, increasing perceived agency and authorship, but also introducing additional revision work due to its persistent structure. One-shot baseline outputs were regarded as highly-polished, and unintended changes were perceived as aesthetic recommendations, while Creo outputs were perceived as provisional and editable. Below, we present four findings showing how staged interaction changes both what users can control and how they relate to the artifact they are creating. Table 1. Creoâs staged generation outperforms GPTâs one-shot prompting across exploration and control metrics (Ο¹ĎΟ¹Ď). Metric Creo GPT Exploration & Anchoring Direction changes (# changes per session) 1.6Âą1.01.6Âą 1.0 0.4Âą0.50.4Âą 0.5 Exploration breadth (# iterations per branch) 3.3Âą1.83.3Âą 1.8 2.5Âą1.72.5Âą 1.7 Concept drift (% unintended drift actions) 14.4Âą10.214.4Âą 10.2 14.2Âą8.914.2Âą 8.9 Intentional pivots (% pivot actions) 3.9Âą3.43.9Âą 3.4 2.9Âą4.32.9Âą 4.3 Control & Predictability Constructive engagement (% construct actions) 17.7Âą8.017.7Âą 8.0 4.3Âą7.64.3Âą 7.6 Evaluation-heavy behavior (% evaluate actions) 35.7Âą5.835.7Âą 5.8 39.4Âą10.639.4Âą 10.6 Behavioral agency (% user-driven actions) 55.0Âą11.855.0Âą 11.8 25.8Âą12.525.8Âą 12.5 Revision burden (ratio of repairs to total actions) 0.1Âą0.10.1Âą 0.1 0.0Âą0.10.0Âą 0.1 Unintended changes (% invariant violations) 10.9Âą10.310.9Âą 10.3 0.0Âą0.00.0Âą 0.0 4.3.1. Provisional representations reduce anchoring and support redirection Participants were less anchored to early outputs in Creo than in GPT, both behaviorally and in how they described the workflow. Creo sessions included more direction changes (major shifts in concept or design trajectory) than GPT sessions (1.6 vs. 0.4 per session), a 4Ă increase, and more iterations (distinct exploration branches) (3.3 vs. 2.5; Table 1). Anchoring was also weaker in Creo: the L2 distance between first and last generated images in CLIP space was higher (M = 0.75, SD = 0.13) than under GPT (M = 0.46, SD = 0.31) as shown in Figure 4, indicating endpoints less aligned with the starting image. Between-session spread in this distance was larger for GPT, consistent with more varied GPT use versus a more standardized Creo pipeline. Participantsâ accounts help explain why. For example, several participants, including P9, P3, and P11, were satisfied with their initial image. In the GPT condition, early outputs were often treated as already complete. For example, P9 prompted a photorealistic kitchen and accepted the first result as âpresentableâ and âpretty good,â noting that the lack of alternatives limited further exploration. P12 similarly constructed a detailed prompt and described the returned image as âsomething very finished right away,â making only minor adjustments and treating it as âclose enough.â In both cases, polished rendering appeared to signal completion. Rather than redirecting the image, participants tended to adjust prompts around the modelâs first proposal. In Creo, participants described the opposite effect. Intermediate outputs were treated as provisional and open to change. After viewing several sketch-like outputs in the viewpoint stage, P1 described the images as âa question mark⌠something I can still go in and shape.â The rough, incomplete quality signaled that the image was still open to change, making it feel more accessible than a fully rendered result. P9 actively sketched on the canvas while developing a living room scene, adding a window and a corner detail that the system then interpreted and cleaned into a more precise drawing. He described these âsketchy imagesâ as the âright level of detail to start⌠not too finalâ, using sketching was not just a way of editing the image but a way of thinking through the design. These representations lowered the perceived cost of change, despite the cost of a model change being effectively the same, and made it easier to revisit earlier decisions without feeling that a finished solution had already been proposed. Figure 4. Creo reduces design anchoring and homogenization. Visual embeddings of the images that users produced using Creo are less tightly clustered than those from GPT. 4.3.2. Staged interaction shifts users from evaluating outputs to constructing them Staged interaction shifted participantsâ behavior from evaluating model outputs toward directly constructing and modifying them. Creo sessions showed substantially more construct actions (direct content specification or modification) than GPT (17.7% vs. 4.3%) indicating that participants more often directly modified the image rather than evaluating generated outputs, and a higher proportion of user-driven actions (55.0% vs. 25.8%), indicating greater direct control over the artifact (Table 1). In contrast, GPT interaction remained more evaluation-heavy (39.4% vs. 35.7%). Participants described this difference in terms of how they interacted with the image. P10, working on a small studio layout, alternated between erasing and redrawing parts of a counter and using masking to introduce a central island. Reflecting on this process, she noted that âdrawing felt more direct than trying to phrase it.â Similarly, P13 used direct manipulation together with masking and prompting to rearrange the sofa in her living room, and said she preferred being able to move the couch âwhere I wanted it instead of describing it again.â This shift in interaction also changed how participants understood their role in the process. Rather than requesting outputs and reacting to them, participants described working on the image over time. P4, who in the GPT condition front-loaded detailed prompts to avoid iteration, contrasted this with Creo, where he could modify individual elements without affecting the whole and being abele to âfinish what Iâve started.â P6 described working with Creo as âa conversation with a partner,â where each stage built on prior decisions. Across multiple stages, he developed a medieval-style kitchen from a black-and-white sketch, adjusted color and lighting, and then returned upstream to add a cat that propagated through the downstream layers. He expressed pride not just in the final image but in being able to recount how it came together: âI could say, first I did this, then I did that.â In contrast, although he described GPT outputs as âgorgeous,â he felt less ownership and said he âwould not be proud to showâ them because they felt too easy to produce. Others expressed similar distinctions: P9 described the process as âI worked with it, not just got something,â while P2 noted that in GPT, âIâm refining the prompt more than the image.â Participants frequently framed their experience in terms of continuity and accumulation rather than repeated generation. As a result, authorship emerged from being able to trace and recount decisions rather than from the final image quality alone. This difference also shaped how participants valued the results. Although GPT outputs were often more polished or photorealistic, they were more likely to be treated as disposable. In contrast, participants expressed stronger attachment to Creo outputs and greater interest in keeping or sharing them, suggesting that perceived ownership was tied to process contribution rather than output quality. 4.3.3. Preserving decisions makes unintended changes visible and repairable While both systems exhibited similar levels of drift (unintended deviations introduced by the model; 14.4% in Creo, 14.2% in GPT), they differed in how these deviations were experienced and handled. While unintended drift occurred at similar rates in both conditions, Creo showed slightly more intentional pivots (3.9% vs. 2.9%), suggesting that users more often redirected the process deliberately rather than accommodating model outputs. Creo showed a higher revision burden (ratio of repair actions; 0.1 vs. âź 0) and non-zero invariant violations, edits that unintentionally altered aspects outside their intended scope, (10.9% vs. 0.0%). . We note that these violations are participant-reported in the transcripts, rather than codings of observed behavior. At first glance, this might suggest that Creo introduces more errors. However, participant accounts indicate a different interpretation. For example, P15 prompted a country-style kitchen but received a more modern result and adopted it, explaining that the model had âactually improved my taste⌠I shifted direction because of it.â Similar reactions were reported by other participants, including P3, P6, and P13, who described baseline model deviations from their prompt as useful suggestions rather than errors. This suggests that the increased revision burden in Creo reflects not worse model performance, but a different interaction structure. By preserving prior decisions, the system makes inconsistencies explicit and localizable, allowing users to correct them directly. In contrast, one-shot generation obscures these inconsistencies by replacing the entire image at each step, and the high-resolution rendering style makes them feel like aesthetic improvements. 4.3.4. Participants used staging as a flexible workspace rather than a fixed pipeline Participants did not follow the staged workflow linearly. A majority (66.7%) revisited earlier stages and skipped others, entering through different starting points such as composition, lighting, or object-level refinement. These differences are reflected in stage usage patterns (Figure 5). Interaction was most concentrated in Composition (39.9% of actions), followed by Color (21.8%), Lighting (19.5%), and Style (18.8%). Participants also entered the workflow through different stages. Some began by focusing on composition or layout, while others prioritized lighting, color, or object-level refinement. For example, participants with backgrounds in architecture or 3D design often adjusted lighting early in the process, while others focused first on spatial arrangement or color relationships. These differences are reflected in stage usage patterns (Figure 5), where interaction concentrated differently across stages depending on the participant. Figure 5. Creo supports peopleâs non-linear creative workflows. Participants P4-P9 frequently explored stages out of order and revisited stages to generate their images. Importantly, this structure appeared to align with participantsâ existing mental models of image creation rather than imposing a new workflow. In the GPT condition, several participants independently decomposed their prompts into components such as viewpoint, composition, color, and style, effectively reconstructing a similar breakdown in natural language. Others externalized this process by using reference images or sketches before prompting, suggesting that they already conceptualized image creation as a sequence of separable decisions. Together, these observations indicate that staging did not constrain users to a fixed pipeline, but instead provided a structured space that supported multiple valid entry points and trajectories. By making different dimensions of the image explicit and independently editable, Creo enabled users to work in ways that reflected their prior experience, goals, and domain knowledge. 5. Discussion and Future Work Our findings suggest that the main contribution of staged image generation is not simply increased controllability. It is a different interaction model. In the one-shot baseline, participants were more often positioned as evaluating and accommodating model proposals. In Creo, they were more often able to build images through a sequence of revisable decisions whose effects remained visible over time. This suggests that an underexplored design dimension in generative systems is the temporal structure of interaction: not only what users can specify, but how a system supports intent formation, revision, and commitment over the course of creation. 5.1. Progressive Commitment as a Temporal Interaction Strategy We interpret these results through the lens of progressive commitment, in which artifacts evolve from coarse, under-specified representations to more detailed ones over time. Prior work in design and sketching shows that early, low-fidelity representations support exploration by keeping alternatives open and reducing the cost of change (Tversky, 2002; Suwa et al., 2000; Buxton, 2010). Our findings suggest that this principle carries over to generative systems. When users interacted with sketch-like intermediate outputs, they were more willing to redirect, revise, and elaborate their ideas. When they interacted with polished one-shot outputs, they were more likely to treat those outputs as proposals to assess or lightly tweak rather than as materials to reshape. In this view, the contribution of Creo is not simply that it decomposes image generation into several steps. It is that each step is paired with a representation suited to the decision at hand. Early stages support structural reasoning through provisional sketches, while later stages introduce appearance-level decisions such as color, lighting and style. This pairing helps users revise earlier choices without restarting the entire process and suggests that effective generative interfaces may depend as much on representational timing as on the number of controls they expose. At the same time, staging introduces new expectations. When visual dimensions are exposed as separable, users expect them to behave independently. Some of the tensions observed in the user study occurred when this expectation was only partially met, such as when a change intended for one stage affected another aspect of the image. This highlights an important design challenge for future generative systems. Progressive interfaces do not only require more structure at the interaction level. They also require behavior that is sufficiently stable and localized to respect that structure. 5.2. Authorship and the Visibility of Creative Contribution Our findings also suggest that perceived authorship in generative systems depends on whether users can recognize their contributions, and chains of decisions in the resulting artifact. In one-shot workflows, participants often described themselves as selecting or evaluating outputs. In contrast, staged interaction made the construction process visible: users could trace how the image evolved through their actions, and describe a sequence of decisions leading to the final result. This aligns with prior work showing that involving users in the creative loop increases perceived ownership (Lubart, 2005; Guzdial and Riedl, 2019; Shneiderman, 2007; Deterding et al., 2017; Oh et al., 2018), but suggests that visibility of contribution may be as important as participation itself. Systems that preserve intermediate decisions and make them inspectable may better support a sense of authorship, even when model assistance remains substantial. An interesting implication of our findings is that these effects persist even when the cost of regeneration is the same for low-fidelity and high-fidelity outputs. Prior work shows that people are less likely to critique high-fidelity representations and more willing to revise low-fidelity ones due to their perceived provisionality (Buxton, 2010; Goel, 1995; Tversky, 2002; Lim et al., 2008). Our results suggest that this dynamic carries over to generative systems: even though generating a new image requires minimal effort, participants were less likely to redirect high-resolution outputs and more likely to explore when working with sketch-like representations. This extends prior findings from traditional design settings to generative AI, where the cost of producing alternatives is negligible but perceptions of finality continue to shape behavior. 5.3. Limitations and Future Work Our prototype reflects one possible decomposition of visual decisions (e.g., viewpoint, composition, color, lighting, style), rather than a universal structure. Different domains may require alternative intermediate representations, and future work should explore how staging can be adapted to domain-specific workflows. In addition, current generative models do not fully support clean separation between visual dimensions. Participants observed unintended interactions across stages, for example changing the color of the sofa could affect the shape, reflecting underlying entanglement in model representations. As staged interfaces make these dimensions explicit, improving controllability and disentanglement becomes increasingly important. More broadly, progressive interfaces may extend beyond image generation to domains such as writing, programming, or decision-making, where users benefit from structuring complex tasks into revisable stages rather than producing final artifacts in one step. Additionally, our results suggest that interaction designâhow outputs are staged, and how users can act on themâmay be as important as model capability in shaping creative outcomes. Finally, our study focuses on short ideation sessions. Many creative workflows unfold over longer timeframes or involve collaboration, where staged generation may support iterative refinement across sessions or coordination between contributors. Exploring progressive interaction in these settings remains an important direction for future work. Overall, progressive commitment provides a promising framework for aligning generative systems with how ideas develop over time, supporting more exploratory, controllable, and personally meaningful forms of humanâAI collaboration. 6. Conclusion Current text-to-image systems generate fully detailed images in a single step, where many visual decisions are made together. As a result, editing one aspect often unintentionally changes others, and early outputs can anchor users before their ideas are fully formed. We introduced Creo, a multi-stage image ideation system that, inspired by how creative workflows isolate visual aspects, structures creation as a sequence of stages while preserving prior decisions. Instead of producing a final image upfront, Creo begins with rough sketches and progressively adds detail. At each stage, users refine specific aspects of the image, while a locking mechanism ensures that earlier decisions remain stable during later edits. In a comparative study, participants were better able to make targeted changes, explore alternatives, and trace how their decisions shaped the final image. This led to a stronger sense of ownership and produced less homogeneous outputs than one-shot generation. These findings suggest that improving generative systems is not only about increasing control, but about structuring how and when decisions are made. Staged generation with mechanisms for preserving decisions can better support exploration, user agency, and diverse creative outcomes. More broadly, our findings suggest that the design of generative systems should focus not only on model capability, but on how interaction is structured over time. By organizing generation around intermediate, revisable representations, progressive interfaces can better support how users form, refine, and take ownership of ideas. We see progressive commitment as a promising direction for building more exploratory, controllable, and human-centered generative tools. References M. Agrawala (2023) Unpredictable black boxes are terrible interfaces. ACM TechTalks. Cited by: §2.2. H. Alzayer, Z. Xia, X. Zhang, E. Shechtman, J. Huang, and M. Gharbi (2025) Magic fixup: streamlining photo editing by watching dynamic videos. ACM Transactions on Graphics 44 (5), p. 1â25. Cited by: §2.2. Artists & Illustrators (2021) How to illustrate a childrenâs book. Artists & Illustrators. Note: Accessed January 2026 External Links: Link Cited by: §3.1. S. Bae, R. Balakrishnan, and K. Singh (2008) ILoveSketch: as-natural-as-possible sketching system for creating 3d curve models. In Proceedings of the 21st annual ACM symposium on User interface software and technology, p. 151â160. Cited by: §2.1. E. J. Bigelow, J. P. McCoy, and T. D. Ullman (2023) Non-commitment in mental imagery. Cognition 238, p. 105498. Cited by: §1, §2.1. S. Brade, B. Wang, M. Sousa, S. Oore, and T. Grossman (2023) Promptify: text-to-image generation through interactive prompt exploration with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, p. 1â14. Cited by: §2.2. J. Brun, P. Le Masson, and B. Weil (2016) Designing with sketches: the generative effects of knowledge preordering. Design Science 2, p. e13. Cited by: §2.1. B. Buxton (2010) Sketching user experiences: getting the design right and the right design. Morgan kaufmann. Cited by: §5.1, §5.2. F. Chevalier, P. Dragicevic, and C. Hurter (2012) Histomages: fully synchronized views for image editing. In Proceedings of the 25th annual ACM symposium on User interface software and technology, p. 281â286. Cited by: §2.1. F. D. K. Ching (2014) Architectural graphics. 6 edition, Wiley, Hoboken, NJ, USA. Cited by: §3.1. N. Cohn (2013) The visual language of comics: introduction to the structure and cognition of sequential images. Bloomsbury Academic, London. Cited by: §1, §2.1. S. Deterding, J. Hook, R. Fiebrink, M. Gillies, J. Gow, M. Akten, G. Smith, A. Liapis, and K. Compton (2017) Mixed-initiative creative interfaces. In Proceedings of the 2017 CHI conference extended abstracts on human factors in computing systems, p. 628â635. Cited by: §5.2. A. R. Doshi and O. P. Hauser (2024) Generative ai enhances individual creativity but reduces the collective diversity of novel content. Science advances 10 (28), p. eadn5290. Cited by: §1, §2.2. A. W. Dow (1914) Composition. Doubleday, Doran, Incorporated. Cited by: §2.1, §3.4.2. W. Eisner (2008) Comics and sequential art. Revised edition edition, W. W. Norton & Company, New York. Cited by: §1, §2.1, §3.4.2. J. E. Fan, W. A. Bainbridge, R. Chamberlain, and J. D. Wammes (2023) Drawing as a versatile cognitive tool. Nature Reviews Psychology 2 (9), p. 556â568. Cited by: §2.1. Y. Feng, X. Wang, K. K. Wong, S. Wang, Y. Lu, M. Zhu, B. Wang, and W. Chen (2023) Promptmagician: interactive prompt engineering for text-to-image creation. IEEE Transactions on Visualization and Computer Graphics 30 (1), p. 295â305. Cited by: §2.2. J. Fernquist, T. Grossman, and G. Fitzmaurice (2011) Sketch-sketch revolution: an engaging tutorial system for guided sketching and application learning. In Proceedings of the 24th annual ACM symposium on User interface software and technology, p. 373â382. Cited by: §2.1. V. Goel (1995) Sketches of thought. MIT press. Cited by: §5.2. G. Goldschmidt (2014) Modeling the role of sketching in design idea generation. In An anthology of theories and models of design: philosophy, approaches and empirical explorations, p. 433â450. Cited by: §2.1. [21] C. Goodwin 1994!. professional vision. American Anthropologist 96 (3), p. 606â633. Cited by: §2.1. T. R. Green (1989) Cognitive dimensions of notations. People and computers V, p. 443â460. Cited by: §1. M. Guzdial and M. Riedl (2019) An interaction framework for studying co-creative ai. arXiv preprint arXiv:1903.09709. Cited by: §5.2. A. Hertzmann (2020) Why do line drawings work? a realism hypothesis. Perception 49 (4), p. 439â451. Cited by: §2.1. A. Hopkins, A. Boggust, and H. Suresh (2025) Chatbot evaluation is (sometimes) ill-posed: contextualization errors in the human-interface-model pipeline. In Proceedings of the Human-Centered Evaluation and Auditing Workshop (HEAL@CHI), Cited by: §2.2. E. Iarussi, A. Bousseau, and T. Tsandilas (2013) The drawing assistant: automated drawing guidance and feedback from photographs. In ACM Symposium on User Interface Software and Technology (UIST), Cited by: §2.1. T. Igarashi, S. Matsuoka, and H. Tanaka (1999) Teddy: a sketching interface for 3d freeform design. In Proceedings of ACM SIGGRAPH, p. 409â416. Cited by: §2.1. D. G. Jansson and S. M. Smith (1991) Design fixation. Design studies 12 (1), p. 3â11. Cited by: §1, §2.2. S. Lee, J. Yoon, S. Lee, J. H. Lee, and S. Bae (2025) 3D sketching + 2d generative ai for car exterior design. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST 2025), Note: Best Demo Honorable Mention External Links: Document Cited by: §2.2. Y. Lim, E. Stolterman, and J. Tenenberg (2008) The anatomy of prototypes: prototypes as filters, prototypes as manifestations of design ideas. ACM Transactions on Computer-Human Interaction (TOCHI) 15 (2), p. 1â27. Cited by: §5.2. M. Livingstone and D. H. Hubel (2002) Vision and art: the biology of seeing. (No Title). Cited by: §2.1. A. Loomis (1947) Creative illustration. Viking Press New York, NY. Cited by: §2.1, §3.4.2. T. Lubart (2005) How can computers be partners in the creative process: classification and commentary on the special issue. International journal of human-computer studies 63 (4-5), p. 365â369. Cited by: §5.2. J. Ma, C. Vu, A. Lyubavina, C. Liu, and J. Li (2025) Computational scaffolding of composition, value, and color for disciplined drawing. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST â25, New York, NY, USA. External Links: ISBN 9798400720376, Link, Document Cited by: §2.1. S. McCloud (1994) Understanding comics: the invisible art. HarperCollins, New York. Cited by: §1, §2.1, §3.4.2. K. Nicolaides (1941) The natural way to draw. Houghton Mifflin, Boston, MA, USA. Cited by: §3.1. P. OâDonovan, A. Agarwala, and A. Hertzmann (2015) DesignScape: design with interactive layout suggestions. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, CHI â15, New York, NY, USA, p. 1221â1224. External Links: ISBN 9781450331456, Link, Document Cited by: §2.1. C. Oh, J. Song, J. Choi, S. Kim, S. Lee, and B. Suh (2018) I lead, you help but only with enough details: understanding user experience of co-creation with artificial intelligence. In Proceedings of the 2018 CHI conference on human factors in computing systems, p. 1â13. Cited by: §5.2. J. A. Okun and S. Zwerman (2010) The ves handbook of visual effects: industry standard vfx practices and procedures. Focal Press, Burlington, MA. Cited by: §1, §2.1. R. Parent (2012) Computer animation: algorithms and techniques. 3rd edition, Morgan Kaufmann, Burlington, MA. Cited by: §1, §2.1. A. Photoshop (2026) Photoshop. Retrieved January. Cited by: §3.4.3. A. T. Purcell and J. S. Gero (1998) Drawings and the design process: a review of protocol studies in design and other disciplines and related research in cognitive psychology. Design studies 19 (4), p. 389â430. Cited by: §2.1, §2.1. A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125. Cited by: §1. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: §1. C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. Ghasemipour, and et al. (2022) Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487. Cited by: §1. V. Sarukkai, L. Yuan, M. Tang, M. Agrawala, and K. Fatahalian (2024) Block and detail: scaffolding sketch-to-image generation. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, p. 1â13. Cited by: §2.2, §2.2, §3.4.2. D. A. SchĂśn (1983) The reflective practitioner: how professionals think in action. Basic Books, New York, NY. Cited by: §2.1, §2.2. X. Shi, L. Wei, N. Zhao, J. Zhao, and R. H. Kazi (2026) Notational animating: an interactive approach to creating and editing animation keyframes. In Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI 2026), Cited by: §2.2. B. Shneiderman (2007) Creativity support tools: accelerating discovery and innovation. Communications of the ACM 50 (12), p. 20â32. Cited by: §5.2. M. Shugrina, J. Lu, and S. Diverdi (2017) Playful palette: an interactive parametric color mixer for artists. ACM Transactions on Graphics (TOG) 36 (4), p. 1â10. Cited by: §2.1. M. Shugrina, W. Zhang, F. Chevalier, S. Fidler, and K. Singh (2019) Color builder: a direct manipulation interface for versatile color theme authoring. In Proceedings of the 2019 CHI conference on human factors in computing systems, p. 1â12. Cited by: §2.1. H. Subramonyam, R. Pea, C. Pondoc, M. Agrawala, and C. Seifert (2024) Bridging the gulf of envisioning: cognitive challenges in prompt based interactions with llms. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â19. Cited by: §2.2. M. Suwa, J. Gero, and T. Purcell (2000) Unexpected discoveries and s-invention of design requirements: important vehicles for a design process. Design studies 21 (6), p. 539â567. Cited by: §5.1. A. Swearngin, A. J. Ko, and J. Fogarty (2018) Scout: mixed-initiative exploration of design variations through high-level design constraints. In Adjunct Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology, p. 134â136. Cited by: §2.1. J. Szuc (2020) Behind the scenes: illustration process tutorial. Jeff Szuc. Note: Accessed January 2026 External Links: Link Cited by: §3.1. M. Tang, Y. Vinker, C. Yan, L. Zhang, and M. Agrawala (2025) Instance segmentation of scene sketches using natural image priors. In ACM SIGGRAPH Conference Proceedings, p. 96:1â96:10. Cited by: §2.2. F. Thomas and O. Johnston (1995) The illusion of life: disney animation. Disney Editions, New York. Cited by: §1, §2.1. B. Tversky (2002) What do sketches say about thinking. In 2002 AAAI Spring Symposium, Sketch Understanding Workshop, Stanford University, AAAI Technical Report S-02-08, Vol. 148, p. 151. Cited by: §2.1, §2.2, §5.1, §5.2. I. M. Verstijnen, C. van Leeuwen, G. Goldschmidt, R. Hamel, and J. Hennessey (1998) Sketching and creative discovery. Design studies 19 (4), p. 519â546. Cited by: §2.1. Y. Vinker, T. R. Shaham, K. Zheng, A. Zhao, J. E Fan, and A. Torralba (2025) SketchAgent: language-driven sequential sketch generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 23355â23368. Cited by: §3.4.5. A. Voynov, K. Aberman, and D. Cohen-Or (2023) Sketch-guided text-to-image diffusion models. In ACM SIGGRAPH 2023 conference proceedings, p. 1â11. Cited by: §2.2. S. Wadinambiarachchi, R. M. Kelly, S. Pareek, Q. Zhou, and E. Velloso (2024) The effects of generative ai on design fixation and divergent thinking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1â18. Cited by: §1, §2.2. B. Williford, A. Doke, M. Pahud, K. Hinckley, and T. Hammond (2019) DrawMyPhoto: assisting novices in drawing from photographs. In Proceedings of the 2019 Conference on Creativity and Cognition, p. 198â209. Cited by: §2.1. K. D. Willis, Y. Pu, J. Luo, H. Chu, T. Du, J. G. Lambourne, A. Solar-Lezama, and W. Matusik (2021) Fusion 360 gallery: a dataset and environment for programmatic cad construction from human design sequences. ACM Transactions on Graphics (TOG) 40 (4), p. 1â24. Cited by: §3.3, §3.4.3. L. Zhang, A. Rao, and M. Agrawala (2023) Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, p. 3836â3847. Cited by: §1, §2.2, §2.2, §3.4.2. L. Zhang, C. Yan, Y. Guo, J. Xing, and M. Agrawala (2025) Generating past and future in digital painting processes. ACM Transactions on Graphics, p. 127:1â127:13. Cited by: §2.2. 7. Appendix Appendix A Creo Workflows: Two entry points to the same abstractions We support two entry points into the Creo workflow: starting from a text prompt or from an existing image (Figure 6). In both cases, Creo uses the same multi-stage generation process, allowing users to progressively manipulate independent representations. However, in the prompt-first workflow, users begin with a natural-language description and start with a sketch. Whereas, in the reference-first workflow, users upload an existing image and Creo reverse-engineers it into the five stages. This maps the finished image to the same intermediate representations used in the prompt-first workflow, allowing users to edit individual stages without regenerating the entire image. Figure 6. Two entry points into Creoâs progressive ideation workflow. Users can begin with an existing image. This image is reverse engineered into the lighting, color and composition stages. Users can edit each of the representations, as shown in the bottom half of the image, moving across stages and propagating decisions across stages. This allows for iterative refinement of visual decisions on existing images. Appendix B Example Workflow: Comic Illustration We walk through a case study to contextualize the design decisions we made in Creo within real creative workflows. In particular, we work with a comic illustrator, exploring how Creo supports their process of designing a webcomic (Figure 2). The illustratorâs goal is to generate a cover image of their comic character, so they instantiate Creo using a prompt-first workflow asking it to generate âa close up of my main characterâ. While this is similar to traditional text-to-image systems, instead of generating a single high-fidelity image, Creo starts by producing a set of six viewpoints displaing the main character from above, below, and profile (Figure 21). These multiple versions scaffold the illustrators divergent thinking, allowing them to exploring different ways of viewing the scene without being influenced by stylistic details. After selecting a viewpoint the matches their vision, the illustrator edits the sketchâs composition. Since Creo intentionally keeps the image as a simple sketch, the illustrator is empowered to draw, erase, move, and transform elements directly or apply masked AI-assisted edits to modify specific regions. The starting viewpoint of the sketch not fully match their vision of the character, so the illustrator adjusts the proportions of the facial features, âmodifying the size of the eyes, reshaping the jawline, and refining the mouthâas well as parts of the back to better align with the intended viewpoint (Figure 2-Stage 2). Since the image is represented as a sketch, they are able to quickly iterate on the characterâs facial features without the overhead of managing shadows or skin texture. With the characterâs features now matching their expectations, the illustrator begins to increase image fidelity by adding color. In the color stage, users can define color palettes manually or generate them from prompts. Since the illustrator draws the same character frequently, they directly select the colors they use for this characters, creating a palette of muted greens and blues (Figure 2-Stage 3). They use direct manipulation tools to manually color in the characterâs hair with the highlights and lowlights they expect. They use AI-assisted tools to fill in colors they are less particular about, converting a rough green stroke over the phone into a neatly colored phone. Next, the illustrator decides to add shading and illumination in the lighting stage. This stageâs representation now reflects depth, making it possible to see how lighting affects the scene. The illustrator explores different light source placements, directions, and intensity, and explores different lighting conditions such as time of day or mood, finally settling on backlighting the character with neon lights (Figure 2-Stage 4). Finally, the illustrator chooses to style the image (Figure 2-Stage 5). While Creo supports various texture and stylistic effects (e.g., photorealistic, watercolor), the illustrator choosing digital painting to match their typical comic illustration style. The composed image is now high-fidelity, but thanks to Creoâs staged workflow, all of the illustratorâs previous choices are maintained. The character still reflects the facial structure they defined in composition and the hair details the illustrator sketched in color. The final image only reflects the illustratorâs design, not an AI modelâs assumptions. Even though their original prompt did not specify decisions like the color of the characterâs hair or the neon background, Creo did not fill that in for them. Instead, the illustrator was able to make each of those decisions as they walked through each of Creoâs stages, allowing them to leverage the power of text-to-image models while maintaining agency over their creative process. Appendix C Example Workflow from the User Study We walk through a case study to contextualize the design decisions we made in Creo within real creative workflows. In particular, we work with an architecture student, exploring how Creo supports their process of designing an interior living space (Figure 7). The participantâs goal is to generate a rendering of their ideal living room, so they instantiate Creo using a structured prompt-first workflow, describing the scene in layers: the objects they want to see, the desired viewpoint, and the overall style. While this is similar to traditional text-to-image systems, instead of generating a single high-fidelity image, Creo starts by producing a set of viewpoint sketches displaying the room from different angles and compositions (Figure 7-Stage 1). These multiple versions scaffold the participantâs divergent thinking, allowing them to explore different spatial arrangements without being influenced by stylistic or material details. The participant gravitates toward a composition that preserves the Charles River view while keeping the interior visible, selecting it as the basis for further refinement. With a viewpoint established, the participant begins editing the composition of the scene. Since Creo intentionally keeps the image as a simple sketch, the participant is empowered to draw, erase, and apply masked AI-assisted edits to modify specific regions. Using the masking tool, they identify an area of the room they want to enrich, and Creo surfaces contextually relevant suggestions based on the region selectedârecommending additions the participant had not explicitly considered, such as a carpet, but found immediately useful (Figure 7-Stage 2). Since the image is represented as a sketch, they are able to quickly iterate on the spatial arrangement without the overhead of managing materials or lighting. With the composition now matching their expectations, the participant begins to increase image fidelity by adding color. In the color stage, users can define color palettes manually or generate them from prompts. The participant selects colors that reflect their personal aestheticâapplying a blue to the couch and introducing orange accentsâusing direct manipulation tools to paint specific regions of the scene (Figure 7-Stage 3). They use AI-assisted tools to reconcile the colors across the scene, noting that the model balances interior and exterior tones in a way that feels spatially coherent. Next, the participant decides to add shading and illumination in the lighting stage. This stageâs representation now reflects depth, making it possible to see how lighting affects the interior space. The participant explores different light source placements and directions, experimenting with multiple simultaneous sources and different times of day, and notes that this level of lighting control is something they had not previously encountered in generative tools (Figure 7-Stage 4). They observe that Creo accurately propagates shadows consistent with the light directions they specify, reflecting an understanding of 3D spatial logic despite operating on a 2D image. Finally, the participant chooses to style the image (Figure 7-Stage 5). While Creo supports various texture and stylistic effects, the participant selects a photorealistic finish to match the kind of architectural rendering they produce in their professional workflow. The composed image is now high-fidelity, but thanks to Creoâs staged workflow, all of the participantâs previous choices are maintained. The room still reflects the spatial composition they defined in the composition stage and the material palette they specified in color. Even though their original prompt did not specify decisions like the color of the couch or the direction of the light, Creo did not fill those in for them. Instead, the participant was able to make each of those decisions as they walked through each of Creoâs stages, allowing them to leverage the power of generative image models while maintaining agency over their creative processâsomething they noted was a meaningful departure from the prompt-and-wait workflows they typically rely on. Figure 7. Creo decomposes image generation into multiple stages. From a prompt, it generates (1) multiple viewpoints, after which the illustrator (2) refines composition, (3) color, (4) lighting, and (5) style in any order. Stages can be completed in any order. Appendix D User Study Protocol and Analysis Details This appendix provides additional methodological details for the user study described in Section 4, including participant recruitment, study procedure, materials, and analysis methods. D.1. Participant Recruitment and Screening Participants were recruited through posts in online communities focused on illustration, art, and design. Interested individuals completed a short pre-screening survey about their creative background, experience with generative image tools, and availability. A total of 275 responses were collected. From these responses, we invited 19 participants whose profiles met our study criteria and represented a range of visual design domains including illustration and concept art, interior design and architectural visualization, comics and storyboarding, graphic design and marketing, and animation or motion design. Participants were required to be at least 18 years old and have prior hands-on experience using generative image tools as part of a creative workflow. The screening survey also collected information about which tools participants used (e.g., Midjourney, Stable Diffusion, DALL¡E, Adobe Firefly), how frequently they used them, and at which stages of their creative process they were typically employed. Participants received a gift card as compensation for their time. D.2. Study Design The study used a within-subjects design in which each participant completed two creative tasks using two different tools: ⢠Progressive staging interface (Creo) ⢠One-shot text-to-image baseline To prevent participants from refining the same concept across conditions, we used two related prompts: ⢠Design your ideal living room ⢠Design your ideal kitchen Each participant completed both prompts, using one prompt per tool. Tool order and prompt order were counterbalanced across participants using four sequences (Table 2). Sequence Tool Order Prompt Order A Creo â Baseline Living Room â Kitchen B Creo â Baseline Kitchen â Living Room C Baseline â Creo Living Room â Kitchen D Baseline â Creo Kitchen â Living Room Table 2. Counterbalancing design used in the study. D.3. Session Structure Each session lasted approximately 60 minutes and followed the structure below: ⢠Consent and study overview (4 minutes) ⢠Screener questions and workflow warm-up discussion (6 minutes) ⢠Tool A introduction (2 minutes) ⢠Tool A task (14 minutes) ⢠Post-condition reflection (7 minutes) ⢠Tool B introduction (2 minutes) ⢠Tool B task (14 minutes) ⢠Post-condition reflection (7 minutes) ⢠Final comparison and wrap-up (4 minutes) Sessions were conducted remotely via video conferencing with screen sharing enabled. D.4. Task Instructions Participants were asked to generate one image per condition using the assigned prompt. During the task, participants were instructed to think aloud and describe what they were noticing, attempting, or deciding between while working. Moderators occasionally used light probes to clarify participantsâ intentions, such as asking what they were trying to change, what they expected to happen, or why they shifted direction. Examples of probes included: ⢠âWhat are you trying right now?â ⢠âWhat made you try that?â ⢠âWhat were you hoping would change?â ⢠âWhy did you decide to keep this result?â After each condition, participants completed a short reflection discussing the outcome they produced, how the idea evolved during the task, and how the workflow compared to their usual creative process. After both conditions were completed, participants participated in a final comparison discussion about differences between the two systems. D.5. Materials and Setup The study used two tools: ⢠Creo, accessed through a browser-based interface. ⢠A prompt-centric one-shot text-to-image interface supporting iterative prompting and image generation. Participants interacted with both tools via screen sharing. Screen and audio recordings were captured for later analysis. For each session we also collected interaction artifacts including generated images, prompt histories, and stage histories produced during the session. We collected both interaction data and participant reflections. D.6. User Study Analysis Details This section describes the data processing, annotation scheme, and metric definitions used in our mixed-methods analysis comparing progressive staging (Creo) and one-shot workflows. Data Sources and Processing We analyzed screen recordings and audio from each session to reconstruct interaction histories, including prompts, edits, generated outputs, and participantsâ verbal reasoning. Sessions were segmented into discrete actions (e.g., prompting, modifying content, evaluating outputs, or correcting unintended changes). D.6.1. Research Questions and Metrics Our analysis is structured around three research questions: RQ1: Exploration vs. Anchoring. We examine how users explore and revise ideas during early ideation. Anchoring is measured via similarity between initial and final outputs and the number of direction changes. Exploration is characterized through iteration structure and the proportions of on-intent, pivot, and drift actions. RQ2: Control and Predictability. We evaluate how disentangling decisions affects control over outcomes. This includes unintended changes (invariant violations), user-driven interaction, and revision effort. RQ3: Non-linear Workflows. We analyze how users appropriate staged interaction, including stage skipping, revisiting earlier stages, and iterative refinement through repropagation. D.6.2. Action Annotation Each action was manually annotated along the following dimensions: Action type. ⢠Construct: Direct specification or modification of content ⢠Evaluate: Inspecting or assessing outputs ⢠Generate: Producing new outputs via prompting or regeneration ⢠Refine: Incremental adjustment of an existing idea ⢠Repair: Correcting unintended changes Intent. ⢠On-intent: Aligns with current design direction ⢠Pivot: Deliberate shift to a new direction ⢠Drift: Unintended deviation introduced by the model Agency. ⢠User-driven: Direct specification or modification ⢠Model-led: Generating and reacting to outputs Additional annotations. ⢠Direction change: Major shift in concept or trajectory ⢠Invariant violation: Unintended changes outside intended scope ⢠Iteration ID: Groups actions into exploration branches D.7. Metric Definitions All metrics are computed at the session level. Exploration and anchoring. ⢠Direction changes: Number of major shifts in trajectory ⢠Exploration breadth: Number of distinct iterations ⢠Concept drift: Proportion of drift actions ⢠Intentional pivots: Proportion of pivot actions Anchoring (image similarity). We compute cosine similarity between embeddings of the first and final generated images using OpenCLIP with L2-normalized embeddings. Control and predictability. ⢠Constructive engagement: Proportion of construct actions ⢠Evaluation-heavy behavior: Proportion of evaluate actions ⢠Perceived agency: Proportion of user-driven actions ⢠Revision burden: Ratio of repair actions ⢠Invariant violations: Proportion of violations Workflow structure. ⢠Iteration structure: Number and distribution of iterations ⢠Stage transitions (Creo): Revisiting and skipping stages ⢠Stage usage: Distribution and adoption across stages Aggregation Metrics are computed per session and aggregated across participants, reporting mean and standard deviation per condition. D.7.1. Qualitative Analysis We conducted a thematic analysis of think-aloud protocols and post-task interviews to understand experiences of authorship, control, predictability, and commitment. Themes were developed iteratively and used to interpret quantitative patterns. D.8. Scope and Limitations Because session durations were not consistently available, all metrics are based on action counts and proportions rather than time-based measures. Additionally, some constructs (e.g., perceived agency) are approximated using behavioral proxies derived from interaction logs. Appendix E System Design The figures depict inputs, outputs, user controls and constraints at each stage. Figure 8 depicts the system design of inputs, outputs and constraints at each stage, while Figure 9 depicts how editing instructions are passed across each of the 5 layers. Figure 8. Diagram depicting how editing instructions and constrains are dynamically passed to each layer. Figure 9. Diagram depicting how editing instructions and constrains passed across layers; revision loops etc.