Paper deep dive
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/19/2026, 4:58:00 AM
Summary
The paper introduces SAGE, a framework for automating storyboard generation in short drama production by evolving directing knowledge from expert demonstrations. SAGE addresses challenges in knowledge acquisition, refinement, and injection by extracting content-free rules, attributing feedback to specific rules, and routing relevant rules to narrative groups. The system was deployed on Virtual Film Studio, significantly reducing authoring time and achieving high acceptance rates. The authors also release PROSE, a dataset of professional screenplays and storyboards.
Entities (7)
Relation Signals (6)
PROSE â contains â Screenplays and Storyboards
confidence 95% ¡ PROSE, the first public dataset pairing screenplays with storyboards by professional directors
SAGE â deployedon â Virtual Film Studio
confidence 95% ¡ Deployed for 14 days on Virtual Film Studio
SAGE â produces â Storyboards
confidence 95% ¡ Storyboards turn screenplays into visual shot plans for automated short drama production... SAGE... learns... directing knowledge
SAGE â uses â Attribution-Guided Rule Evolution
confidence 95% ¡ SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge
SAGE â departsfrom â Manual Storyboarding
confidence 90% ¡ Manual storyboarding constrains throughput... SAGE... automates this step
SAGE â reduces â Authoring Time
confidence 90% ¡ production team recorded over 83 percent less authoring time per episode
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.17468v1
- Canonical: https://arxiv.org/abs/2608.17468v1
Trouble viewing inline? Open PDF directly â
Full Text
71,403 characters extracted from source content.
Expand or collapse full text
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule EvolutionCCS: Computing methodologies Artificial intelligence Maolin Ran Affiliation: Shanghai Jiao Tong University , Shanghai , China email: maolinr03@sjtu.edu.cn , Xiaoyang Lu Affiliation: Shanghai Jiao Tong University , Shanghai , China email: xiaoyangl@sjtu.edu.cn , Jiaqi Liu Affiliation: Shanghai Jiao Tong University , Shanghai , China email: jkliu189@gmail.com , Jian Wang Affiliation: CreativeFitting , Shanghai , China email: jim.wang@creativefitting.ai , Weiwen Liu Affiliation: Shanghai Jiao Tong University , Shanghai , China email: wwliu@sjtu.edu.cn , Jianghao Lin Affiliation: Shanghai Jiao Tong University , Shanghai , China email: linjianghao@sjtu.edu.cn , Yong Yu Affiliation: Shanghai Jiao Tong University , Shanghai , China email: yyu@sjtu.edu.cn and Weinan Zhang Note: Corresponding authors. Affiliation: Shanghai Jiao Tong University , Shanghai , China email: wnzhang@sjtu.edu.cn Abstract. Storyboards decompose screenplays into shot-by-shot visual plans that drive automated short drama production. Because high-quality storyboarding rests on the tacit expertise of professional directors, it remains a capacity bottleneck at industrial scale. Large language models can automate this step, yet existing ways of equipping them with directorial knowledge face three challenges: (1) Knowledge acquisition: the craft stays implicit in exemplars or must be authored by hand, so explicit knowledge exists only where a human writes it. (2) Knowledge refinement: authored knowledge is never evaluated against execution outcomes, and opaque generation prevents feedback from being attributed to the knowledge behind each decision. (3) Knowledge injection: injecting everything exceeds the usable context, yet hand-picking knowledge for every narrative group does not scale. In light of these challenges, we present SAGE (Skill with Attribution-Guided Evolution). SAGE is a deployed framework that learns, attributes, and evolves directing knowledge from expert demonstrations. It first extracts content-free rules by contrasting each training screenplay with its expert storyboard. During generation, the model declares which rules each narrative group adopts. Joining these records with localized feedback yields attribution at the level of individual rules, which drives targeted rule updates. Evolved rules are then consolidated into scenario packages accessed through a routing index. Each group therefore retrieves only a bounded set of scenario packages matched to its situation, without expert intervention. On 18 test episodes across three genres, SAGE scored 77.877.8 on an expert-validated rubric, exceeding professional directorsâ 77.177.1. Deployed for 14 days in Virtual Film Studio, a commercial short drama production platform, it produced 1,344 narrative group outputs. Of these, 87.2%87.2\% were accepted without substantive edits, and the production team recorded a drop of over 83%83\% in authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards authored by professional directors, spanning 68 episodes at https://github.com/creDreams/PROSE. Keywords: self-evolving agents, skill learning, credit assignment, storyboard generation 1. Introduction Figure 1. Camera language decisions on one screenplay beat from His Toyboy EP001, a held-out test episode, under an identical screenplay, backbone LLM, and storyboard schema. SAGE alone selects a low-angle push-in that reinforces Victorâs dominance, whereas Few-shot and CoT keep the angle flat and Vanilla widens to a composition that weakens the power dynamic. Only the textual records are system outputs; thumbnails and threat labels are post-hoc readings, excluded from evaluation.Five stacked rows compare four systems on one screenplay beat. The top row states the input screenplay line. The next four rows correspond to SAGE, Few-shot, CoT, and Vanilla. Each of these rows holds a textual storyboard record, three camera language fields for shot scale, camera angle, and camera movement, a greyscale thumbnail, and a threat level indicator. The SAGE row lists close-up, low angle, and push-in. The other three rows list flat angles and a static camera. A closing row summarizes the comparison. As video generation models mature, automated film production systems increasingly transform screenplays into videos (14; 40). In these systems, the storyboard is the intermediate artifact connecting creative intent to video synthesis (14; 53). It specifies, shot by shot, the visual content, shot scale, camera angle, and camera movement. Its quality directly constrains the fidelity, coherence, and cinematic expressiveness of the downstream video. Manual storyboarding constrains throughput across the short drama industry. Tens of thousands of serialized episodes ship annually on platforms such as Douyin and TikTok, yet the timeline from concept to release often spans only weeks or even days, which forces creators into multiple roles at once to meet the efficiency demands of the industry (4). Storyboards are still authored by hand, and this manual stage caps the speed of short drama production (37; 32). The bottleneck persists because storyboarding rests on tacit directorial expertise (26). A director jointly decides shot rhythm, visual description, shot scale, camera angle, and camera movement, yet the craft behind these decisions is rarely articulated as explicit principles. On the commercial platform studied in this work, a director spends more than one hour on a single episode. Large language models (LLMs) can automate this process. Benchmarks nevertheless show that frontier models lack professional competence in camera language (17; 34). Figure 1 illustrates this deficit on one dramatic beat, where direct generation defaults to a flat angle that does not convey the sceneâs power dynamic. Effective automation therefore requires equipping LLMs with this domain knowledge. Prior efforts inject such knowledge into LLMs through three mechanisms. Few-shot prompting supplies screenplay and storyboard exemplars, and chain-of-thought (CoT) prompting (36) adds reasoning chains written by experts. Skills (3) package a stable workflow together with separate knowledge files, a design that has become an industrial practice. These mechanisms let LLMs exploit directorial knowledge, yet three challenges remain. C1: Knowledge acquisition. In existing mechanisms, knowledge is either implicit in exemplars or authored by hand. Few-shot exemplars encode the craft implicitly, so the model must re-induce it at inference time and never obtains an inspectable, reusable form. CoT chains and skill files state the knowledge explicitly, but directors must author every chain and every file themselves. Explicit knowledge is thus available only where a human writes it, and the challenge is to extract it automatically from expert demonstrations. C2: Knowledge refinement. Once authored, the knowledge stays fixed and is never evaluated against execution outcomes, so its defects remain undetected. Refinement must therefore use execution feedback. Feedback alone is nevertheless insufficient, because generation is opaque: it does not reveal which piece of the injected knowledge shaped which decision. Existing pipelines judge knowledge at extraction time or by votes over whole trajectories, without tracing its use in generation (49; 43), and self-correction without such localization is reported to degrade performance (13). Effective refinement therefore requires generation to be traceable, so that feedback can be attributed to the knowledge responsible for each decision. C3: Knowledge injection. At inference time, the knowledge given to the model must be selected automatically. Current practice secures relevance by manual selection, and prior automatic retrieval still operates over knowledge units defined by hand (10). Manual selection cannot scale when every narrative group demands its own decision over thousands of entries, and injecting everything exceeds the usable context. We present SAGE (Skill with Attribution-Guided Evolution), a framework that resolves the three challenges on the skill substrate, which already separates a stable workflow from an explicit knowledge base. For acquisition, SAGE contrasts each training screenplay with its expert storyboard and extracts content-free rules automatically (C1). For refinement, generation declares which rules each narrative group adopts. Joining these rule-adoption records with localized feedback yields three evolution operations: misfiring rules are revised, new rules are added for coverage gaps, and unused rules are retired (C2). For injection, evolved rules are consolidated by semantic clustering into scenario packages with a routing index, so each narrative group retrieves only the packages matched to its situation (C3). On a test set of 18 episodes spanning three drama genres, SAGE scored 77.877.8, exceeding the professional directorsâ 77.177.1. All scores come from an LLM rubric whose agreement with three professional directors exceeds inter-human agreement (§4.6). It also outperformed strong baselines given reasoning chains authored by directors and exemplars from adjacent episodes. Attribution-guided iteration contributed 3.63.6 points over rule warm-start, whereas iteration without attribution peaked early and then fell below its starting point. The consolidated rules transferred unchanged to three other backbone LLMs, with relative gains of 8.1%8.1\% to 21.2%21.2\%. Our contributions are: ⢠Framework. SAGE treats the knowledge base of an LLM skill as learnable parameters and evolves it from expert demonstrations. To our knowledge, SAGE was the first framework to evolve such a knowledge base under credit assignment at the granularity of individual rules, rather than a single validation score over the whole skill document. The framework spans evolution, consolidation, and deployment, and it runs in commercial production. ⢠Mechanism. A rule-level attribution mechanism that assigns credit over knowledge expressed in natural language. It links authoring to execution feedback, a connection that static skills lack. Our ablation and iteration studies show that iteration without it peaks early and then declines. ⢠Dataset. We release PROSE, to our knowledge the first open dataset pairing screenplays with storyboards authored by professional human directors11 1 Publicly available at https://github.com/creDreams/PROSE.. It spans 68 episodes across three professionally produced series of distinct genres. Prior resources instead provide shot annotations that models reverse-engineer from finished videos (32; 53). ⢠Deployment. We deployed SAGE in Virtual Film Studio (VFS), a commercial short drama production platform, and report a 14-day study on three unseen ongoing dramas. Of 1,344 narrative group outputs, 87.2%87.2\% entered production without substantive edits. Authoring time per episode fell from over one hour to roughly 10 minutes, a reduction of over 83%83\%. Storyboarding therefore shifted from manual authoring to review, which is what removes the capacity bottleneck. 2. Related Work LLM-based storyboard generation. FilmAgent (40) and MovieAgent (38) coordinate role-playing agents. FilMaster (14) instead retrieves camera language conventions from 440K film clips. Neither learns an explicit, inspectable knowledge base from professional storyboard demonstrations. DramaDirector (53) is the closest system to our task, since it fine-tunes an LLM planner with SFT and GRPO for short drama storyboards, but it keeps cinematic knowledge implicit in model weights. A parallel line generates image storyboards (39; 7), addressing visual consistency rather than the directorial decomposition studied here, and earlier engine-based previsualization renders shot candidates under manually specified rules (28). Benchmark studies consistently find that frontier models lack professional competence in camera language (17; 34), which motivates explicit knowledge injection. Storyboard datasets. SkyScript-100M (32) and DramaBoard (53) pair short drama scripts with shot-level annotations. Their storyboards, however, are reverse-engineered from finished videos, so they capture what ended up on screen rather than the decisions the director made. PROSE instead releases the directorsâ original pre-production storyboards, the decision traces that demonstration-supervised evolution requires. Experiential knowledge in self-evolving agents. Self-evolving agents (11) improve from their own experience; we focus on the branch that evolves context rather than weights. Self-Refine (20) and Reflexion (30) iterate on individual outputs or store trajectory-level reflections, without accumulating knowledge that persists beyond the task. Unaided self-correction is also known to be unreliable (13). A second family distills experience into persistent natural-language knowledge, expressed as failure-derived rules, state-conditioned guidelines, causal abstractions, or distilled insights and procedural memory (43; 10; 21; 49; 8). A third family stores capabilities as executable skills. Voyager (33) grows a code skill library through environment feedback, and successors extend the idea to other interactive domains (31; 51; 35); industrial standards package such procedural instructions with resources (3), and memory architectures manage the state generically (24; 48). Across this family the maintenance signal is coarse: items are added or voted on from trajectory outcomes, without tracking which item influenced which output. Sound and harmful items therefore receive the same credit, and the noise grows as the store accumulates entries. AutoManual (5) comes closest, since its planner cites the rules it engages, but it learns from binary episodic rewards and revises rules by post-hoc judgment over whole trajectories. SAGE instead aligns against expert demonstrations, the only supervision available without an environment that verifies success, and joins per-group adoption records with localized feedback (§4.4). Agent-Pro (47) and AgentEvolver (45) also use reflection or self-attribution, but operate over policy prompts and RL action steps rather than declarative knowledge items. Natural-language component optimization. Manually supplied context is brittle, since in-context learning is sensitive to exemplar relevance and ordering, and selecting effective exemplars is itself a retrieval problem (18; 50; 19; 29). Treating the natural-language components of a frozen-LLM system as optimizable parameters began with instruction and evolutionary prompt search (54; 41; 9; 12). The idea then matured into gradient-like frameworks. ProTeGi (27) edits prompts along critique-derived textual gradients, while TextGrad (44) and Trace (6) backpropagate language feedback through computation graphs and execution traces. DSPy (15) compiles declarative pipelines with learnable instructions, and later frameworks train agent functions as weights (46) or show that language-level reflection can outperform RL in sample efficiency (1). Most recently the paradigm has reached skills. EvoSkill (2) edits skill folders from failure analysis, and SkillOpt (42) treats skill documents as trainable state. In both, the learning signal remains a single scalar validation score over the whole skill, so edits are retained or discarded in bulk and an individual ruleâs contribution is never measured. SAGE shares the view of knowledge as learnable parameters, but maintains rule-adoption records that assign credit to individual rules, so one round can revise, add, and retire different rules at once. Our experiments show that this granularity keeps iteration productive where document-level optimization plateaus (§4.2). 3. Methodology We present SAGE, a framework that enables the knowledge component of an LLM skill to evolve autonomously from expert demonstrations. 3.1. Problem Formulation Storyboard generation. A screenplay S consists of scenes with dialogue, action descriptions, and character information. The goal is to produce a storyboard B=â¨b1,âŚ,bnâŠB= b_1,âŚ,b_n , an ordered sequence of shots. Each shot bib_i is a structured tuple specifying visual content, shot scale, camera angle, camera movement, characters, and dialogue. Operationally, each scene is partitioned into narrative groups; every group yields a storyboard segment of one or more shots, and these segments are merged into B. Producing B requires joint decisions across five professional dimensions: shot rhythm, visual description, shot scale, camera angle, and camera movement. These dimensions encode tacit directorial knowledge that is difficult to specify exhaustively in a static prompt. Skill as workflow and knowledge. We define a skill as a pair ÎŁ=(,â) =(W,R). The workflow W is a fixed multi-step procedural scaffold. It specifies how the screenplay is partitioned, which knowledge is retrieved at each step, and how partial outputs are merged. The knowledge base âR holds the declarative rules that the workflow consumes. This decomposition follows emerging industrial standards for agent skills (3). Each rule rââr is a pair (1) r=(cond,prac),r=(cond,prac), where cond describes the narrative situation in which the rule applies, such as a shock reaction within a high-intensity dialogue. The second element prac is an executable directive, such as not inserting a breathing shot between that shock reaction and the follow-up question. Each rule is tagged with one of the five dimensions. Rules are further constrained to be content-free: they must not contain concrete shot content or verbatim material from expert storyboards, which prevents data leakage and encourages generalization. Learning view. The workflow W is easy to fix by design, whereas skill quality is dominated by âR. We therefore cast rule acquisition as optimization. Let gÎŁâ(S)g_ (S) denote the storyboard generated by an LLM equipped with skill ÎŁ . Let falignâ(gÎŁâ(S),Bâ)â[0,100]f_align(g_ (S),B^*)â[0,100] measure how closely that storyboard matches the expert reference BâB^* across the five dimensions. Given a corpus of expert demonstrations =(Sj,Bjâ)j=1mD=\(S_j,B_j^*)\_j=1^m, training seeks (2) ââ=argâĄmaxââ(S,Bâ)âźâ[falignâ(g(,â)â(S),Bâ)].R = _R\;E_(S,B^*) [f_align (g_(W,R)(S),B^* ) ]. This formulation yields a direct analogy to standard machine learning. The rule set âR plays the role of learnable parameters, and falignf_align acts as the training objective. The test metric is a separate reference-free quality score fqualf_qual, which measures absolute professional quality without access to BâB^*. Two properties distinguish this setting from gradient-based learning. The âparametersâ are discrete natural-language rules, and the optimization signal must be routed to individual rules through an explicit attribution mechanism (§3.3). 3.2. Framework Overview Figure 2. Overview of SAGE. Stage 1 (Evolution): for each training episode, a four-phase loop generates storyboards with per-group rule attribution, evaluates them against the expert reference, and revises the rule set through attribution-guided diagnosis. Stage 2 (Consolidation): rules evolved across all episodes are deduplicated, embedded, and clustered into scenario packages with a routing index. Stage 3 (Inference): on unseen screenplays, narrative groups are routed to their matched scenario packages, whose rules are injected into generation; no expert reference is required.A horizontal pipeline of four blocks. The leftmost block holds the screenplay and the expert storyboard. The Evolution block encloses a cycle of four numbered phases, namely rule extraction, attributed generation, alignment evaluation, and attribution-guided diagnosis, with rule-adoption records at the center and a loop labeled T rounds. The Consolidation block lists deduplication, condition embedding, clustering, and scenario labeling. The Knowledge Base block holds a routing index and scenario packages. The rightmost Inference block lists grouping, top-k routing, rule-injected generation, and merging into a storyboard. As shown in Figure 2, SAGE operates in three stages. All three share a unified generation pipeline of narrative grouping, scenario routing, rule injection, attributed generation, and segment merging. Stage 1 (evolution, §3.3) refines a per-episode rule set through a four-phase loop whose key ingredient is rule-level attribution. Stage 2 (consolidation, §3.4) deduplicates and clusters the large, redundant union of per-episode rule sets into scenario packages with a routing index. Stage 3 (inference, §3.5) routes each group of an unseen screenplay to its top-k packages and generates a segment with the retrieved rules injected. 3.3. Stage 1: Attribution-Guided Rule Evolution The evolution stage refines the rule set for each training episode over T rounds. Each round executes the four phases shown in Figure 3, and Algorithm 1 summarizes the loop. Figure 3. One round of attribution-guided rule evolution, on a real trace from Beyond the Wall. Attributed generation (Phase 2) records which rules each group adopted; diagnosis (Phase 4) joins these records with dimension-level feedback to decide, per rule, whether to revise, add, or retire.The upper part shows the four phases of one evolution round in sequence, annotated with a real trace in which a camera angle rule scores 63 out of 100 and is marked for revision. The lower part expands rule-level credit assignment. A rule-adoption record and a localized feedback item meet at a join operation, which produces a rule-level diagnosis. Three outcomes branch from the diagnosis, namely a revised rule, a newly added rule, and a retired rule. Phase 1: Rule Extraction. In the first round, an initial rule set â(1)R^(1) is extracted by contrasting screenplay with expert storyboard. The screenplay is first partitioned into a two-level hierarchy, in which scenes at level 1 are split into narrative groups at level 2. A group is a dialogue exchange, an action sequence, or an emotional beat, and is the minimal unit of analysis. For every group, the extractor examines how the director decomposed it into shots, then induces content-free rules (cond,prac)(cond,prac) that explain the observed decisions. Subsequent rounds inherit the revised rule set from the previous roundâs Phase 4. Phase 2: Attributed Generation. The model generates one segment per group from three inputs. These are the screenplay, a director vocabulary defining the legal shot scales, angles, and movements, and the current rule set â(t)R^(t). The expert storyboard is withheld. The defining feature of this phase is the rule-adoption record: for each group u, the model declares the adopted rules Auââ(t)A_u ^(t) alongside the segment it produces. The full attribution map (t)=(u,Au)uâgroupsâĄ(S)A^(t)=\(u,A_u)\_u (S) makes every generation decision traceable to the rules that informed it. Phase 3: Alignment Evaluation. The storyboard is scored against the expert reference by falignf_align, which produces an overall score, five per-dimension scores, and natural-language feedback that localizes each deviation. One such deviation reads that the confrontation in scene 4 lacks a re-establishing two-shot after three consecutive close-ups. The expert storyboard is visible only to the evaluator, never to the generator. Appendix C details the review protocol. Phase 4: Attribution-Guided Diagnosis. Diagnosis joins the feedback with (t)A^(t) to perform rule-level credit assignment. Deviations fall into two classes with distinct remedies. In Class A (misfiring rule), a dimension deviates in groups where some rule r is adopted. That rule is implicated, so its condition or practice is revised. In Class B (coverage gap), a dimension deviates in groups where no adopted rule governs it. No existing rule is at fault, so a new rule is induced from the feedback. Rules never adopted throughout the episode are retired, which keeps the set minimal. At most 30%30\% of rules may be modified per round, a cap that prevents destructive oscillation. The revised set â(t+1)R^(t+1) seeds the next round. Without attribution, feedback can only be assigned at the episode level. The optimizer then knows that a dimension scored poorly but not which rule caused it, so revisions become undirected rewrites. Our ablations identify this shift from episode-level to rule-level credit assignment as the condition for sustained improvement (§4). Algorithm 1 Attribution-Guided Rule Evolution for a Single Episode 0: screenplay S, expert storyboard BâB^*, rounds T 0: evolved rule set â(T+1)R^(T+1) 1: UâGroupâĄ(S)U (S) two-level narrative grouping 2: â(1)âExtractRulesâĄ(S,Bâ,U)R^(1) (S,B^*,U) Phase 1 3: for t=1t=1 to T do 4: (B(t),(t))âAttrGenâĄ(S,U,â(t))(B^(t),A^(t)) (S,U,R^(t)) Phase 2 5: (s(t),F(t))âfalignâ(B(t),Bâ)(s^(t),F^(t))â f_align(B^(t),B^*) Phase 3 6: for all deviations dâF(t)dâ F^(t) do 7: if ârâAuâ\,râ A_u governing dim(d) (d) for the group u of d then 8: revise r Class A: misfiring rule 9: else 10: â(t)ââ(t)âŞInduceRuleâĄ(d)R^(t) ^(t)âŞ\InduceRule(d)\ Class B: gap 11: end if 12: end for 13: retire rules never adopted in (t)A^(t) 14: â(t+1)âR^(t+1)â revised set â¤30%⤠30\% of rules modified 15: end for 3.4. Stage 2: Rule Consolidation Evolution is per-episode by design. The union across the corpus reaches the order of 10210^2 rules per episode and several thousand in total, and is redundant and in places contradictory. Consolidation compresses it into a retrievable knowledge base in three steps. Deduplication. A two-level union-find procedure runs within each dimension. At level 1, rules whose condition embeddings exceed a cosine similarity of 0.900.90 are merged into a condition group. At level 2, within each condition group, rules whose practice embeddings also exceed the threshold are collapsed to their semantic centroid. Practices below the threshold are preserved as alternative practices of a single multi-practice rule. One pass thus resolves both redundancy, where condition and practice coincide, and latent contradiction, where a shared condition maps to divergent practices. The threshold is deliberately conservative because the two error directions are not symmetric. Merging rules that differ in meaning destroys knowledge irrecoverably, whereas failing to merge equivalent rules only leaves redundancy, since divergent practices survive as alternatives of one rule instead of being collapsed into a centroid. Embedding and clustering. Each deduplicated rule is represented by its condition embedding, capturing the narrative situation it targets. Embeddings are â2 _2-normalized, reduced with UMAP (22), and clustered with k-means. A grid search selects the number of clusters and the UMAP hyperparameters, jointly scoring dimension coverage, cluster size compliance, size uniformity, and silhouette quality. The number of scenario packages is therefore determined by the data rather than fixed a priori. Rules triggered by similar situations thus become co-located and co-retrieved, regardless of their source episode or drama. Scenario packaging. An LLM agent labels each cluster with a human-readable scenario name, such as emotional climax under psychological pressure, together with a short applicability description. The agent then materializes the cluster as a scenario package, a document that groups the clusterâs rules by dimension. A compact routing index is built alongside, listing every packageâs name, description, and rule inventory. Appendix A shows the resulting cluster structure and an example package. 3.5. Stage 3: Scenario-Aware Inference At deployment the evolved skill runs on unseen screenplays with no expert reference and no iteration, reusing the training pipeline. The screenplay is partitioned into the same two-level hierarchy used during evolution. For each group, the model matches its situation against the routing index and selects the top-k packages, recording a justification per match. We set k=3k=3 to match the three reference episodes supplied to Few-shot and CoT, which equalizes the injection budget across knowledge-injection methods. Routing over the index rather than scanning all rules keeps the injected context bounded as the knowledge base grows. Each groupâs segment is then generated independently and in parallel, conditioned on the groupâs screenplay content, the retrieved packages, and the director vocabulary. Generation also emits a rule-adoption record, which preserves traceability in deployment. Finally, segments are concatenated in screenplay order and their shots renumbered into the final storyboard. Because packages encode situation-conditioned knowledge rather than model-specific tricks, the consolidated base is backbone-agnostic and can be injected into other LLMs unchanged. Appendix B traces one real group through routing and generation. 4. Experiments We evaluate SAGE around four research questions. (RQ1) Does the evolved skill close the quality gap to professional directors, and how does it compare with strong prompting and skill optimization baselines? (RQ2) How much does each component contribute, namely rule warm-start, iteration, and attribution? (RQ3) Does attribution make quality improve over rounds instead of fluctuating? (RQ4) Is the consolidated knowledge base portable across backbone LLMs? 4.1. Experimental Setup Dataset. PROSE comprises three professionally produced short drama series of distinct genres, namely a sci-fi suspense series of 20 episodes (Beyond the Wall), an urban romance series of 23 (His Toyboy), and an emotional healing series of 25 (My Cure). Each episode pairs a screenplay, comprising a synopsis, character profiles, and a scene-level script, with the storyboard authored by the seriesâ professional director. We held out 6 episodes per series, 18 in total, as the test set. The consolidated knowledge base is built exclusively from rules evolved on the remaining 50 training episodes. Evaluation protocol. All systems were scored by a reference-free quality rubric on the five dimensions of shot rhythm, visual description, shot scale, camera angle, and camera movement. Each dimension uses a 100-point scale, and the overall score is their average. Scoring used Claude Opus 4.6 under a fixed rubric prompt, whose score anchors are given in Appendix D. This metric is distinct from the alignment score falignf_align used as the training signal: quality measures how good a storyboard is against professional standards, whereas alignment measures how close it is to a specific expert reference. The scorerâs reliability is validated against human experts in §4.6. Baselines. We compared six alternatives under the same backbone (Claude Opus 4.6) and output schema. Director is the human storyboard, which serves as the expert reference. Vanilla generates directly with no external knowledge. Few-shot is conditioned on ⨠, storyboard⊠pairs from the three nearest neighboring episodes of the same series, never the target episode itself. CoT adds reasoning chains that encode decomposition thinking authored by directors on those same reference episodes. EvoSkill (2) and SkillOpt (42) are representative skill optimization methods, reimplemented faithfully on the same test set. The strong baselines thus receive demonstrations from adjacent episodes, whereas SAGE uses none at inference, which makes the comparison conservative for our method. Implementation. Rule evolution ran T=10T=10 rounds per training episode with at most 30%30\% of rules modified per round. Conditions were embedded with Qwen3-Embedding-8B; consolidation yielded 5555 scenario packages from 2,0362,036 deduplicated rules. Inference routed each group to its top-3 packages. Unless stated otherwise, SAGE results use the round-5 knowledge base, which is the best-performing round on the test set. The no-attribution ablation is likewise reported at its own best round (§4.4), so this oracle round selection is applied symmetrically and characterizes each variantâs upper bound. 4.2. Main Results (RQ1) Table 1. Main comparison on the 18-episode test set (quality scores, 100-point scale). Bold: best among AI systems; underline: exceeds the human director. Method Rhythm Visual Scale Angle Move. Overall Director 80.1 75.1 81.7 76.8 71.6 77.1 Vanilla 68.2 67.4 74.4 62.1 53.8 65.2 Few-shot 73.1 73.3 78.2 68.0 64.1 71.2 CoT 79.4 80.8 79.5 72.1 68.4 76.0 EvoSkill 73.9 72.2 74.3 62.8 59.7 68.7 SkillOpt 78.9 74.9 79.5 74.4 68.1 75.1 SAGE (ours) 79.2 85.1 78.8 74.6 71.2 77.8 Table 1 reports the main comparison, from which three findings emerge. Expert-level quality. At 77.877.8 overall, SAGE was the only AI system to exceed the human director at 77.177.1. It was also the closest system to the director on camera angle and camera movement, the two dimensions on which Vanilla scored lowest. Knowledge versus exemplars. Both prompting baselines received demonstrations from adjacent episodes and reasoning authored by directors, yet Few-shot reached only 71.271.2 and CoT 76.076.0. SAGE encodes the same knowledge explicitly, instead of leaving it latent in exemplars for the model to induce anew. Generic skill optimization. EvoSkill and SkillOpt both scored below SAGE, with their largest deficits on camera movement, the dimension with the lowest scores overall. SAGE thus gained most where Vanilla was weakest, and its visual description even surpassed the director. We attribute this to evolved rules that enforce compositional completeness in lighting, blocking, and framing, which human storyboards often leave implicit. 4.3. Ablation Study (RQ2) Table 2. Ablation on the 18-episode test set. Each row adds one component. Configuration Overall Î A Vanilla (no external knowledge) 65.2 â- B + rule warm-start (no iteration) 74.2 +9.0+9.0 C + iteration (no attribution) 75.4 +1.2+1.2 D + attribution (full SAGE) 77.8 +2.4+2.4 Table 2 isolates each component. The contrastive warm-start in row B provided the largest single gain, since rules extracted by contrasting screenplays with expert storyboards already capture substantial explicit knowledge. Iteration without attribution in row C added little, because episode-level feedback cannot identify which rules to fix. Adding attribution in row D more than doubled the iteration benefit, with its largest gains on shot scale and visual description, the two dimensions where row C remained weakest. This pattern confirms the argument of §3.3: rule-level credit assignment converts iteration from perturbation into optimization. 4.4. Iteration Dynamics (RQ3) A line chart with evolution round from 1 to 10 on the horizontal axis and overall quality score from 73 to 78 on the vertical axis. Both lines start together at 74.2, marked by a dotted horizontal reference line. The solid line for the setting with attribution climbs to a peak of 77.8 at round 5 and then stays close to that level. The dashed line for the setting without attribution peaks at 75.4 at round 3 and then declines to 73.8, ending below the starting level. Figure 4. Quality on the test set over 10 evolution rounds. Both settings share the same round-1 rule set at 74.2. With attribution, quality rises to 77.8 by round 5 and stays within 77.4 to 77.7 through round 10; without attribution, it peaks at 75.4 in round 3 and degrades to 73.8 by round 10.A line chart with evolution round from 1 to 10 on the horizontal axis and overall quality score from 73 to 78 on the vertical axis. Both lines start together at 74.2, marked by a dotted horizontal reference line. The solid line for the setting with attribution climbs to a peak of 77.8 at round 5 and then stays close to that level. The dashed line for the setting without attribution peaks at 75.4 at round 3 and then declines to 73.8, ending below the starting level. Figure 4 tracks quality across 10 rounds from an identical warm-start set. With attribution, quality rose over the first five rounds and then held stable through round 10. Without attribution, it peaked earlier at a lower value and ended below its starting point, so attribution changes both the magnitude and the stability of improvement. 4.5. Cross-Model Generalization (RQ4) A grouped bar chart over three backbone models, namely GPT-5.4, GLM-5.2, and DeepSeek V4-Pro. Each model has a Vanilla bar and a bar for the same model with SAGE rules injected. The Vanilla scores are 71.6, 60.4, and 59.0. The scores with SAGE rules are 77.4, 71.7, and 71.5. Relative improvements of 8.1 percent, 18.7 percent, and 21.2 percent are printed above the second bar of each group. Figure 5. Cross-model transfer of the consolidated knowledge base, evolved entirely with Claude Opus 4.6. Injecting the unchanged scenario packages improves every transfer target; labels above the SAGE bars report relative improvements over Vanilla.A grouped bar chart over three backbone models, namely GPT-5.4, GLM-5.2, and DeepSeek V4-Pro. Each model has a Vanilla bar and a bar for the same model with SAGE rules injected. The Vanilla scores are 71.6, 60.4, and 59.0. The scores with SAGE rules are 77.4, 71.7, and 71.5. Relative improvements of 8.1 percent, 18.7 percent, and 21.2 percent are printed above the second bar of each group. If the evolved rules encode directorial domain knowledge rather than backbone-specific tricks, they should transfer to other LLMs unchanged. We injected the identical scenario packages, evolved entirely with Claude Opus 4.6, into three backbones and modified nothing else in the pipeline. As Figure 5 shows, every target improved by 8.1%8.1\% to 21.2%21.2\% relative, and GPT-5.4 approached the source modelâs own quality.22 2 GLM-5.2 was evaluated on 16 of the 18 episodes owing to two routing failures, with the matching Vanilla episodes excluded for parity. Two patterns are notable. First, weaker backbones benefited more, because the rules supply structure that compensates for missing domain knowledge, whereas the strongest backbone gained least from the highest baseline. Second, visual description transferred most universally, which makes compositional checklists the most portable evolved knowledge. Absolute scores nonetheless tracked backbone capability, since rules supply domain knowledge but not generation capability. 4.6. Validity of the Automatic Scorer Table 3. Scorer validity on 18 director storyboards: agreement of Claude Opus 4.6 with the consensus of three professional directors, vs. inter-human agreement, measured by Linâs C. Dimension Claude vs. human consensus Inter-human Shot rhythm 0.717 0.689 Visual description 0.893 0.664 Shot scale 0.674 0.708 Camera angle 0.873 0.809 Camera movement 0.873 0.627 Mean 0.806 0.699 All reported scores come from an LLM scorer, a paradigm whose reliability and biases are well documented (52; 25). We therefore validated it against human judgment. Three professional directors and the scorer independently scored the 18 director storyboards on the five dimensions, yielding 90 score pairs. Table 3 reports Linâs concordance correlation coefficient (16). The scorerâs agreement with the human consensus exceeded inter-human agreement on four of the five dimensions, the sole exception being shot scale. Human scores were on average lower by a small margin, a systematic offset that does not affect relative rankings. Self-preference bias (25) is also unlikely to favor our method, since all AI systems in Table 1 share the scorerâs backbone and any such bias applies uniformly. We conclude the scorer is a reliable proxy for expert judgment in this domain. 5. Production Deployment CreativeFitting is an AI native entertainment company based in Shanghai. It operates Reel.AI, among the first AI generated short drama apps distributed to overseas audiences on the App Store and Google Play, and VFS, its in-house creation platform on which over a thousand creators produce content. To test whether offline gains translate into production value, we deployed SAGE in VFS, using the same framework trained on a larger proprietary corpus of director demonstrations. The evaluation covered three ongoing productions disjoint from the public 68-episode corpus. Over 14 days, platform logs recorded 12 production users, 1,344 narrative group outputs, and 2,038 generation and revision operations. We computed acceptance at the narrative group level, the unit of independent generation. Following the production teamâs operational criterion, an output is accepted if it can enter downstream production without substantive edits, where changes limited to asset references, formatting, punctuation, or wording count as non-substantive. Table 4. Production acceptance on three ongoing dramas. Each output corresponds to one narrative group and is a segment that may contain multiple shots. Prod. A Prod. B Prod. C Overall Narrative group outputs 559 453 332 1,344 Accepted outputs 460 380 332 1,172 Acceptance (%) 82.3 83.9 100.0 87.2 Table 4 shows that 1,172 of 1,344 outputs were accepted without substantive edits, an overall rate of 87.2%87.2\% that ranged from 82.3%82.3\% to 100.0%100.0\% across the three productions. The production team further reported that typical authoring time per episode fell from over one hour to roughly 10 minutes, an approximately sixfold acceleration. The acceptance rate is computed from platform interaction logs, whereas the turnaround was tracked by the production team over the same period. Two properties of the deployed pipeline keep the residual manual effort bounded. The unit of acceptance coincides with the unit of generation, so a rejected output calls for a local regeneration of one narrative group rather than a revision pass over the episode. The system also emits rule-adoption records at inference time (§3.5), so every rejected output stays traceable to the rules that informed it. 6. Conclusion Professional storyboarding depends on directorial knowledge that experts cannot exhaustively articulate, so every existing injection path relies on manual externalization. SAGE removes this dependence by evolving the knowledge component of a skill from the expert demonstrations released in PROSE. Rule-adoption records route feedback to individual rules, and the evolved rules are consolidated into scenario packages for deployment without a reference. The evolved knowledge exceeded the professional directors on our test set, transferred unchanged to three other backbones, and held these gains in a production deployment on live dramas. Our iteration study also generalizes beyond storyboarding. The granularity of credit assignment determines whether knowledge evolution converges: rule-level attribution reached a stable optimum, whereas episode-level feedback declined. Systems that treat natural-language knowledge as learnable parameters therefore need to localize feedback to individual items. This requirement is architectural rather than domain specific, since any pipeline whose generation step declares the knowledge it consumed can route feedback to that knowledge. References Agrawal et al. (2026) L. A. Agrawal, S. Tan, D. Soylu, N. Ziems, R. Khare, K. Opsahl-Ong, A. Singhvi, H. Shandilya, M. J. Ryan, M. Jiang, C. Potts, K. Sen, A. G. Dimakis, I. Stoica, D. Klein, M. Zaharia, and O. Khattab GEPA: reflective prompt evolution can outperform reinforcement learning. In The Fourteenth International Conference on Learning Representations (ICLR), Note: Oral. arXiv:2507.19457 Cited by: §2. Alzubi et al. (2026) S. Alzubi, N. Provenzano, J. Bingham, W. Chen, and T. Vu EvoSkill: automated skill discovery for multi-agent systems. arXiv preprint arXiv:2603.02766. Cited by: §2, §4.1. Anthropic (2025) Anthropic Equipping agents for the real world with agent skills. Note: https://w.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skillsEngineering blog; Agent Skills open standard. Accessed 2026-07-13 Cited by: §1, §2, §3.1. Cao et al. (2026) G. Cao, T. He, Y. Liu, and RAY LC Audience in the loop: viewer feedback-driven content creation in micro-drama production on social media. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, External Links: Document Cited by: §1. Chen et al. (2024) M. Chen, Y. Li, Y. Yang, S. Yu, B. Lin, and X. He AutoManual: constructing instruction manuals by LLM agents via interactive environmental learning. In Advances in Neural Information Processing Systems 37 (NeurIPS), Cited by: §2. Cheng et al. (2024) C. Cheng, A. Nie, and A. Swaminathan Trace is the next AutoDiff: generative optimization with rich feedback, execution traces, and LLMs. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2406.16218 Cited by: §2. Dinkevich et al. (2025) D. Dinkevich, M. Levy, O. Avrahami, D. Samuel, and D. Lischinski Story2Board: a training-free approach for expressive storyboard generation. arXiv preprint arXiv:2508.09983. Cited by: §2. Fang et al. (2026) R. Fang, Y. Liang, X. Wang, J. Wu, S. Qiao, P. Xie, F. Huang, H. Chen, and N. Zhang Memp: exploring agent procedural memory. In Findings of the Association for Computational Linguistics: ACL 2026, Note: arXiv:2508.06433 Cited by: §2. Fernando et al. (2024) C. Fernando, D. Banarse, H. Michalewski, S. Osindero, and T. Rocktäschel Promptbreeder: self-referential self-improvement via prompt evolution. In Proceedings of the 41st International Conference on Machine Learning (ICML), p. 13481â13544. Note: arXiv:2309.16797 Cited by: §2. Fu et al. (2024) Y. Fu, D. Kim, J. Kim, S. Sohn, L. Logeswaran, K. Bae, and H. Lee AutoGuide: automated generation and selection of context-aware guidelines for large language model agents. In Advances in Neural Information Processing Systems 37 (NeurIPS), Cited by: §1, §2. Gao et al. (2026) H. Gao, J. Geng, W. Hua, M. Hu, X. Juan, H. Liu, S. Liu, J. Qiu, X. Qi, Y. Wu, H. Wang, et al. A survey of self-evolving agents: what, when, how, and where to evolve on the path to artificial super intelligence. Transactions on Machine Learning Research. Note: arXiv:2507.21046 Cited by: §2. Guo et al. (2024) Q. Guo, R. Wang, J. Guo, B. Li, K. Song, X. Tan, G. Liu, J. Bian, and Y. Yang Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. In The Twelfth International Conference on Learning Representations (ICLR), Note: arXiv:2309.08532 Cited by: §2. Huang et al. (2024) J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou Large language models cannot self-correct reasoning yet. In The Twelfth International Conference on Learning Representations (ICLR), Cited by: §1, §2. Huang et al. (2025) K. Huang, Y. Huang, X. Wang, Z. Lin, X. Ning, P. Wan, D. Zhang, Y. Wang, and X. Liu FilMaster: bridging cinematic principles and generative AI for automated film generation. arXiv preprint arXiv:2506.18899. External Links: Document Cited by: §1, §2. Khattab et al. (2024) O. Khattab, A. Singhvi, P. Maheshwari, Z. Zhang, K. Santhanam, S. Vardhamanan, S. Haq, A. Sharma, T. T. Joshi, H. Moazam, H. Miller, M. Zaharia, and C. Potts DSPy: compiling declarative language model calls into state-of-the-art pipelines. In The Twelfth International Conference on Learning Representations (ICLR), Note: arXiv:2310.03714 External Links: Link Cited by: §2. Lin (1989) L. I. Lin A concordance correlation coefficient to evaluate reproducibility. Biometrics 45 (1), p. 255â268. Cited by: §4.6. Liu et al. (2025) H. Liu, J. He, Y. Jin, D. Zheng, Y. Dong, et al. ShotBench: expert-level cinematic understanding in vision-language models. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2506.21356 Cited by: §1, §2. Liu et al. (2022) J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen What makes good in-context examples for GPT-3?. In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, p. 100â114. Cited by: §2. Lu et al. (2022) Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp Fantastically ordered prompts and where to find them: overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), p. 8086â8098. Cited by: §2. Madaan et al. (2023) A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark Self-refine: iterative refinement with self-feedback. In Advances in Neural Information Processing Systems 36 (NeurIPS), Cited by: §2. Majumder et al. (2023) B. P. Majumder, B. Dalvi Mishra, P. Jansen, O. Tafjord, N. Tandon, L. Zhang, C. Callison-Burch, and P. Clark CLIN: a continually learning language agent for rapid task adaptation and generalization. arXiv preprint arXiv:2310.10134. Cited by: §2. McInnes et al. (2018) L. McInnes, J. Healy, and J. Melville UMAP: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. Cited by: §3.4. Murch (2001) W. Murch In the blink of an eye: a perspective on film editing. 2nd edition, Silman-James Press, Los Angeles, CA. Cited by: Appendix D. Packer et al. (2023) C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez MemGPT: towards LLMs as operating systems. arXiv preprint arXiv:2310.08560. Cited by: §2. Panickssery et al. (2024) A. Panickssery, S. R. Bowman, and S. Feng LLM evaluators recognize and favor their own generations. In Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Cited by: §4.6. Polanyi (1966) M. Polanyi The tacit dimension. Doubleday, Garden City, NY. Cited by: §1. Pryzant et al. (2023) R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng Automatic prompt optimization with âgradient descentâ and beam search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 7957â7968. Note: arXiv:2305.03495 Cited by: §2. Rao et al. (2023) A. Rao, X. Jiang, Y. Guo, L. Xu, L. Yang, L. Jin, D. Lin, and B. Dai Dynamic storyboard generation in an engine-based virtual environment for video production. In ACM SIGGRAPH 2023 Posters, New York, NY, USA, p. 1â2. Note: arXiv:2301.12688 External Links: Document Cited by: §2. Rubin et al. (2022) O. Rubin, J. Herzig, and J. Berant Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), p. 2655â2671. Cited by: §2. Shinn et al. (2023) N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems 36 (NeurIPS), p. 8634â8652. Cited by: §2. Tan et al. (2025) W. Tan, W. Zhang, X. Xu, H. Xia, Z. Ding, B. Li, B. Zhou, J. Yue, J. Jiang, Y. Li, R. An, M. Qin, C. Zong, L. Zheng, Y. Wu, X. Chai, Y. Bi, T. Xie, P. Gu, X. Li, C. Zhang, L. Tian, C. Wang, X. Wang, B. F. Karlsson, B. An, S. Yan, and Z. Lu Cradle: empowering foundation agents towards general computer control. In Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267, Note: arXiv:2403.03186 Cited by: §2. Tang et al. (2024) J. Tang, Q. Jia, Y. Xie, Z. Gong, X. Wen, J. Zhang, Y. Guo, G. Chen, and J. Yang SkyScript-100m: 1,000,000,000 pairs of scripts and shooting scripts for short drama. arXiv preprint arXiv:2408.09333. Cited by: 3rd item, §1, §2. Wang et al. (2024) G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar Voyager: an open-ended embodied agent with large language models. Transactions on Machine Learning Research. Note: arXiv:2305.16291 Cited by: §2. Wang et al. (2025a) X. Wang, S. Xu, X. Shan, Y. Zhang, M. Diao, X. Duan, Y. Huang, K. Liang, and Z. Ma CineTechBench: a benchmark for cinematographic technique understanding and generation. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025), p. 60372â60408. Note: arXiv:2505.15145 External Links: Document Cited by: §1, §2. Wang et al. (2025b) Z. Z. Wang, J. Mao, D. Fried, and G. Neubig Agent workflow memory. In Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267, Note: arXiv:2409.07429 Cited by: §2. Wei et al. (2022) J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35 (NeurIPS), p. 24824â24837. Note: arXiv:2201.11903 Cited by: §1. Wei et al. (2025) Z. Wei, H. Wu, L. Zhang, X. Xu, Y. Zheng, P. Hui, M. Agrawala, H. Qu, and A. Rao CineVision: an interactive pre-visualization storyboard system for directorâcinematographer collaboration. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST), External Links: Document Cited by: §1. Wu et al. (2025) W. Wu, Z. Zhu, and M. Z. Shou Automated movie generation via multi-agent CoT planning. arXiv preprint arXiv:2503.07314. Cited by: §2. Xie et al. (2024) J. Xie, J. Feng, Z. Tian, K. Q. Lin, Y. Huang, et al. Learning long-form video prior via generative pre-training. arXiv preprint arXiv:2404.15909. Cited by: §2. Xu et al. (2025) Z. Xu, L. Wang, J. Wang, Z. Li, S. Shi, et al. FilmAgent: a multi-agent framework for end-to-end film automation in virtual 3d spaces. arXiv preprint arXiv:2501.12909. Cited by: §1, §2. Yang et al. (2024) C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen Large language models as optimizers. In The Twelfth International Conference on Learning Representations (ICLR), Note: arXiv:2309.03409 Cited by: §2. Yang et al. (2026) Y. Yang, Z. Gong, W. Huang, Q. Yang, Z. Zhou, Z. Huang, Y. Li, X. Gao, Q. Dai, B. Liu, K. Qiu, Y. Yang, D. Chen, X. Yang, and C. Luo SkillOpt: executive strategy for self-evolving agent skills. arXiv preprint arXiv:2605.23904. Cited by: §2, §4.1. Yang et al. (2023) Z. Yang, P. Li, and Y. Liu Failures pave the way: enhancing large language models through tuning-free rule accumulation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 1751â1777. Cited by: §1, §2. Yuksekgonul et al. (2025) M. Yuksekgonul, F. Bianchi, J. Boen, S. Liu, P. Lu, Z. Huang, C. Guestrin, and J. Zou Optimizing generative AI by backpropagating language model feedback. Nature 639, p. 609â616. Note: Framework known as TextGrad; preprint: arXiv:2406.07496 Cited by: §2. Zhai et al. (2025) Y. Zhai, S. Tao, C. Chen, et al. AgentEvolver: towards efficient self-evolving agent system. arXiv preprint arXiv:2511.10395. Cited by: §2. Zhang et al. (2024a) S. Zhang, J. Zhang, J. Liu, L. Song, C. Wang, R. Krishna, and Q. Wu Offline training of language model agents with functions as learnable weights. In Proceedings of the 41st International Conference on Machine Learning (ICML), p. 60315â60335. Note: arXiv:2402.11359 Cited by: §2. Zhang et al. (2024b) W. Zhang, K. Tang, H. Wu, M. Wang, Y. Shen, G. Hou, Z. Tan, P. Li, Y. Zhuang, and W. Lu Agent-pro: learning to evolve via policy-level reflection and optimization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), p. 5348â5375. Cited by: §2. Zhang et al. (2025) Z. Zhang, Q. Dai, X. Bo, C. Ma, R. Li, X. Chen, J. Zhu, Z. Dong, and J. Wen A survey on the memory mechanism of large language model-based agents. ACM Transactions on Information Systems 43 (6). Note: arXiv:2404.13501 External Links: Document Cited by: §2. Zhao et al. (2024) A. Zhao, D. Huang, Q. Xu, M. Lin, Y. Liu, and G. Huang ExpeL: LLM agents are experiential learners. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 19632â19642. Cited by: §1, §2. Zhao et al. (2021) Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh Calibrate before use: improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning (ICML), p. 12697â12706. Cited by: §2. Zheng et al. (2025) B. Zheng, M. Y. Fatemi, X. Jin, Z. Z. Wang, A. Gandhi, Y. Song, Y. Gu, J. Srinivasa, G. Liu, G. Neubig, and Y. Su SkillWeaver: web agents can self-improve by discovering and honing skills. arXiv preprint arXiv:2504.07079. Cited by: §2. Zheng et al. (2023) L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track, Cited by: §4.6. Zhou et al. (2026) H. Zhou, S. Liu, J. Chen, X. Zou, L. Xia, and L. Nie DramaDirector: geometry-guided short drama generation. arXiv preprint arXiv:2606.24107. External Links: Document Cited by: 3rd item, §1, §2, §2. Zhou et al. (2023) Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba Large language models are human-level prompt engineers. In The Eleventh International Conference on Learning Representations (ICLR), Note: arXiv:2211.01910 Cited by: §2. Appendix A Rule Consolidation Details Figure 6 shows that rules learned across episodes and series form coherent, well-separated clusters (a). Each cluster becomes a self-contained scenario package, routable by its natural-language description and organized by professional dimension (b). This is how evolved knowledge is stored and retrieved at inference time (§3.4). Two properties are worth noting. First, clusters mix rules from all three series rather than separating by source, indicating that the learned conditions describe narrative situations rather than one dramaâs idiosyncrasies. The example package contains 38 rules contributed by all three series, and the recovered scenario carries no trace of its source episodes. Second, the distribution across dimensions is intentionally uneven: camera angle dominates this power-asymmetry package, whereas an emotional-release package concentrates on shot scale and rhythm. Consolidation therefore preserves each situationâs dimensional signature rather than balancing dimensions artificially. (a) (b) Figure 6. Rule consolidation. (a) UMAP projection of rule-condition embeddings, colored by cluster. (b) A representative scenario package shown in translation: a routable natural-language description over rules grouped by professional dimension. The package contains 38 rules in total; representative rules are shown for space.Panel (a) is a two-dimensional UMAP scatter plot of rule condition embeddings. Points form many small compact groups, each drawn in a distinct color, separated by empty space. Panel (b) shows one scenario package as a document. A header names the package and gives its applicability description, the rule count of 38, the three contributing series, and the rule counts per dimension. Three framed entries follow, one each for camera angle, shot scale, and shot rhythm, and each entry states a condition and a practice. A footer notes that 35 further rules are omitted. Appendix B Scenario-Aware Inference Example Figure 7 traces one real narrative group from a test episode through the scenario-aware inference pipeline of §3.5. A two-person confrontation is matched against the routing index and dispatched to its top-3 scenario packages, whose rules are injected into generation; the produced shots then carry rule-adoption records back to the packages that informed them. No expert reference is involved. The three retrieved packages are complementary rather than redundant. One supplies the angle vocabulary for power asymmetry, another governs the rhythm of a verbal exchange, and the third covers the framing of a physical intrusion. Their rules therefore act on different shots of the same segment, and the adoption records make this division visible after the fact, which is what allows a rejected output to be traced to the rule that shaped it. Figure 7. A real routing example from Beyond the Wall EP001, a held-out test episode, translated. Each shot is marked with the primary package whose rules it adopted, and one representative rule per package is shown.A left to right flow in three columns. The left column describes one narrative group, a two-person confrontation, with its dialogue lines. The middle column lists the three scenario packages it was routed to, each with a identifier, a name, a condition, and a practice. The right column shows the generated shots, each labeled with its shot scale, camera angle, and camera movement. Colored connecting lines trace which package informed which shot. Appendix C Alignment Review Skill The alignment score falignf_align of §3.1 is produced by a review skill that compares a generated storyboard against the director reference. Figure 8 reproduces that skill file, translated and condensed to fit the column. The reviewer receives both artifacts, whereas the generator never sees the reference. Comparison proceeds over the five professional dimensions, and the director storyboard is treated as the sole correct target on every one of them. Two design choices make the resulting signal usable for rule level diagnosis. First, the skill discards all surface variation. Table layout, column order, and wording are excluded from the comparison, so a deviation reflects a directorial decision rather than a formatting artifact. Second, the report records only deviations. Praise and hedging are suppressed, and each deviation must cite the shot numbers at which the two storyboards diverge. A deviation without such a citation cannot be attributed and is therefore rejected. Phase 4 of evolution consumes these localized deviations and joins them with the rule adoption records, as described in §3.3. # Role You are a professional storyboard review expert. Compare each AI storyboard against the directorâs original storyboard and score its degree of alignment. # Review dimensions Follow the five dimensions below. Ignore table format, column order, and wording entirely. 1. Rhythm --- shot splitting density, timing of inserted reaction shots, emotional breathing room. 2. Visual description --- purely visual content, physical action, micro expression and physiological reaction such as a swallow or a tremble. Ignore all interiority and literary embellishment. 3. Shot scale --- scale selection, including the directorâs preference for close-up and extreme close-up on body detail. 4. Camera angle --- subjective and objective viewpoint shifts, over the shoulder framing, Dutch angle, and the power dynamics built by low and high angles. 5. Camera movement --- visual stability of the move, such as a locked off frame or a slow push in. # Scoring Each dimension uses a 100-point scale; the overall score is their mean. A higher score means closer alignment with the director. The directorâs manual storyboard is the only correct alignment target (gold standard). # Workflow 1. Ask the user for the paths of the AI storyboards and of the director benchmark file before starting. 2. Write a Python script to read the CSV files and analyse the differences. 3. Never print a whole table, which overflows the context. Slice the data or search for action keywords to locate the dramatic peaks, then compare reaction close-ups and oppressive compositions there. 4. Emit the report in the two prescribed parts. # Schema pitfalls The two sources carry different headers, so resolve each field through a cascade of fallbacks. The AI output merges scale, viewpoint, and composition into one column, whereas the director file splits scale and angle apart. Read every file as utf-8-sig so that a byte order mark cannot corrupt the first header. Guard against generated files that hold a header but no rows. # Output Part 1. A summary table: one row per plan, with the episode, the overall score, the five dimension scores, and a one-line justification. Part 2. A deviation analysis, under strict discipline: â Pain points only. Never state an advantage of the AI plan, a weakness of the director, or any word of praise. â Cite shot numbers. For example: in Shot 15 the director uses a Dutch low angle with an over-the-shoulder framing to convey Victorâs pressure, whereas plan A stays at eye level and loses the spatial hierarchy entirely. â Report explicitly whether the plan misses the directorâs preferred visual grammar, the handheld breathing quality, or the listenerâs reaction shot. Alignment Review Skill (condensed from the skill file used for falignf_align) Figure 8. The skill file behind falignf_align, translated and condensed to fit the column. Section headings and the wording of every directive follow the original. The director storyboard is the gold standard and stays hidden from the generator.A framed transcript of the alignment review skill file, set in a monospaced font under a colored title bar. Sections appear in order. A role section casts the reader as a storyboard review expert. A review dimensions section enumerates the five professional dimensions and instructs the reviewer to ignore table format, column order, and wording. Further sections state the scoring procedure and the report format. Appendix D Quality Scoring Rubric The quality metric fqualf_qual of §4.1 scores a storyboard on absolute professional standards without any reference. Figure 9 reproduces the scoring prompt, translated and condensed to fit the column. Its grounding is the editing priority of Walter Murch, which ranks emotion and story above rhythm and continuity (23). Every dimension is read in the context of vertical short drama, where narrative density is high and the opening seconds govern retention. Each dimension carries a 100 point scale divided into five bands of twenty points, and the overall score is their unweighted mean. Every band is anchored by the observable evidence expected at that level rather than by an adjective, which keeps runs comparable. The bands share one structure across the five dimensions, so a score reads the same wherever it appears. Three protocol rules govern the output. Every credit or deduction cites specific shot numbers. Judgment follows short drama practice instead of feature film convention. Section 4.6 validates the rubric against three professional directors. # Role You are a senior storyboard director and cinematographer with twenty years of experience, fluent in vertical short drama production. Score the given storyboard on five dimensions, 100 points each, with no reference storyboard available. Ground every judgment in: Murchâs six criteria for editing --- emotion 51%, story 23%, rhythm 10%, eye trace 7%, planarity 5%, spatial continuity 4%; the function of a storyboard as the route map from script to screen, where every frame carries a narrative purpose; and short drama traits --- dense narration, mobile first, the first seconds deciding retention, a twist every 30 to 60 seconds. # Score anchors (five bands of twenty points, the same structure on every dimension; the rhythm dimension is shown) 81â100 complete rhythmic arc of setup, escalation, climax, and breath; density gradient precisely matched to tension; ASL near 1.5 to 2s at the climax and 3 to 5s in dialogue; deliberate breathing shots; the first three shots form an effective hook. 61â80 rhythm varies plausibly and the gradient is broadly right; breathing shots exist but sit imprecisely. 41â60 basic fast and slow variation, yet mechanical; breathing shots absent or misplaced; the opening lacks a hook. 21â40 scattered variation with no density logic; action and dialogue are barely distinguished. 0â20 monotonous throughout; no breathing shot; the climax is indistinguishable from the setup, and shots are merely enumerated. # Core criteria per dimension (abridged) Rhythm â density gradient, breathing design, hook structure, arc within a scene, motivation for each cut. Visual description â executable composition, light direction and quality, spatial layering, precision of character description, environmental storytelling, purely visual content only. Shot scale â spectrum coverage from ELS to ECU, emotional weight matching, progression logic, the 30-degree rule, restraint on ECU. Camera angle â narrative motivation, power dynamics through low and high angles, the 180-degree axis, POV placement, restraint on Dutch angle. Camera movement â trigger, path, and stop point for every move; static against moving contrast; semantics of direction; type variety. Vertical rules override film convention: 9:16 framing, a naturally high share of CU and MCU, and push-in to CU as the signature reveal. # Output format Report a score table over the five dimensions with a grade each and the overall mean. Grades follow the bands: excellent, good, fair, pass, fail. Then, per dimension, give the reasoning with shot numbers as evidence, the strongest case, the main defect, and a concrete suggestion targeting it. Close with a verdict naming the single most important strength and weakness. # Cautions Judge by the standard of a working short drama director rather than relaxing it because the plan already beats most AI output. Every deduction must cite shot numbers. Stay in the short drama context and never import feature film standards. Executability comes first: the bottom line is whether a crew could shoot the plan without further discussion. A few standout shots do not offset a systemic defect, since consistency across the episode matters more than a local highlight. The five dimensions must remain discriminative, so do not assign one score to all of them. Quality Scoring Rubric (condensed from the prompt used for fqualf_qual) Figure 9. The scoring prompt behind fqualf_qual, translated and condensed to fit the column. The five bands are shown for the rhythm dimension; the other four share the same structure and are abridged to their core criteria.A framed transcript of the quality scoring prompt, set in a monospaced font under a colored title bar. A role section casts the reader as a senior storyboard director and states that no reference storyboard is available. A grounding paragraph lists Murch's six criteria for editing with their weights, the function of a storyboard, and the traits of short drama. A score anchors section then gives five bands of twenty points for the rhythm dimension, each band stating the observable evidence expected at that level.