Paper deep dive
Grounding Machine Creativity in Game Design Knowledge Representations: Empirical Probing of LLM-Based Executable Synthesis of Goal Playable Patterns under Structural Constraints
Hugh Xuechen Liu, Kıvanç Tatar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/13/2026, 12:30:43 AM
Summary
This paper investigates the feasibility of using Large Language Models (LLMs) to synthesize executable Unity game code from goal-based design patterns. The authors evaluate two models (DeepSeek-Coder-V2-Lite-Instruct and Qwen2.5-Coder-7B-Instruct) using a direct generation baseline versus an intermediate representation (IR)-conditioned pipeline. The study finds that while IR-based conditioning improves structural grounding, all generated artifacts failed to compile, revealing significant bottlenecks in project-level and engine-level grounding, as well as persistent 'hygiene' failures in code generation.
Entities (5)
Relation Signals (3)
IR v0.2-runtime-evidence → conditions → LLM
confidence 95% · pipelines conditioned on a human-authored Unity-specific intermediate representation (IR)
DeepSeek-Coder-V2-Lite-Instruct → evaluatedon → Goal Playable Concepts
confidence 90% · We compare a direct generation baseline... with pipelines conditioned on a human-authored Unity-specific intermediate representation (IR)
Unity → imposesconstraintson → LLM
confidence 90% · generated artifacts must satisfy Unity's syntactic and architectural requirements
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Creatively translating complex gameplay ideas into executable artifacts (e.g., games as Unity projects and code) remains a central challenge in computational game creativity. Gameplay design patterns provide a structured representation for describing gameplay phenomena, enabling designers to decompose high-level ideas into entities, constraints, and rule-driven dynamics. Among them, goal patterns formalize common player-objective relationships. Goal Playable Concepts (GPCs) operationalize these abstractions as playable Unity engine implementations, supporting experiential exploration and compositional gameplay design. We frame scalable playable pattern realization as a problem of constrained executable creative synthesis: generated artifacts must satisfy Unity's syntactic and architectural requirements while preserving the semantic gameplay meanings encoded in goal patterns. This dual constraint limits scalability. Therefore, we investigate whether contemporary large language models (LLMs) can perform such synthesis under engine-level structural constraints and generate Unity code (as games) structured and conditioned by goal playable patterns. Using 26 goal pattern instantiations, we compare a direct generation baseline (natural language -> C# -> Unity) with pipelines conditioned on a human-authored Unity-specific intermediate representation (IR), across three IR configurations and two open-source models (DeepSeek-Coder-V2-Lite-Instruct and Qwen2.5-Coder-7B-Instruct). Compilation success is evaluated via automated Unity replay. We propose grounding and hygiene failure modes, identifying structural and project-level grounding as primary bottlenecks.
Tags
Links
- Source: https://arxiv.org/abs/2603.07101v1
- Canonical: https://arxiv.org/abs/2603.07101v1
Trouble viewing inline? Open PDF directly →
Full Text
64,187 characters extracted from source content.
Expand or collapse full text
Grounding Machine Creativity in Game Design Knowledge Representations: Empirical Probing of LLM-Based Executable Synthesis of Goal Playable Patterns under Structural Constraints Hugh Xuechen Liu, Kıvanc ̧ Tatar Chalmers University of Technology and University of Gothenburg xuechen@chalmers.se, tatar@chalmers.se Abstract Creatively translating complex gameplay ideas into exe- cutable artifacts (e.g., games as Unity projects and code) re- mains a central challenge in computational game creativity. Gameplay design patterns provide a structured representa- tion for describing gameplay phenomena, enabling designers to decompose high-level ideas into entities, constraints, and rule-driven dynamics. Among them, goal patterns formalize common player–objective relationships. Goal Playable Con- cepts (GPCs) operationalize these abstractions as playable Unity engine implementations, supporting experiential explo- ration and compositional gameplay design. We frame scalable playable pattern realization as a problem of constrained executable creative synthesis: generated arti- facts must satisfy Unity’s syntactic and architectural require- ments while preserving the semantic gameplay meanings en- coded in goal patterns. This dual constraint limits scalabil- ity. Therefore, we investigate whether contemporary large language models (LLMs) can perform such synthesis under engine-level structural constraints and generate Unity code (as games) structured and conditioned by goal playable pat- terns Using 26 goal pattern instantiations, we compare a direct generation baseline (natural language → C# → Unity) with pipelines conditioned on a human-authored Unity-specific intermediate representation (IR), across three IR configu- rations and two open-source models (DeepSeek-Coder-V2- Lite-Instruct and Qwen2.5-Coder-7B-Instruct). Compilation success is evaluated via automated Unity replay. We propose grounding and hygiene failure modes, identifying structural and project-level grounding as primary bottlenecks. Introduction Computational Creativity (C) investigates how computa- tional systems generate novel, coherent, and valuable arti- facts (Boden 2004; Colton, Charnley, and Pease 2011). In computational game creativity—C “within and for” digi- tal games (Liapis, Yannakakis, and Togelius 2014)—an ad- ditional constraint is unavoidable: artifacts must be exe- cutable and playable. We frame this challenge as creative realization (Costikyan 2002): given structured game de- sign knowledge, can a computational system reliably in- stantiate it as operational and playable content? Executable viability is not merely an engineering gate but a neces- sary condition for artifact existence under real engine con- straints. A second motivation is scale: human-authored playable concept implementations (Kultima et al. 2020; Lyu, Holopainen, and Bj ̈ ork 2023) are rich but do not scale; LLM-driven pipelines that reliably instantiate design knowl- edge as executable artifacts open pathways to large-scale ex- ploration of gameplay design spaces. We adopt a representation-first perspective (Shaker, To- gelius, and Nelson 2016; Togelius et al. 2011; Liapis, Smith, and Shaker 2016): rather than treating LLM generation as direct text-to-code synthesis, we ask whether explicit design knowledge encoded as a structured intermediate representa- tion (IR) can scaffold reliable creative realization. Game- play design patterns (Bjork and Holopainen 2005) provide a structured representation for recurring interaction struc- tures; among these, goal patterns describe how player objec- tives are constituted through entities, constraints, and rule- governed interactions. Instantiating a goal pattern in Unity requires coherent object configuration, correct component attachment, and runtime wiring—making goal-pattern real- ization a stringent testbed for creative realization under real engine constraints. Using 26 goal pattern reference implementations (Lyu, Holopainen, and Bj ̈ ork 2023; Kultima et al. 2020), we compare a direct generation baseline against pipelines con- ditioned on IR v0.2-runtime-evidence across three config- urations (free, min, full) and two open-source mod- els: DeepSeek-Coder-V2-Lite-Instruct (Dai et al. 2024) and Qwen2.5-Coder-7B-Instruct (Hui et al. 2024; Yang et al. 2024). Every generated artifact failed to compile; rather than treating this as a negative result, we analyze the struc- tured distribution of failures as empirical evidence of where grounding breaks down. Contributions. We contribute: (1) Creative realization as a C problem framing—situating executable pattern instantiation within computational creativity and establish- ing compile-grounded viability as a necessary condition for artifact existence; (2) An execution-grounded evaluation pipeline—an end-to-end workflow from HPC inference to Unity batch replay and per-seed metric aggregation; (3) A Unity-specific IR for knowledge injection—IR v0.2- runtime-evidence encoding project-level structural conven- tions and goal-pattern semantic knowledge, evaluated un- der three conditioning configurations (free, min, full); (4) A structured failure taxonomy—empirical analysis of grounding and hygiene failure modes, revealing project- arXiv:2603.07101v1 [cs.AI] 7 Mar 2026 level and engine-level grounding as primary bottlenecks for knowledge-conditioned LLM generation in Unity. Background Computational Creativity and Constrained Generative Synthesis Computational creativity research concerns systems that generate novel, valuable, and surprising artifacts (Boden 2004; Ventura 2016; Colton, Wiggins, and others 2012). Bo- den’s framework distinguishes combinatorial, exploratory, and transformational creativity (Boden 2004); exploratory creativity, most relevant here, involves traversing a struc- tured conceptual space to produce artifacts novel within its boundaries. In generative systems, constraints serve a dual role: bounding the conceptual space where creativity can take place. (?). This is particularly acute in executable cre- ative synthesis, where artifacts must satisfy both semantic design intentions and syntactic execution requirements— unlike open-ended generation, executable artifacts admit ob- jective validity criteria: they either run or they do not. Co-creative systems distribute creative agency across hu- man and machine contributors (Liapis, Smith, and Shaker 2016; Deterding et al. 2017), with each contributing what the other cannot efficiently provide. Prior work demon- strates that human-authored constraints can substantially improve coherence of machine-generated content (Liapis, Smith, and Shaker 2016). Our work instantiates this model: human authors contribute domain knowledge as a struc- tured IR while the model contributes generative breadth; where this boundary breaks down is a central empirical question. Prior computational game creativity work has ap- proached generation through procedural content generation (PCG) (Shaker, Togelius, and Nelson 2016) or LLM-based content generation (Todd et al. 2023; Chen et al. 2023; Sudhakaran et al. 2023); our work differs in targeting instan- tiation of a design concept as a complete executable artifact, foregrounding grounding as central to creative synthesis. LLMs for Code Generation and Executable Artifact Synthesis Large language models have demonstrated substantial ca- pability from function-level synthesis (Chen et al. 2021) to repository-level editing (Jimenez et al.2023; Jiang et al. 2026), extended to structured generation conform- ing to domain-specific schemas (Poesia et al.2022) and creative domains balancing novelty with formal con- straints (Yuan et al. 2022). However, LLM-based code generation exhibits characteristic failure modes for complex, architecture-dependent artifacts: models generate locally plausible code that fails at integration—syntactically cor- rect components that are architecturally incompatible within the target project (Jimenez et al. 2023; Chen et al. 2023). In game engine contexts this is compounded by engine- specific conventions underrepresented in pretraining cor- pora: Unity’s component model, scene graph architecture, and MonoBehaviour lifecycle impose constraints that differ substantially from general-purpose programming patterns. Intermediate representations (IR) address this complex- ity by decomposing generation into semantically mean- ingful intermediate steps with more tractable targets (Yin and Neubig 2017; Austin et al.2021), particularly ef- fective when the IR externalizes stable domain knowledge and reduces the burden on the model to rediscover struc- tural conventions from context alone (Dohan et al. 2022; Gao et al. 2023). Goal Playable Concepts as Coupled Game Design Representation Gameplay design patterns describe recurring interaction structures for analysis and reuse in game design (Bjork and Holopainen 2005).As intermediate-level design knowl- edge (H ̈ o ̈ ok and L ̈ owgren 2012), they occupy a position be- tween abstract theories and concrete instances, encoding in- teraction relations rather than surface aesthetics. Because gameplay design constitutes second-order design (Tekinbas and Zimmerman 2003)—designers control rules rather than experience directly—pattern descriptions are inherently ab- stract and resistant to direct operationalization. Patterns con- stitute a design language (Rheinfrank and Evenson 1996) with densely interconnected instantiation and modulation re- lations (Bjork and Holopainen 2005), meaning no single pat- tern can be instantiated without invoking others and impos- ing non-trivial compositional requirements on any genera- tive system. Amongthese,goalpatternsformalizeplayer– objectiverelationshipsfocusingonimperative interaction-level goals (Bjork and Holopainen 2005; Debus, Zagal, and Cardona-Rivera 2020)—a bounded, structured design space amenable to both exploratory and combinatorial creative operations in Boden’s sense (Boden 2004). Goal Playable Concepts (GPCs) (Lyu, Holopainen, and Bj ̈ ork 2023) 1 operationalize goal patterns as playable Unity implementations, coupling textual descriptions with interactive instantiations. Neither format is sufficient alone: patterns lack interactivity while playable concepts cannot encode abstract relational structure; their coupling is struc- turally necessary as each serves as contextual scaffolding for interpreting the other (Rheinfrank and Evenson 1996; Kultima et al.2020; Consalvo 2017).From a data perspective,GPCs expose design knowledge across three modalities: natural language descriptions, graph- structured relational knowledge, and executable Unity code—providing a grounded source from which to derive an IR encoding both structural conventions and semantic gameplay intentions. Grounding LLM Generation in Structured Domain Knowledge Grounding constrains model outputs to conform to external knowledge structures, reducing the gap between fluent gen- eration and valid artifacts (Ji et al. 2023). Staged pipelines with IRs improve conformance by externalizing stable do- main knowledge (Dohan et al. 2022; Gao et al. 2023; 1 https://gameplaydesignpatterns.itch.io/ Chen et al. 2021; Jiang et al. 2026). In game engine con- texts, this is compounded by engine-specific conventions underrepresented in pretraining corpora—Unity’s compo- nent model, MonoBehaviour lifecycle, and scene graph ar- chitecture impose constraints that differ substantially from general-purpose programming patterns. Problem Setting and Method Problem Setting We consider 26 Unity scene implementations, each a ref- erence instantiation of a distinct goal pattern from the Goal Playable Concepts collection, co-existing in an au- thored Unity project containing prefab assets, MonoBe- haviour scripts, and scene-level configuration. Generation targets a new Unity Editor script instantiating the scene for the target pattern, performed on top of this existing layer rather than from scratch. Evaluation measures compilation viability under real Unity engine constraints; semantic fi- delity evaluation remains future work. Pipeline Overview We compare two pipelines across four configurations. No- schema baseline: the model receives the goal pattern as a natural language markdown document and generates a Unity Editor C# script directly. IR-conditioned pipeline: gener- ation proceeds in two steps—first generating an IR JSON from the pattern description, then translating it to C#— under three conditioning levels: free (no schema con- straints), min (minimal field skeleton), and full (complete IR v0.2-runtime-evidence schema with four hard referential- integrity constraints). Prompt templates are provided verba- tim in Appendix: Prompt Templates (Verbatim). IR v0.2-runtime-evidence The IR mediates between goal pattern description and ex- ecutable Unity realization, derived from the same Unity project used as the evaluation target to ground it simultane- ously in concrete implementation structure and goal pattern semantics. Schema bootstrapping and freeze. An initial static draft defined six top-level fields from three hand-authored pat- terns (Ownership, Delivery, Alignment). Static YAML pars- ing proved insufficient—critical mechanics emerge from prefab instantiation and runtime script logic—so the schema was extended with runtime params and refined through five versioned iterations (Appendix: IR iteration record). Frequency analysis across all 26 patterns confirmed all seven top-level fields are Core-tier (100% coverage); the schema was frozen as v0.2-runtime-evidence prior to any generation runs. Extraction pipeline. The IR for each pattern is generated programmatically from static analysis of the Unity scene YAML: PrefabInstance blocks are resolved to canonical as- set names via a curated GUID map; MonoBehaviour blocks to C# class names via a script GUID map; serialized field values are extracted into runtime params. Semantic links and rules entries—encoding conditional runtime relations evidenced in script source—are assembled by a pattern-specific configuration function. Three scenes were hand-authored as bootstrapping examples; the remaining 23 were batch-processed. Schema definition and dual grounding. The frozen schema defines seven fields: scene, objects, scripts, params (always ), runtime params, links, and rules, governed by four hard constraints (per-instance script binding,no aggregate placeholders,required evidence type on all rules, conditional relation label convention). The IR carries two complementary grounding layers: a structural layer grounded in the concrete project implementation and a semantic layer grounded in the goal pattern, providing both the implementation informa- tion needed for compilable Unity code and the pattern information needed for correct gameplay instantiation. Full schema and annotated example are in Appendix: IR v0.2-runtime-evidence Summary. Models and Configurations Two code-specialized open-source models were selected: DeepSeek-Coder-V2-Lite-Instruct (Dai et al.2024) and Qwen2.5-Coder-7B-Instruct (Hui et al. 2024; Yang et al. 2024) chosen for code specialization, open weights, and HPC batch inference feasibility via vLLM. In addition, DeepSeek-Coder-V2-Lite-Instruct and Qwen2.5-Coder-7B- Instruct report the highest pass@1 percentage on Hu- manEval among recent open-source models (Jiang et al. 2026; Chen et al.2021).The evaluation comprises 2 models× 4 configurations× 520 records = 4,160 total records (520 = 26 patterns × 20 seeds), run on NVIDIA A40 GPUs via vLLM (Kwon 2025) v0.7.3 under Apptainer 2 .Decoding is done with the following settings: tem- perature = 0.2, top p = 0.95, maxmodellen = 3072; maxtokens = 2048 (C#) and 4096 (IR). Evaluation Protocol Primary metric: M1 Compile Success. M1 is binary compile viability from Unity batch replay logs: each gen- erated C# script is written to a temporary asset file, Unity 2022.2.23f1 is invoked in batchmode 3 , and the Editor log is scanned for compiler error codes. Records exceeding 120 seconds are terminated and marked compile timeout. Compilation is a minimal operational threshold: without it, an artifact cannot exist as an executable candidate. Failure analysis and reproducibility. Compiler error codes are extracted from Unity batch logs and analyzed to characterize failure modes; timeout records are reported sep- arately. Pattern data, playable implementations, pipeline code, and reproduction instructions are available in the sup- plementary repository. 4 2 https://github.com/apptainer/apptainer 3 https://docs.unity3d.com/2019.1/ Documentation/Manual/CommandLineArguments. html 4 https://anonymous.4open.science/r/ llm-goal-playable-pattern-E312/README.md Results M1 Compile Success No generated artifact achieved successful compila- tion across either model or any pipeline configuration: pass@k = 0.0 for all k ∈1, . . . , 20 across all 26 patterns, two models, and four configurations. A notable secondary finding is the sharp increase in compilation timeout rate under IR-conditioned configura- tions. Under no schema, timeout rates range from 37.5% (DeepSeek) to 51.5% (Qwen); under IR-conditioned config- urations they rise monotonically with schema detail, reach- ing 96–99% under with schemafull. Compilation timeout is recorded when Unity’s BatchRunner watchdog terminates a compilation-and-domain-reload cycle after 120 seconds. The monotonic increase may suggest that IR-conditioned generation produces structurally more complex C# outputs that systematically exhaust the compilation budget. Error code data for IR-conditioned configurations should therefore can be interpreted as partial compilation evidence. Failure Distribution Observed Compiler Errors Table 1 lists all 41 C# com- piler error codes observed across the evaluation, with their compiler-reported message templates and observed frequen- cies by configuration. Inspection of the error messages reveals two distinct fail- ure types. A first group of 13 errors—CS0115, CS0117, CS0122, CS0234, CS0239, CS0246, CS0311, CS0315, CS0509, CS0619, CS1061, CS1624, CS8121—all involve references to types, members, namespaces, or inheritance structures that do not exist in the target Unity project or engine API. These errors indicate that the model gener- ated code assuming the existence of constructs that are ab- sent from the actual codebase. The required knowledge is dual in nature: accurate knowledge of the target project’s Unity implementation structure (prefab identifiers, script class names, component bindings) and accurate knowledge of the goal pattern vocabulary (which mechanics and rela- tions the pattern requires). Both layers are encoded in the IR; their absence in the no schema condition is precisely what these errors reflect. We term these grounding failures. A second group of 28 errors—CS0029, CS0101, CS0103, CS0111, CS0116, CS0136, CS0165, CS0263, CS0595, CS1001, CS1002, CS1003, CS1010, CS1012, CS1013, CS1022, CS1026, CS1029, CS1040, CS1041, CS1056, CS1503, CS1513, CS1519, CS1525, CS1529, CS2001, CS8803—reflects syntax corruption, duplicate declarations, formatting leakage, and type coercion errors. These errors are independent of the model’s knowledge of the project or pattern information: they would occur even if all referenced types and constructs existed. We term these hygiene failures — errors independent of grounding that are in principle ad- dressable through constrained decoding or output sanitiza- tion — borrowing the notion of hygiene from programming language theory (Kohlbecker et al. 1986) and software en- gineering practice (Tilbrook and McMullen 1990). Grounding Failures Grounding failures (as short as G failures) reflect the model’s inability to map goal-pattern semantics onto the actual implementation primitives avail- able in the target Unity project. Under noschema, this category accounts for 121 out of 329 total logged error in- stances (36.8%). The dominant codes are CS0115 (40 logs), CS0122 (17 logs), CS0234 (20 logs), and CS0246 (33 logs), collectively indicating failures at three grounding layers: Failures occur at three layers: project-level (CS0122, CS0246; 17 and 33 logs under no schema), engine API (CS0117, CS0234, CS0311, CS0315, CS0619, CS1061, CS8121; present across all configurations), and architec- tural (CS0115, CS0239, CS0509, CS1624; 45 logs under no schema, nearly absent under IR-conditioned generation). The IR-conditioned configurations show a marked re- duction in grounding failures relative to no schema. Un- der withschemafree, the grounding-sensitive total falls to 18 (8.3% of logged errors), driven primarily by the near- elimination of CS0115, CS0234, and CS0122.Under with schemamin and withschemafull, however, ground- ing failure totals rise to 53 (21.8%) and 45 (20.7%) respec- tively, driven primarily by the persistence of CS0246 (29 and 31 logs respectively) and an increase in CS1061 (10 and 12 logs, compared to 3 under no schema). CS0311 and CS0315 also appear under withschemamin (5 logs each) but are ab- sent or near-absent in other configurations. These interpre- tations must be qualified by the high timeout rates under all IR-conditioned configurations: the majority of records did not reach a complete compilation cycle, and the error code data reflects partial logs only. Hygiene Failures Hygiene failures (as short as H fail- ures) reflect pipeline hygiene and output formatting is- sues rather than grounding deficits.Under no schema, this category accounts for 208 out of 329 total logged er- ror instances (63.2%), dominated by duplicate declaration errors (CS0101, CS0111, CS0263), type coercion errors (CS1503, CS1513), and unassigned-local errors (CS0165). The composition shifts markedly under IR-conditioned gen- eration. Under withschemafree, withschemamin, and withschemafull, hygiene failures account for 198 (91.7%), 190 (78.2%), and 172 (79.3%) of logged errors respec- tively. Codes associated with output formatting and sanitizer rejection—CS1029 (marker comment leakage), CS1001, CS1003, CS1013, and CS1529—become dominant, replac- ing the duplicate declaration and unassigned-local errors that characterise no schema output. This compositional shift in- dicates progress: surface formatting failures are address- able through post-processing, whereas duplicate declara- tions and unassigned-local errors reflect deeper generation- side issues. CS1029 saturates at 40 log files across all three IR-conditioned configurations, suggesting systematic failure to strip IR-related markup. CS0263 and CS0165, promi- nent under noschema (20 and 13 logs respectively), are en- tirely absent under IR-conditioned generation. These inter- pretations are subject to a caveat: high timeout rates under IR-conditioned configurations mean that compilation errors manifesting late in the build cycle are systematically unob- served, and the true error distribution may differ from partial logs. Table 1: All observed C# compiler error codes, message templates, and logcount by configuration. Sorted by error code number. NS = no schema CodeCompiler messageNSFreeMinFullTotal CS0029Cannot implicitly convert X (e.g., string) to Y (e.g., string[])21003 CS0101Namespace already contains a definition for X (e.g., PlayerController)2518152280 CS0103The name X (e.g., whatToConceal) does not exist in the current context1416141559 CS0111Member X (e.g., Update) already defined with same parameter types161491857 CS0115 X (e.g., EvadeGoal.IsCompleted()) is not a suitable method for override4000040 CS0116Namespace cannot directly contain members such as fields or methods00011 CS0117 X (e.g., EditorGUILayout) does not contain a definition for Y (e.g., Vector2ArrayField) 22015 CS0122 X (e.g., RescueGoal.Rescue()) is inaccessible due to its protection level1700017 CS0136Local X cannot be declared; name used in enclosing scope20002 CS0165Use of unassigned local variable X (e.g., interferableGoal)1300013 CS0234Type X (e.g., Rigging) does not exist in namespace Y (e.g., UnityEngine.Animations) 2000020 CS0239Cannotoverridesealedmember X(e.g., TrackAsset.CreatePlayable()) 40004 CS0246The type or namespace X (e.g., GuardGoal) could not be found33112931104 CS0263Partial declarations of X (e.g., GainInformation) must not specify different base classes 2000020 CS0311Type X (e.g., ActionDescription) cannot be used as type parameter T in Y (e.g., AddComponent<T>()) 10517 CS0315Type X (e.g., UnityEngine.Color) cannot be used as type parameter T; no boxing conversion 00505 CS0509Cannot derive from sealed type X (e.g., ClipCaps)10001 CS0595Invalid real literal0103114 CS0619 X (e.g., AddComponent(string)) is obsolete00101 CS1001Identifier expected124241766 CS1002 ; expected903214 CS1003Syntax error, X (e.g., ,) expected420221258 CS1010Newline in constant60006 CS1012Too many characters in character literal10001 CS1013Invalid number020171148 CS1022Type or namespace definition, or end-of-file expected603211 CS1026 ) expected20002 CS1029 #error directive (e.g., BatchRunner sanitizenocsharpstart)19404040139 CS1040Preprocessor directive must be first non-whitespace on line00314 CS1041Identifier expected; X (e.g., in) is a keyword01102 CS1056Unexpected character20002 CS1061 X (e.g., PlayerController) does not contain a definition for Y (e.g., health) 34101229 CS1503Argument N: cannot convert from X (e.g., string) to Y (e.g., UnityEngine.Object) 2312026 CS1513 expected2403229 CS1519Invalid token in member declaration10001 CS1525Invalid expression term10001 CS1529 using clause must precede all other elements424261872 CS1624Body of X (e.g., MoveGameElement()) cannot be an iterator block because return type is Y (e.g., void) 01001 CS2001Source file X could not be found694827 CS8121Expression of type X (e.g., Component) cannot be handled by pattern of type Y (e.g., ScriptableObject) 00303 CS8803Top-level statements must precede namespace declarations701210 Total3292162432171005 Cross-Model Comparison As shown in Table 2, un- der no schema, where timeout rates are sufficiently low to support reliable comparison (Qwen: 51.5%, DeepSeek: 37.5%), both models produce grounding failures at compa- rable absolute levels (Qwen: 64, DeepSeek: 57). However, their grounding failure profiles are compositionally distinct: Qwen produces CS0234 (20 vs. 0) and CS0239 (4 vs. 0), while DeepSeek produces CS0122 (17 vs. 0) and CS0246 at a higher rate (20 vs. 13). Both models converge on CS0115 (20 each), supporting the interpretation that architectural grounding failure is structural rather than model-specific. Hygiene failures diverge more sharply: DeepSeek shows higher rates of duplicate declaration and structural errors— CS0101 (20 vs. 5), CS0263 (20 vs. 0), CS0165 (13 vs. 0), and CS1513 (20 vs. 4)—while Qwen produces CS1029 (19 vs. 0) and CS1503 (19 vs. 4) at substantially higher rates. Under with schemafree (timeout 86–92%),Qwen is dominated by CS1001/CS1003/CS1013/CS1529 while DeepSeek shows higher CS0101/CS0111; CS1029 satu- rates at 20 log files for both. Under with schemamin and withschemafull (timeout≥96%), quantitative comparison is unreliable; the most consistent pattern is CS0101/CS0111 persistently higher for DeepSeek and CS0103/CS1013 for Qwen, with CS1029 saturating at 20 per model across all IR-conditioned configurations. Pattern-Level Failure Distribution Tables 3 aggregate both models (40 logs per pattern per configuration (20 seeds × 2 models)). Aggregation is appropriate at the configura- tion level given the cross-model consistency in grounding failure categories. However, pattern-level grounding failure counts in some cases are model-specific rather than shared behaviour; per-model breakdown is provided in Appendix: Pattern-Level Error Distribution by Model. Across all configurations, G failures are distributed un- evenly across patterns. Under no schema, 22 of 26 pat- terns exhibit at least one G failure; 4 patterns (1Ownership, 3Eliminate, 19Exploration, 25LastManStanding) show zero G failures with all logged failures attributable to H causes. Under withschemafree, G failures are reduced to 7 patterns, with CS1061, CS0117, CS0246 and CS1624 as the only remaining G codes. Under with schemamin and withschemafull, G failures reappear across more patterns, with CS0246 dominant in nearly all affected cases; however, these distributions are based on partial logs due to high time- out rates and should be interpreted with caution. A consistent cross-configuration observation is that 1Ownership exhibits persistent G failures across all three setups (G log counts: 9, 14, 16 under with schemafree, min, and fullrespectively), despite showing zero G failures under noschema, in contrast to structurally similar patterns such as 2Collection and 3Eliminate which show G failures only under noschema or not at all. The interpretation of this cross-configuration pattern is deferred to Section Dis- cussion. Table 2:Per-model logcount by error code and configuration.G = groundingfailure;H = hygiene failure.NS = no schema,Free = withschemafree, Min = withschemamin, Full = withschemafull. QwenDeepSeek ClassCodeNSFreeMinFullNSFreeMinFull G CS01152000020000 CS011722010000 CS0122000017000 CS0234200000000 CS023940000000 CS024613315192081412 CS031110410010 CS031500500000 CS050910000000 CS061900100000 CS1061301100492 CS162401000000 CS812100300000 H CS002921000000 CS0101511420171418 CS010371413157210 CS011160121014816 CS011600000001 CS013600002000 CS0165000013000 CS0263000020000 CS0595010300001 CS100112014904108 CS100220007032 CS10032201492083 CS101010005000 CS101210000000 CS10130201490032 CS102210005032 CS102600002000 CS1029192020200202020 CS104000000031 CS104101000010 CS105620000000 CS1503190204100 CS1513400020032 CS151910000000 CS152510000000 CS152942014904129 CS200124344514 CS880320005012 Table 3: Pattern-level errors by configuration (both models combined, 40 logs per pattern). G = grounding failure; H = hygiene failure. Timeout rates: NS 37–51%, Free 86–92%, Min 96%, Full 97–99%. noschemawithschemafreewithschemaminwithschemafull PatternTotGHTop GTotGHTop GTotGHTop GTotGHTop G 1Ownership18018—31922CS0246(6), CS1061(3) 711457CS0246(11), CS1061(3) 491633CS0246(15), CS0117(1) 2 Collection23815CS0115(2), CS0234(2) 47047—50050—47839CS0246(6), CS1061(2) 3Eliminate37037—46244CS0117(2)43736CS0246(4), CS8121(2) 43043— 4 Capture412714CS0234(16), CS0115(11) 47047—38038—35035— 5Overcome514CS0234(1)25025—441529CS0246(7), CS1061(5) 541044CS0246(5), CS1061(5) 6 Evade503218CS0115(16), CS0246(15) 46046—38434CS0246(2), CS1061(2) 37037— 7 Stealth52502CS0115(24), CS0246(19) 41041—36036—44638CS0246(3), CS1061(3) 8HerdAttract24618CS0115(3), CS0234(2) 43142CS0246(1)36135CS0246(1)36036— 9Conceal53746CS0115(3), CS0239(2) 47047—40040—44044— 10Rescue31229CS0122(17), CS0246(5) 45045—49049—46046— 11Delivery542727CS0234(16), CS0115(11) 63063—44440CS0246(4)41041— 12Guard17161CS0234(8), CS0115(7) 41041—45045—36036— 13 Race963CS0234(4), CS0115(1) 26224CS0246(2)36432CS0246(3), CS1061(1) 27423CS0246(4) 14 Alignment21516CS0234(3), CS0115(2) 37037—47047—551045CS0246(5), CS1061(4) 15 Configuration514CS0246(1)45045—45936CS0246(5), CS0315(4) 56056— 16Traverse332112CS0234(19), CS0115(2) 32131CS1061(1)36036—41041— 17Survive972CS0115(4), CS0234(2) 41140CS1624(1)431132CS0246(8), CS0311(2) 38038— 18Connection20614CS0115(2), CS0234(2) 37037—48147CS0246(1)39039— 19Exploration404—61061—38335CS0246(1), CS0311(1) 40337CS0246(3) 20Reconnaissance1165CS0246(6)44242CS0246(2)37136CS1061(1)35332CS0246(2), CS1061(1) 21Contact30282CS0246(11), CS0234(9) 53053—37037—35530CS0246(5) 22Enclosure752CS0246(2), CS0115(1) 53053—46244CS0246(2)46145CS0246(1) 23 GainCompetence523CS0246(1), CS0311(1) 44044—38038—42141CS0246(1) 24 GainInformation551045CS0239(4), CS0115(3) 62062—43439CS0246(4)41041— 25 LastManStanding15015—44044—37235CS0246(2)35035— 26 KingoftheHill37433CS0246(2), CS0115(1) 35035—38038—33033— Discussion IR as representational scaffold. IR does not replace gen- erative exploration; it constrains and channels it. The data reveal an asymmetric grounding effect: IR conditioning nearly eliminates architectural grounding failures (CS0115: 40 logs under no schema, 0 under all IR-conditioned config- urations), confirming that the IR’s structural layer success- fully transfers MonoBehaviour and inheritance conventions. However, CS0246 (hallucinated project-specific types) per- sists across all configurations (no schema: 33, free: 11, min: 29, full: 31), suggesting that project-level ground- ing is not fully resolved by schema conditioning alone and may require richer or more targeted knowledge injection. The full configuration’s explicit hard constraints did not eliminate grounding failures, and the monotonic increase in compilation timeout with schema detail (37–51% under no schema to 97–99% under withschemafull) suggests a further problem: IR-conditioned generation produces struc- turally more complex C# outputs that systematically exhaust the Unity compilation budget. The IR is simultaneously nec- essary for grounding and costly for compilation tractability. This tension—too sparse for reliable grounding without it, too complex for reliable compilation with it—defines the current boundary of the approach. Human-machine division of labor. Human-machine co- creativity research distinguishes between systems where humans and machines occupy complementary generative roles, each contributing what the other cannot efficiently provide (Liapis, Smith, and Shaker 2016; Deterding et al. 2017). Our pipeline instantiates this division explic- itly: human authors contribute domain knowledge, concep- tual structure, and representational schema grounded in ex- pert game design practice; the model contributes generative breadth across the space of possible realizations. Neither role is substitutable by the other—the IR cannot be auto- matically derived without human design knowledge, and the scale of instantiation cannot be achieved through manual au- thoring alone. The current results reveal an asymmetry in this division: the human-authored schema successfully encodes what ex- ists in the project, but does not yet fully specify how Unity’s architectural conventions govern the use of those el- ements. Pattern-level analysis further reveals that ground- ing difficulty is not uniformly distributed: under no schema, grounding failure counts range from 0 to 50 across patterns, suggesting that some goal patterns impose systematically higher grounding demands than others. This suggests that the productive boundary between human and machine con- tribution may need to be located at the pattern level rather than uniformly across the task space—some patterns may require richer or more targeted knowledge transfer than oth- ers before reliable machine-side realization becomes possi- ble. Co-creative system design, in this framing, is not only a question of task allocation but of knowledge boundary ne- gotiation grounded in the specific demands of individual de- sign patterns. The immediate next step is not to improve the IR schema but to relocate the human/machine boundary: either through constrained decoding that enforces syntactic hygiene on the machine side, or through scene-level generation that sidesteps the full compilation problem altogether. In either case, the failure taxonomy established here provides the di- agnostic foundation for that relocation—identifying not just that creative realization fails, but precisely where and why. Future directions. The present work establishes compile- grounded viability as a necessary precondition for deeper creative evaluation; future work will extend toward struc- tural adherence, gameplay meaning preservation, and near- pattern confusion analysis (e.g., distinguishing Stealth from Rescue (Lyu, Holopainen, and Bj ̈ ork 2023)), with human intervention requirements as a further dimension of compu- tational creativity assessment. The failure taxonomy iden- tifies two orthogonal intervention targets addressable inde- pendently: grounding failures via a GNN embedding over the gameplay design pattern graph (Peng et al. 2023), in- jectable via PEFT (Liu et al. 2023; Gururangan et al. 2020) or RAG (Lewis et al. 2020); hygiene failures via constrained decoding (Ma and Hu 2025; Zarrieß, Voigt, and Sch ̈ uz 2021; Geng et al. 2023) or rule-based sanitization. Separating the two inference steps—NL→IR for semantic interpreta- tion and IR→C# for syntactic realization—may better match model capability to task demand. The monotonic time- out increase further suggests flattening the prefab layer into explicit enumerable structures, reframing goal-pattern real- ization as a single scene generation problem and shifting the generative challenge from syntactic code correctness to structured compositional assembly. Limitations and scope. Pattern-level model asymmetry (Section Pattern-Level Failure Distribution) is reported as an observation; its interpretation requires per-pattern semantic analysis beyond the current scope. Broader claims regard- ing comparative model performance, configuration optimal- ity, or semantic fidelity remain future work. The evaluation is scoped to a single engine (Unity) and a single pattern cat- egory (goal patterns); generalization to other engines or pat- tern types is not claimed. Conclusion We frame goal-pattern instantiation as a creative realiza- tion problem: converting structured game design knowl- edge into executable digital artifacts under real engine con- straints. Using 26 Unity reference instantiations, we estab- lish an execution-grounded evaluation pipeline and analyze where creative realization succeeds or fails across ground- ing and hygiene failure modes. IR-conditioned generation provides a principled representational interface for knowl- edge injection, encoding both project-level structural con- ventions and goal-pattern semantic knowledge derived from human-authored GPC implementations. Uniform compile failure across all configurations reveals that project-level and engine-level grounding remain primary bottlenecks for knowledge-conditioned LLM generation in Unity, estab- lishing the analytical foundation for deeper structural and semantic evaluation of knowledge-conditioned generative game systems. Acknowledgement The batch compilation pipeline adapts the write-to-asset ap- proach introduced in AICommand by Keijiro Takahashi 5 . We thank Staffan Bj ̈ ork and Jussi Holopainen for their in- put on goal playable concepts and related background. 5 https://github.com/keijiro/AICommand References [Austin et al. 2021] Austin, J.; Odena, A.; Nye, M.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C.; Terry, M.; Le, Q.; et al. 2021. Program synthesis with large lan- guage models. arXiv preprint arXiv:2108.07732. [Bjork and Holopainen 2005] Bjork, S., and Holopainen, J. 2005. Patterns in game design, volume 11. Charles River Media Hingham. [Boden 2004] Boden, M. A. 2004. The creative mind: Myths and mechanisms. Routledge. [Chen et al. 2021] Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. D. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al.2021.Evaluat- ing large language models trained on code. arXiv preprint arXiv:2107.03374. [Chen et al. 2023] Chen, D.; Wang, H.; Huo, Y.; Li, Y.; and Zhang, H.2023.Gamegpt: Multi-agent collabo- rative framework for game development. arXiv preprint arXiv:2310.08067. [Colton, Charnley, and Pease 2011] Colton, S.; Charnley, J. W.; and Pease, A. 2011. Computational creativity the- ory: The face and idea descriptive models. In ICCC, 90–95. Mexico City. [Colton, Wiggins, and others 2012] Colton,S.;Wiggins, G. A.; et al. 2012. Computational creativity: The final fron- tier? In Ecai, volume 12, 21–26. Montpelier. [Consalvo 2017] Consalvo, M. 2017. When paratexts be- come texts: De-centering the game-as-text. Critical Studies in Media Communication 34(2):177–183. [Costikyan 2002] Costikyan, G. 2002. I have no words & i must design: Toward a critical vocabulary for games. In Computer Games and Digital Cultures Conference Proceed- ings. [Dai et al. 2024] Dai, D.; Deng, C.; Zhao, C.; Xu, R.; Gao, H.; Chen, D.; Li, J.; Zeng, W.; Yu, X.; Wu, Y.; et al. 2024. Deepseekmoe: Towards ultimate expert specializa- tion in mixture-of-experts language models. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), 1280–1297. [Debus, Zagal, and Cardona-Rivera 2020] Debus, M. S.; Za- gal, J. P.; and Cardona-Rivera, R. E. 2020. A typology of imperative game goals. Game Studies 20(3). [Deterding et al. 2017] Deterding, S.; Hook, J.; Fiebrink, R.; Gillies, M.; Gow, J.; Akten, M.; Smith, G.; Liapis, A.; and Compton, K. 2017. Mixed-initiative creative interfaces. In Proceedings of the 2017 CHI conference extended abstracts on human factors in computing systems, 628–635. [Dohan et al. 2022] Dohan, D.; Xu, W.; Lewkowycz, A.; Austin, J.; Bieber, D.; Lopes, R. G.; Wu, Y.; Michalewski, H.; Saurous, R. A.; Sohl-Dickstein, J.; et al. 2022. Language model cascades. arXiv preprint arXiv:2207.10342. [Gao et al. 2023] Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; and Neubig, G. 2023. Pal: Program-aided language models. In International confer- ence on machine learning, 10764–10799. PMLR. [Geng et al. 2023] Geng, S.; Josifoski, M.; Peyrard, M.; and West, R. 2023. Grammar-constrained decoding for struc- tured nlp tasks without finetuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, 10932–10952. [Gururangan et al. 2020] Gururangan, S.; Marasovi ́ c, A.; Swayamdipta, S.; Lo, K.; Beltagy, I.; Downey, D.; and Smith, N. A. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th an- nual meeting of the association for computational linguis- tics, 8342–8360. [H ̈ o ̈ ok and L ̈ owgren 2012] H ̈ o ̈ ok, K., and L ̈ owgren, J. 2012. Strong concepts: Intermediate-level knowledge in interac- tion design research.ACM Transactions on Computer- Human Interaction (TOCHI) 19(3):1–18. [Hui et al. 2024] Hui, B.; Yang, J.; Cui, Z.; Yang, J.; Liu, D.; Zhang, L.; Liu, T.; Zhang, J.; Yu, B.; Dang, K.; et al. 2024. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186. [Ji et al. 2023] Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023. Survey of hallucination in natural language genera- tion. ACM computing surveys 55(12):1–38. [Jiang et al. 2026] Jiang, J.; Wang, F.; Shen, J.; Kim, S.; and Kim, S. 2026. A survey on large language models for code generation. ACM Transactions on Software Engineering and Methodology 35(2):1–72. [Jimenez et al. 2023] Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K. 2023. Swe- bench: Can language models resolve real-world github is- sues? arXiv preprint arXiv:2310.06770. [Kohlbecker et al. 1986] Kohlbecker, E.; Friedman, D. P.; Felleisen, M.; and Duba, B. 1986. Hygienic macro expan- sion. In Proceedings of the 1986 ACM Conference on LISP and Functional Programming, 151–161. [Kultima et al. 2020] Kultima, A.; Park, S.; Lassheikki, C.; and Kauppinen, T. 2020. Designing games as playable concepts: five design values for tiny embedded educational games. In Digital Games Research Association Conference. Digital Games Research Association (DIGRA). [Kwon 2025] Kwon, W. 2025. vLLM: An Efficient Inference Engine for Large Language Models. Ph.D. Dissertation, UC Berkeley. [Lewis et al. 2020] Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; K ̈ uttler, H.; Lewis, M.; Yih, W.-t.; Rockt ̈ aschel, T.; et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33:9459–9474. [Liapis, Smith, and Shaker 2016] Liapis, A.; Smith, G.; and Shaker, N. 2016. Mixed-initiative content creation. In Pro- cedural content generation in games. Springer. 195–214. [Liapis, Yannakakis, and Togelius 2014] Liapis, A.;Yan- nakakis, G. N.; and Togelius, J. 2014. Computational game creativity. ICCC. [Liu et al. 2023] Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; and Neubig, G. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys 55(9):1–35. [Lyu, Holopainen, and Bj ̈ ork 2023] Lyu, Z.; Holopainen, J.; and Bj ̈ ork, S. 2023. Goal playable concepts coupling game- play design patterns with playable concepts. In Proceed- ings of the 26th International Academic Mindtrek Confer- ence, 57–66. [Ma and Hu 2025] Ma, F., and Hu, A. J. 2025. Logically constrained decoding. In Valentino, M.; Ferreira, D.; Thaya- paran, M.; Ranaldi, L.; and Freitas, A., eds., Proceedings of The 3rd Workshop on Mathematical Natural Language Pro- cessing (MathNLP 2025), 150–167. Suzhou, China: Asso- ciation for Computational Linguistics. [Peng et al. 2023] Peng, C.; Xia, F.; Naseriparsa, M.; and Osborne, F. 2023. Knowledge graphs: Opportunities and challenges.Artificial intelligence review 56(11):13071– 13102. [Poesia et al. 2022] Poesia, G.; Polozov, O.; Le, V.; Tiwari, A.; Soares, G.; Meek, C.; and Gulwani, S. 2022. Syn- chromesh: Reliable code generation from pre-trained lan- guage models. arXiv preprint arXiv:2201.11227. [Rheinfrank and Evenson 1996] Rheinfrank, J., and Even- son, S. 1996. Design languages. In Bringing design to software. 63–85. [Shaker, Togelius, and Nelson 2016] Shaker, N.; Togelius, J.; and Nelson, M. J. 2016. Procedural content generation in games. Springer. [Sudhakaran et al. 2023] Sudhakaran, S.; Gonz ́ alez-Duque, M.; Freiberger, M.; Glanois, C.; Najarro, E.; and Risi, S. 2023. Mariogpt: Open-ended text2level generation through large language models. Advances in Neural Information Processing Systems 36:54213–54227. [Tekinbas and Zimmerman 2003] Tekinbas, K. S., and Zim- merman, E. 2003. Rules of play: Game design fundamen- tals. MIT press. [Tilbrook and McMullen 1990] Tilbrook, D. M., and Mc- Mullen, J. 1990. Washing behind your ears: Principles of software hygiene. In Proceedings of the EurOpen Confer- ence. [Todd et al. 2023] Todd, G.; Earle, S.; Nasir, M. U.; Green, M. C.; and Togelius, J. 2023. Level generation through large language models. In Proceedings of the 18th International Conference on the Foundations of Digital Games, 1–8. [Togelius et al. 2011] Togelius, J.; Yannakakis, G. N.; Stan- ley, K. O.; and Browne, C. 2011. Search-based procedural content generation: A taxonomy and survey. IEEE Trans- actions on Computational Intelligence and AI in Games 3(3):172–186. [Ventura 2016] Ventura, D. 2016. Beyond computational intelligence to computational creativity in games. In 2016 IEEE conference on computational intelligence and games (CIG), 1–8. IEEE. [Yang et al. 2024] Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M.; Xue, M.; Ni, N.; Zhang, P.; Wang, P.; Peng, R.; Men, R.; Gao, R.; Lin, R.; Wang, S.; Bai, S.; Tan, S.; Zhu, T.; Li, T.; Liu, T.; Ge, W.; Deng, X.; Zhou, X.; Ren, X.; Zhang, X.; Wei, X.; Ren, X.; Fan, Y.; Yao, Y.; Zhang, Y.; Wan, Y.; Chu, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; and Fan, Z. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671. [Yin and Neubig 2017] Yin, P., and Neubig, G. 2017. A syntactic neural model for general-purpose code generation. In Proceedings of the 55th Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Pa- pers), 440–450. [Yuan et al. 2022] Yuan, A.; Coenen, A.; Reif, E.; and Ip- polito, D. 2022. Wordcraft: story writing with large lan- guage models. In Proceedings of the 27th International Con- ference on Intelligent User Interfaces, 841–852. [Zarrieß, Voigt, and Sch ̈ uz 2021] Zarrieß, S.; Voigt, H.; and Sch ̈ uz, S. 2021. Decoding methods in neural language gen- eration: A survey. Information 12(9):355. Table 4: IR schema iteration history from initial draft to frozen v0.2-runtime-evidence. #Iteration What ChangedWhyImpact 1v0 static draft Defined six top-level fields: scene, objects, scripts, params, links, rules. Minimal structured representation be- tween pattern description and Unity code generation. Established baseline schema consumed by all pipeline stages. 2MVP nar- row- ing Deferred params extraction; pipeline operates on objects, scripts, links, rules only. params emitted as. params requires GUID resolution (scene → prefab → script → .cs), blocking the initial pipeline. Unblocked end-to-end generation with- out serialized-field extraction. 3Runtime ex- ten- sion Added PrefabInstance/PrefabAsset object types, script-defined rules, and runtime params field. Static .unity parsing misses prefab- drivengameplay;corebehaviour emerges from prefab instantiation and runtime script logic. IR captures the actual gameplay loop; generation can reason about spawned entities and runtime configuration. 4Per- instance con- straint scripts[].object id must refer- ence a real objects[].id.No implicit aggregate placeholders (e.g. circle all). Aggregate placeholders create ambigu- ous references unresolvable during code generation or evaluation. Enforces 1:1 script-to-object binding; enables automatic referential integrity validation. 5Evidence- aware se- man- tics Conditionalrelationlabels(e.g. cantriggergamewinifaligned). Required evidencetypeon every rules[] entry.Optional confidence field. Unconditional labels over-assert deter- minism for conditional code paths. Evi- dence attribution improves trust calibra- tion in generated output. Generation produces more accurate causal claims; evaluation can filter or weight rules by evidence type and con- fidence. Appendix: IR iteration record See Table 4 Appendix: Full 26-Pattern Set with IR Statistics for Freezing Schema See Table 5 and Table 6 Appendix: IR v0.2-runtime-evidence Summary Top-level fields (all required). scene: string objects: [ id, name, type , ...] scripts: [ id, object_id, class_name , ...] params: runtime_params: "<id>": ... , ... links: [ source, target, relation, evidence_type? , ...] rules: [ id, type, description, pattern, evidence_type, confidence? , ...] Hard constraints. 1) scripts[].object_id must reference objects[].id 2) scripts are per-instance (no sharing across objects) 3) no implicit aggregate placeholders 4) rules[].evidence_type is required in direct_code, scene_override, inferred Annotated example: 1Ownership. The following is the reference IR extracted from the Unity project for the Owner- ship goal pattern. Structural fields (objects, scripts, runtimeparams) are derived from static scene YAML analysis; semantic fields (links relations, rules) are hand-authored to encode runtime behaviour evidenced in script source code. "scene": "1_Ownership", "objects": [ "id": "118191953", "name": "Canvas", "type": "GameObject" , "id": "519420028", "name": "Main Camera", "type": "GameObject" , "id": "1570331856", "name": "Text (Legacy)", "type": "GameObject" , "id": "1012039051484866332", "name": "Goal Manager", "type": "PrefabInstance" , "id": "1112099645", "name": "Player", "type": "PrefabInstance" , "id": "775309098", "name": "Boundary", "type": "PrefabInstance" , "id": "8743824104122491932", "name": "Game Manager", "type": "PrefabInstance" , "id": "9011082862537914474", "name": "Spawn Manager", "type": "PrefabInstance" , "id": Table 5: IR schema frequency tiers across all 26 patterns (v2). Thresholds: Core≥ 80%, Common≥ 40%, Optional < 40%. TierItemCoverage Top-level fields Core scene, objects, scripts, params, runtimeparams, links, rules 26/26 Object types Core GameObject, PrefabInstance26/26 Optional PrefabAsset4/26 Script classes Core GameManager26/26 Common SpawnManager19/26 Common GoalManager18/26 Link relations Core has prefabinstance, hascomponent 26/26 OptionalAll pattern-specific relations≤7/26 Rule types Core wincondition26/26 Common triggercount12/26 Runtime param keys Common spawnStart, spawnCount, spawnRepeat, spawnRangeX, spawnRangeY 18–19/26 Common goalCount, setGoal, currentCount 17–18/26 OptionalAll pattern-specific keys≤4/26 Rule evidence types Core direct code26/26 "prefab_057536c2a19bd9e4b8cdb1cb044a64f1", "name": "OwnershipObject","type": "PrefabAsset" ], "scripts": [ "id": "da1b...", "object_id": "9011082862537914474", "class_name": "SpawnManager" , "id": "74fe...", "object_id": "1012039051484866332", "class_name": "GoalManager" , "id": "bf0c...", "object_id": "prefab_057536c2a19bd9e4b8cdb1cb044a64f1", "class_name": "ChangeColor" , "id": "game_manager", "object_id": "8743824104122491932", "class_name": "GameManager" ], "params": , "runtime_params": "da1b...": "spawnStart": true, "spawnCount": 8, "spawnRepeat": false, "spawnPrefabGuid": "057536c2..." Table 6: Full set of 26 goal patterns with IR structural statis- tics (v0.2-runtime-evidence). All 26 scenes pass referential integrity and evidence-type validation. PatternObjectsScriptsLinksRules 1 Ownership94123 2 Collection74133 3 Eliminate75133 4 Capture76143 5 Overcome85144 6 Evade107163 7 Stealth1610223 8 Herd/Attract1816353 9 Conceal1913283 10 Rescue1813294 11 Delivery1316353 12 Guard1929603 13 Race197183 14 Alignment1311294 15 Configuration2811233 16 Traverse197183 17 Survive106114 18 Connection5372 19 Exploration93112 20 Reconnaissance106153 21 Contact74103 22 Enclosure108173 23 Gain Competence6382 24 Gain Information96143 25 Last Man Standing209233 26 King of the Hill169233 Total32122153378 , "74fe...": "goalCount": 8, "setGoal": true , "links": [ "source": "scene", "target": "9011082862537914474", "relation": "has_prefab_instance" , "source": "9011082862537914474", "target": "da1b...", "relation": "has_component" , "source": "da1b...", "target": "prefab_057536c2...", "relation": "spawns_prefab", "evidence_type": "direct_code" , "source": "bf0c...", "target": "74fe...", "relation": "increments_current_count_on_trigger", "evidence_type": "direct_code" , "source": "74fe...", "target": "game_manager", "relation": "can_trigger_game_win_if_count_met", "evidence_type": "direct_code" ], "rules": [ "id": "rule_spawn_on_start", "type": "spawn", "description": "SpawnManager spawns OwnershipObject prefab on Start when spawnStart is true.", "pattern": "Ownership", "evidence_type": "direct_code", "confidence": 1.0 , "id": "rule_collect_changes_color_and_counts", "type": "trigger_count", "description": "When Player enters OwnershipObject trigger, object color is set to player color and GoalManager.currentCount increases by 1.", "pattern": "Ownership", "evidence_type": "direct_code", "confidence": 1.0 , "id": "rule_win_when_count_reaches_goal", "type": "win_condition", "description": "If GoalManager.setGoal is true and currentCount equals goalCount, GameManager.GameWin() is called.", "pattern": "Ownership", "evidence_type": "direct_code", "confidence": 1.0 ] Appendix: Prompt Templates (Verbatim) No-schema prompt [pattern: <PATTERN_ID>] [method: no_schema] Generate a Unity Editor script that implements the playable concept described below. Output only raw C# code. <PATTERN_MD> With-schema prompt (IR to C#) [pattern: <PATTERN_ID>] [method: <METHOD>] Generate a Unity Editor script that instantiates a scene matching the following engine-specific Intermediate Representation (IR). Thereafter, you may refer to it as IR. Output only raw C# code. <IR_JSON> IR maker prompt (free) [pattern: <PATTERN_ID>] [method: with_schema_free] Generate an engine-specific Intermediate Representation (IR) JSON for the playable concept described below. Thereafter, you may refer to it as IR. Output ONLY valid JSON. No extra text. <PATTERN_MD> IR maker prompt (min skeleton) [pattern: <PATTERN_ID>] [method: with_schema_min] Generate an engine-specific Intermediate Representation (IR) JSON for the playable concept described below. Thereafter, you may refer to it as IR. Output ONLY valid JSON. No extra text. Required top-level fields: "scene" -- string "objects" -- [ "id", "name", "type" , ... ] "scripts" -- [ "id", "object_id", "class_name" , ... ] "params" -- "runtime_params" -- "<id>": ... , ... "links" -- [ "source", "target", "relation" , ... ] "rules" -- [ "id", "type", "description", "pattern", "evidence_type" , ... ] <PATTERN_MD> IR maker prompt (full schema) [pattern: <PATTERN_ID>] [method: with_schema_full] Generate an engine-specific Intermediate Representation (IR) JSON for the playable concept described below. Thereafter, you may refer to it as IR. Output ONLY valid JSON. No extra text. Follow the IR v0.2-runtime-evidence schema precisely. Top-level fields (all required): "scene" -- string, scene identifier "objects" -- array of "id", "name", "type" type in "GameObject", "PrefabInstance", "PrefabAsset" "scripts" -- array of "id", "object_id", "class_name" one entry per component instance on one object "params" -- always "runtime_params" -- object keyed by scripts[].id; values are flat field: value maps "links" -- array of "source", "target", "relation", "evidence_type"? "rules" -- array of "id", "type", "description", "pattern", "evidence_type", "confidence"? Hard constraints: 1. Every scripts[].object_id MUST reference a real objects[].id (no dangling refs). 2. Scripts are per-instance; no shared script entries across objects. 3. Every entity must be listed explicitly in objects (no aggregate placeholders). 4. Every rules[] entry MUST include evidence_type in "direct_code", "scene_override", "inferred" . <PATTERN_MD> Appendix: Pattern-Level Error Distribution by Model Tables 7–14 provide per-model breakdowns of G and H fail- ure counts for all four configurations. Table 7: Pattern-level errors: no schema, DeepSeek-Coder- V2-Lite (timeout 37–51%; 20 logs per pattern). G = ground- ing failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership18018— 2Collection000— 3 Eliminate36036— 4Capture12012— 5 Overcome404— 6 Evade30300CS0115(15), CS0246(15) 7 Stealth38380CS0115(19), CS0246(19) 8HerdAttract220CS0115(1), CS0246(1) 9 Conceal40040— 10Rescue17170CS0122(17) 11Delivery27027— 12Guard101— 13 Race000— 14 Alignment000— 15 Configuration202— 16Traverse12012— 17Survive101— 18Connection202— 19Exploration101— 20 Reconnaissance101— 21Contact990CS0246(9) 22Enclosure220CS0246(2) 23GainCompetence202— 24 GainInformation36036— 25 LastManStanding202— 26KingoftheHill17017— Table 8: Pattern-level errors: noschema, Qwen2.5-Coder- 7B (timeout 37–51%; 20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership000— 2 Collection23815CS0115(2), CS0234(2) 3Eliminate101— 4Capture29272CS0234(16), CS0115(11) 5 Overcome110CS0234(1) 6Evade20218CS0115(1), CS0234(1) 7Stealth14122CS0234(7), CS0115(5) 8HerdAttract22418CS0115(2), CS0234(2) 9 Conceal1376CS0115(3), CS0239(2) 10Rescue1459CS0246(5) 11Delivery27270CS0234(16), CS0115(11) 12 Guard16160CS0234(8), CS0115(7) 13Race963CS0234(4), CS0115(1) 14Alignment21516CS0234(3), CS0115(2) 15 Configuration312CS0246(1) 16Traverse21210CS0234(19), CS0115(2) 17Survive871CS0115(4), CS0234(2) 18 Connection18612CS0115(2), CS0234(2) 19Exploration303— 20Reconnaissance1064CS0246(6) 21Contact21192CS0234(9), CS0115(7) 22 Enclosure532CS0115(1), CS0234(1) 23 GainCompetence321CS0246(1), CS0311(1) 24GainInformation19109CS0239(4), CS0115(3) 25 LastManStanding13013— 26KingoftheHill20416CS0246(2), CS0115(1) Table 9: Pattern-level errors: withschemafree, DeepSeek- Coder-V2-Lite (timeout 86–92%; 20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership1192CS0246(6), CS1061(3) 2 Collection19019— 3 Eliminate24024— 4 Capture19019— 5 Overcome808— 6Evade14014— 7Stealth19019— 8HerdAttract18018— 9Conceal20020— 10 Rescue19019— 11Delivery14014— 12Guard17017— 13Race927CS0246(2) 14 Alignment17017— 15 Configuration18018— 16 Traverse16115CS1061(1) 17Survive19019— 18Connection18018— 19Exploration19019— 20Reconnaissance17017— 21 Contact14014— 22Enclosure20020— 23GainCompetence19019— 24GainInformation25025— 25 LastManStanding19019— 26 KingoftheHill18018— Table 10: Pattern-level errors: withschemafree, Qwen2.5- Coder-7B (timeout 86–92%; 20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership20020— 2Collection28028— 3Eliminate22220CS0117(2) 4 Capture28028— 5Overcome17017— 6Evade32032— 7Stealth22022— 8HerdAttract25124CS0246(1) 9Conceal27027— 10 Rescue26026— 11Delivery49049— 12Guard24024— 13Race17017— 14Alignment20020— 15Configuration27027— 16 Traverse16016— 17Survive22121CS1624(1) 18Connection19019— 19Exploration42042— 20Reconnaissance27225CS0246(2) 21 Contact39039— 22 Enclosure33033— 23 GainCompetence25025— 24GainInformation37037— 25LastManStanding25025— 26KingoftheHill17017— Table 11:Pattern-level errors:withschemamin, DeepSeek-Coder-V2-Lite (timeout 96%;20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership361323CS0246(10), CS1061(3) 2Collection19019— 3Eliminate19019— 4Capture20020— 5Overcome231112CS0246(6), CS1061(5) 6 Evade20416CS0246(2), CS1061(2) 7Stealth18018— 8HerdAttract19019— 9Conceal20020— 10 Rescue20020— 11Delivery20020— 12Guard18018— 13Race15213CS0246(1), CS1061(1) 14Alignment24024— 15 Configuration19019— 16Traverse18018— 17Survive18018— 18Connection30030— 19Exploration18315CS0246(1), CS0311(1), CS1061(1) 20 Reconnaissance16016— 21Contact19019— 22Enclosure19019— 23GainCompetence20020— 24 GainInformation26422CS0246(4) 25 LastManStanding17017— 26KingoftheHill18018— Table 12: Pattern-level errors: withschemamin, Qwen2.5- Coder-7B (timeout 96%; 20 logs per pattern). G = ground- ing failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership35134CS0246(1) 2Collection31031— 3Eliminate24717CS0246(4), CS8121(2) 4Capture18018— 5 Overcome21417CS0246(1), CS0311(1) 6Evade18018— 7Stealth18018— 8HerdAttract17116CS0246(1) 9Conceal20020— 10 Rescue29029— 11Delivery24420CS0246(4) 12Guard27027— 13Race21219CS0246(2) 14Alignment23023— 15Configuration26917CS0246(5), CS0315(4) 16 Traverse18018— 17Survive251114CS0246(8), CS0311(2) 18Connection18117CS0246(1) 19 Exploration20020— 20 Reconnaissance21120CS1061(1) 21 Contact18018— 22Enclosure27225CS0246(2) 23GainCompetence18018— 24GainInformation17017— 25 LastManStanding20218CS0246(2) 26 KingoftheHill20020— Table 13: Pattern-level errors: withschemafull, DeepSeek- Coder-V2-Lite (timeout 97–99%; 20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership251015CS0246(10) 2 Collection19019— 3Eliminate22022— 4 Capture17017— 5 Overcome30426CS0246(2), CS1061(2) 6 Evade18018— 7Stealth18018— 8HerdAttract17017— 9Conceal21021— 10 Rescue20020— 11Delivery21021— 12Guard18018— 13Race835CS0246(3) 14 Alignment20119CS0246(1) 15 Configuration28028— 16 Traverse22022— 17Survive20020— 18Connection19019— 19Exploration20020— 20Reconnaissance17116CS0246(1) 21 Contact16016— 22Enclosure20020— 23GainCompetence23122CS0246(1) 24GainInformation22022— 25 LastManStanding18018— 26 KingoftheHill19019— Table 14: Pattern-level errors: withschemafull, Qwen2.5- Coder-7B (timeout 97–99%; 20 logs per pattern). G = grounding failure; H = hygiene failure PatternTotalGHTop 2 G codes 1Ownership24618CS0246(5), CS0117(1) 2Collection28820CS0246(6), CS1061(2) 3Eliminate21021— 4Capture18018— 5 Overcome24618CS0246(3), CS1061(3) 6Evade19019— 7Stealth26620CS0246(3), CS1061(3) 8HerdAttract19019— 9 Conceal23023— 10Rescue26026— 11Delivery20020— 12Guard18018— 13Race19118CS0246(1) 14Alignment35926CS0246(4), CS1061(4) 15 Configuration28028— 16Traverse19019— 17Survive18018— 18Connection20020— 19 Exploration20317CS0246(3) 20 Reconnaissance18216CS0246(1), CS1061(1) 21 Contact19514CS0246(5) 22Enclosure26125CS0246(1) 23GainCompetence19019— 24 GainInformation19019— 25 LastManStanding17017— 26KingoftheHill14014—