Paper deep dive
AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment
Yixuan Wang, Yue Huang, Hong Qian, Yunzhao Wei, Yifei Ding, Wenkai Wang, Zhi Liu, Zhongjing Huang, Aimin Zhou, Jiajun Guo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/21/2026, 11:26:06 AM
Summary
AlphaContext is an evolutionary tree-based psychometric context generator designed to automate the creation of high-quality, long-form creativity assessment contexts. The system addresses the scarcity of expert-designed contexts and the limitations of current LLM-based generators (such as weak narrative coherence and limited diversity) through a four-stage pipeline: the HyperTree Outline Planner for hierarchical planning, an MCTS-based Context Generator for structural and semantic quality, an Evolutionary Context Optimizer using MAP-Elites for stylistic diversity, and an Assessment-Guided Evolution Refiner that uses virtual participant simulation to ensure assessment validity. Experimental results on the CreaTE dataset demonstrate that AlphaContext outperforms several competitive LLM-based methods across multiple quality metrics.
Entities (10)
Relation Signals (7)
AlphaContext â contains â HyperTree Outline Planner
confidence 100% ¡ AlphaContext comprises four modules: the HyperTree Outline Planner, the MCTS-based Context Generator, the Evolutionary Context Optimizer and the Assessment-Guided Evolution Refiner.
AlphaContext â contains â MCTS-based Context Generator
confidence 100% ¡ AlphaContext comprises four modules: the HyperTree Outline Planner, the MCTS-based Context Generator...
AlphaContext â contains â Evolutionary Context Optimizer
confidence 100% ¡ AlphaContext comprises four modules: ... the Evolutionary Context Optimizer and the Assessment-Guided Evolution Refiner.
AlphaContext â contains â Assessment-Guided Evolution Refiner
confidence 100% ¡ AlphaContext comprises four modules: ... the Assessment-Guided Evolution Refiner.
HyperTree Outline Planner â uses â HyperTree
confidence 100% ¡ the HyperTree Outline Planner formalizes expert-designed outlining as a rule-guided hypertree
Evolutionary Context Optimizer â uses â MAP-Elites
confidence 100% ¡ the Evolutionary Context Optimizer evolves contexts with MAP-Elites
MCTS-based Context Generator â uses â Monte Carlo Tree Search
confidence 100% ¡ The MCTS-based Context Generator fills the outline via MCTS
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Creativity has become a core competence in the era of LLMs and human-AI collaboration, underpinning innovation in real-world problem solving. Crucially, the systematic improvement of creativity necessitates scientifically valid assessment instruments. Psychometric research recognizes context-based assessment as an effective way to measure creative thinking. However, high-quality expert-designed contexts remain scarce. Existing LLM-based generators often struggle with insufficient assessment cues, weak narrative coherence, limited stylistic diversity, and poor support for creative thinking. To address these challenges, we propose AlphaContext, an evolutionary tree-based psychometric context generator for creativity assessment. First, the HyperTree Outline Planner formalizes expert-designed outlining as a rule-guided hypertree and performs top-down hierarchical planning. The MCTS-based Context Generator fills the outline via MCTS to balance global structure and local quality. Then, the Evolutionary Context Optimizer evolves contexts with MAP-Elites by repeatedly updating niche elites to jointly improve diversity and quality. Finally, the Assessment-Guided Evolution Refiner simulates virtual participants with diverse styles and recycles weak contexts for further evolution. Experiments show that AlphaContext yields an average improvement of 8% over competitive methods across 6 quality metrics.
Tags
Links
- Source: https://arxiv.org/abs/2604.18398v1
- Canonical: https://arxiv.org/abs/2604.18398v1
Trouble viewing inline? Open PDF directly â
Full Text
88,038 characters extracted from source content.
Expand or collapse full text
AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment Yixuan Wang 1 * Yue Huang 1 * Hong Qian 1,2â Yunzhao Wei 1 Yifei Ding 1 Wenkai Wang 1 Zhi Liu 1,2 Zhongjing Huang 1 Aimin Zhou 1,2 Jiajun Guo 1 1 East China Normal University, Shanghai, China 2 Shanghai Innovation Institute, Shanghai, China yxwang, yhuang@stu.ecnu.edu.cn, hqian@cs.ecnu.edu.cn Abstract Creativity has become a core competence in the era of LLMs and humanâAI collaboration, underpinning innovation in real-world problem solving. Crucially, the systematic improvement of creativity necessitates scientifically valid as- sessment instruments. Psychometric research recognizes context-based assessment as an ef- fective way to measure creative thinking. How- ever, high-quality expert-designed contexts re- main scarce. Existing LLM-based generators often struggle with insufficient assessment cues, weak narrative coherence, limited stylistic di- versity, and poor support for creative thinking. To address these challenges, we propose Alpha- Context, an evolutionary tree-based psychome- tric context generator for creativity assessment. First, the HyperTree Outline Planner formal- izes expert-designed outlining as a rule-guided hypertree and performs top-down hierarchical planning. The MCTS-based Context Genera- tor fills the outline via MCTS to balance global structure and local quality. Then, the Evolution- ary Context Optimizer evolves contexts with MAP-Elites by repeatedly updating niche elites to jointly improve diversity and quality. Fi- nally, the Assessment-Guided Evolution Re- finer simulates virtual participants with diverse styles and recycles weak contexts for further evolution. Experiments show that AlphaCon- text yields an average improvement of 8% over competitive methods across 6 quality metrics. 1 Introduction Creativity is typically defined as the ability to gener- ate novel and appropriate ideas (Runco and Jaeger, 2012), and it is a crucial skill that drives social innovation and scientific discovery (Sternberg and Lubart, 1999). As AI increasingly takes over rou- tine tasks, creativity is becoming an even more important driver of original contributions and trans- formative breakthroughs (Lee, 2022). * Equal Contribution â Corresponding Author Context: AI Partner Agentle melodystreamedfromthe smartspeakerbyEmmaâsbed,whilethe ambient lightinggraduallyshiftedtoa fresh, pale blue.ThiswastheâMorning Wake-Up ServiceâprovidedbyherAI partner,Q-Bot.Itcouldaccurately detect Emmaâs sleep stagesandbegin herdayuniquelywiththemostsuitable combinationofsoundandlight.Atits core,Q-Botwasanintelligent system connected to the cloudâitshuman-like form merelyservedasaninterface for real-world interaction [......] Pleaseidentifythepotentialchallenges arisingfromthewidespreadadoptionof âAIPartnerâ,listingasmanyasyoucan. Design the Context [Anchor] [Open Task] Context Design Process Design Principles [Arts & Aesthetics][Basic Needs][Business & Commerce][Comm unication][Culture & Religion][.....] [Anchor][Scene Setting] [Characters & Interaction] [Conflict & Challenge][Open Task] AI Partner companionship humanâAI ethics... Creative Association Challenge Cue Answer human-like form [Ethics & Morality] Participant1 felt lonely [Psychological Health] Participant2 Complete the Assessment Creativity Context Assessment a fresh,pale blue [Arts & Aesthetics] detect sleep stages [Physical Health] Answer 1 Answer 2 Answer 1 Answer 2 Figure 1: Workflow of creativity context assessment. Experts first design an âAI Partnerâ psychometric con- text containing implicit challenge cues. Participants then complete an open-ended task to identify potential challenges based on the context in their responses. Given this growing importance, scientifically as- sessing creativity has become a key issue in psycho- metrics and intelligent education (Jr. et al., 2024; Wang et al., 2025b; Guo and Woulfin, 2016). In cre- ativity assessment, context-based paradigms have been widely adopted (Barbot et al., 2019b). Stern- bergâs triarchic theory of intelligence emphasizes that creative thinking arises when faced with novel situations (Sternberg, 1984). Therefore, future- oriented contexts, due to their inherent high uncer- tainty and broad imaginative space, are considered ideal stimuli for eliciting creativity thinking (Koh and Leung, 2019), as shown in Figure 1. The Fu- ture Problem Solving Program (FPSP) (Torrance et al., 1976) provides authoritative evidence that well-designed future contexts can reliably elicit creative thinking (Crabbe, 1982). Therefore, gen- erating high-quality contexts is essential for de- veloping valid and reliable creativity assessments arXiv:2604.18398v1 [cs.CL] 20 Apr 2026 for humans. However, current practice faces a sig- nificant bottleneck in productivity. High-quality creativity assessment contexts still rely on expert craftsmanship. In recent years, rapid advances in LLMs have substantially improved the generation of stories and dialogues (Li et al., 2025a; Feng et al., 2025; Wu et al., 2025a), making the automated construction of psychometric contexts for the assessment of cre- ativity increasingly plausible. Although much prior work has examined the creativity of LLMs, there has been far less research on whether LLMs can generate valid creativity assessment contexts for humans. However, psychometric contexts differ from general narratives, and both directly influ- encing LLMs and reusing story-generation frame- works still face two key challenges. The first key challenge is to simultaneously enforce implicit assessment cues and global narrative coherence. Psychometric cues for creative thinking are often embedded implicitly in textual details. Existing methods struggle to precisely control the alignment between cues and themes (Bai et al., 2025), thus failing to satisfy the fine-grained constraints re- quired for psychometric content design and nar- rative structuring. The second key challenge is to improve diversity while ensuring both context quality and measurement validity at limited cost. For a given theme, future problem contexts require diverse types and styles to adapt to different as- sessment populations (Novikov et al., 2025), yet increasing diversity typically raises generation and refinement costs. Moreover, creativity assessment contexts require reliable quality and elicitation va- lidity. Traditional expert workflows as shown in Figure 1, rely on expensive human studies and it- erative rework (Wu et al., 2025b), while current methods lack efficient validation and optimization loops. Our work has implications for employing LLMs to automatically generate valid and reliable creativity assessment contexts for humans and AI. To address these two challenges, this paper pro- poses AlphaContext, an evolutionary tree-based psychometric context generator for creativity as- sessment. To tackle the first challenge, the Hyper- Tree Outline Planner formalizes context outlining as a rule-guided hypertree, mapping expert reason- ing into a searchable outline space. The MCTS- based Context Generator then performs Monte Carlo Tree Search (MCTS) generation under the outline, balancing global structural coherence and local semantic quality to produce seed contexts. To handle the second challenge, the Evolutionary Con- text Optimizer conducts evolutionary search with MAP-Elites in a task-specific behavioral space, it- eratively expanding stylistic diversity via niche- wise elite updates. Finally, the Assessment-Guided Evolution Refiner simulates virtual participant re- sponses and iteratively refines weak contexts to better elicit creative thinking. Experimental results show that AlphaContext substantially outperforms baselines across multiple evaluation metrics. 2 Related Work 2.1 Language Model based Creativity Assessment The NLP community has shown growing inter- est in psychometrically grounded creativity assess- ment. Luchini et al. (Luchini et al., 2025) fine- tuned RoBERTa and GPT-2 to automatically score creativity responses. Since assessment contexts are expert-designed and costly to scale, CPIG investi- gates using LLMs to automatically generate items for a classic free-response creativity test, examin- ing whether LLMs can generate valid creativity assessments for humans. However, CPIG targets short items and does not address long-form Future Problem contexts, which require discourse-level coherence and implicit assessment cues. In parallel, other works evaluate the creativ- ity of LLMs (Si et al., 2025; Fang et al., 2025). LiveIdeaBench (Ruan et al., 2024) uses single- keyword prompts to assess scientific creative think- ing. AidanBench (Mclaughlin et al., 2024) mea- sures novelty, non-redundancy, and coherence un- der open-ended creative questions. Nevertheless, these benchmarks rely on keyword triggers or expert-designed contexts. Moreover, the scarcity of high-quality long-form contexts limits both eval- uation protocols and model improvement. Overall, the key bottleneck is generating high- quality long-form psychometric contexts that en- able valid creativity assessment for humans and also support more comparable evaluation of LLM creativity. To address this gap, we propose Alpha- Context, an evolutionary tree-based psychometric context generator for creativity assessment. 2.2 LLM-based Story Generation LLMs have significantly advanced automated nar- rative generation (Bai et al., 2025; Wu et al., 2025a; Lee et al., 2025), enabling the generation of co- herent long-form stories. DOC (Yang et al., 2023) (a) Model Input (b) Tree-based ContextGeneration Context Query MCTS-based Context Generator Evolutionary Context Optimizer Assessment-Guided Evolution Refiner Designafuture-oriented creativityproblemcontext titledâAIPartnerâabout humanâAIcompanionship andautonomy,focusingon ethics,emotionalreliance, privacy,andgovernance ofpervasivepersonalAI assistants. Title Theme Seed Context Operators Insertion Deletion Replacement Mutation Evolve Archive Evaluation Elite Update Simulated Participants Output Context (c) Context Evolution and Refinement HT-Decide HT-Select HT-Expand HT-Construct HyperTree Outline Planner ... LLM LLM LLM LLM LLM MCTS-Backpropagate MCTS-Select MCTS-Expand MCTS-Evaluate Talkative Normal Quiet ... Answers LLM Scorer Creativity Scores Outline Figure 2: The procedure of the proposed AlphaContext. (a) Given a context queryQ, (b) the HyperTree Outline Planner and MCTS-based Context Generator generate seed contexts, and (c) the Evolutionary Context Optimizer and Assessment-guided Evolution Refiner improve diversity and quality, yielding assessment-ready contexts. adopts the outline-first strategy and then expands the outline into detailed text. STORYTELLER (Li et al., 2025a) introduces a plot node mechanism based on the subject-verb-object (SVO) structure and a dynamic interaction module, further improv- ing narrative coherence and logical consistency. However, most existing story generation methods typically focus on entertainment and fluency, of- ten failing to satisfy the quality and validity re- quirements of psychometric assessment contexts. Although S-GEN (Feng et al., 2025) explores gen- erating psychological social stories for autism in- terventions, this setting differs fundamentally from creativity assessment settings. Creativity assess- ment contexts must maintain coherent long-form narratives while implicitly placing assessment cues that elicit creative thinking and support psychomet- ric validity. Consequently, generating high-quality long-form psychometric contexts for creativity as- sessment remains an open challenge. 3 Preliminaries Creativity Context Generation. Given a con- text design queryQthat specifies the title and theme, together with a pre-trained LLMĎ Î¸ , our goal is to construct an assessment-ready archive A = C k |A| k=1 of contexts for creativity assess- ment. Each context is represented asC = (T,O), whereOis a structured outline, andTis the re- sulting context text guided byO. We adopt a plan- generate-evolve pipeline. The planner produces O = Ď Î¸ (ÎŚ(Q))under predefined instructionsÎŚ. The generator realizesTguided by the outlineO. The evolve stage iteratively refines and diversifies candidates by updatingA. HyperTree Structure. In conventional tree-based planning, each edge links a parent node to a single child node. By contrast, a HyperTree introduces directed hyperedges, where a parent node connects to a set of child nodes via one edge, enabling hier- archical divide-and-conquer by jointly organizing discourse structure and assessment-cue placement for outline planning. Formally, we define a Hy- perTree asH = (N,Q,R), whereQdenotes the query,Nis the node set, andRis a set of expansion rules. GivenQ, the HyperTree is generated hier- archically according to the rule setR. Compared with ordinary trees, this structure better aligns with expert practices in creativity context design. 4 The Proposed AlphaContext Overview. As shown in Figure 2, AlphaCon- text comprises four modules: the HyperTree Out- line Planner, the MCTS-based Context Genera- tor, the Evolutionary Context Optimizer and the Assessment-Guided Evolution Refiner. Given a queryQ, the HyperTree Outline Planner places expert outlining as a rule-guided hypertree over a libraryR. The MCTS-based Context Generator then performs a sentence-level search to fill in the outline. To cover the multi-solution space under the same theme, the Evolutionary Context Opti- mizer applies MAP-Elites to explore diverse styles in a task-specific behavior space while improv- ing within-niche quality. Finally, the Assessment- Guided Evolution Refiner simulates participant re- sponses and feeds weak contexts back for further evolution. 4.1 HyperTree Outline Planner Experts plan contexts holistically and refine them hierarchically, motivating a HyperTree representa- tion. We propose the HyperTree Outline Planner to cast outline design as a HyperTree search, where directed hyperedges support hierarchical divide- and-conquer over structure and cue placement. For- mally, we define a HyperTree asH = (N,Q,R). GivenQ, HyperTreeHis generated hierarchically underR. Each nodenâ Ncorresponds to a struc- tural unit, and each ruler â Rexpands a parent node into a set of child nodes throughr : n p 7â n c , wheren p is the parent node andn c denotes the corresponding child nodes. The HyperTree (HT) planner proceeds in four phases: HT-Select, HT- Expand, HT-Construct, and HT-Decide. HT-Select. Given the current hypertreeH, its dis- tinct branches are mapped onto a set of hyperchains L 1 ,...,L g , wheregdenotes the number of hy- perchains. To control the search scale, an LLM is used to evaluate and prune candidate hyperchains. Each candidate hyperchain is scored by the LLM and the optimal hyperchainsL â are selected. The divisible nodes are then identified under the rule setR. For each selected hyperchain, we choose its most promising divisible leaf node for expansion using an LLM selector. Thus, this phase consists of two steps: selecting the hyperchainsL â and se- lecting the divisible leaf noden â i in each chosen hyperchain L â i âL â . HT-Expand. In this phase, given the selected node n â i in each hyperchainL â i , its applicable expansion rules are retrieved asR(n â i ) =r âR| r : n â i 7â n c . For each ruler â R(n â i ), candidate child groupsn c are generated. Each group is treated as a single branch and is appended toL â i as the expansion outcome of n â i . HT-Construct. The planner iterates Select and Ex- pand step by step, growingHfrom the root node across depth levels over a set of selected hyper- chains. Construction stops when no divisible nodes remain or the iteration limit is reached, yielding a hypertree that compactly stores multiple candidate hyperchains. HT-Decide. After constructingH, the LLM glob- ally evaluates the candidate hyperchains and de- cides the optimal hyperchain as the final outlineO. This decision jointly considers structural validity for creativity assessment context design and nar- rative consistency with the input title and theme. More details can be found in Appendix G. 4.2 MCTS-based Context Generator In the generation stage, we cast creativity context writing as a sentence-level decision process guided by an outlineO. Given an input promptxand an outlineO, we build a separate search tree for each discourse section. A context is represented as C = (T,O), whereT = (t 1 ,...,t p )is the gener- ated sentence sequence. A node at depthpiss p = t p ,N (s p ),V (s p ),O , wheret p is the current text, N (s p )is the visit count, andV (s p )is the estimated value. An LLM policyĎ Î¸ proposes the next can- didate sentences, and an LLM evaluator provides quality feedback. The generator follows the stan- dard MCTS loop: MCTS-Select, MCTS-Expand, MCTS-Evaluate, and MCTS-Backpropagate. MCTS-Select. Starting from the roots 0 , the search recursively selects the child with the highest explo- ration potential according to the UCT score. Specif- ically, the UCT score of the nodes p is defined as UCT(s p ) = V (s p ) + c q lnN(q) N(s p ) . Here,V (s p )de- notes the value score ofs p ,N (s p )is its visit count, andN (q)is the visit count of its parent nodeq. cis a hyper-parameter that balances exploitation (V (s p )) and exploration (the second term). MCTS-Expand.Given the selected nodes p , we expand it by samplingUcandidate next sen- tences from the policy modelĎ Î¸ :t (u) p+1 âź Ď Î¸ (¡| x, t 1:p , O), u = 1,...,U. Here,t 1:p denotes the previously generated sentences, enabling parallel exploration of diverse narrative realizations and cue instantiations within O. MCTS-Evaluate. We evaluate each expanded node to assign its node valueV (s p+1 ). For long- form creativity context generation, we adopt a dual-horizon valuation mechanism at the evalua- tion phase to balance reliability and computational cost. Given an expanded childs p+1 , we first apply multi-aspect immediate scoring with an evaluator: V imm (s p+1 ) = Ě S(s p+1 ) 1â S ha (s p+1 ) .(1) Here, Ě S(s p+1 )is a weighted average of cue align- mentS sc , imagery vividnessS im , and discourse co- herence S co with P i Ď i = 1. S ha (s p+1 ) measures hallucination risk. To mitigate myopic decisions, whenV imm (s p+1 ) < Ď, we sample a short continu- ation and re-evaluate the concatenated fragment to obtain a more stable value estimate, a lightweight look-ahead for this node. This look-ahead is trig- gered for low-scoring nodes to reduce myopic er- rors, while high-confidence nodes directly use the immediate evaluation to save sampling budget. For more details, please refer to our Appendix F. MCTS-Backpropagate. The obtained evaluation scorer e is propagated back along the simulated path to all ancestor nodess j (0⤠j ⤠p), updating the visit counts and value estimates: N new (s j ) = N old (s j ) + 1, (2) V new (s j ) = V old (s j )N old (s j ) + r e N new (s j ) . (3) After multiple simulations, the tree concentrates on trajectories that better satisfy the outline, improve coherence, and reduce hallucination risk. We then extract the highest-value root-to-leaf path as a seed context to initialize the evolutionary module. 4.3 Evolutionary Context Optimizer We introduce a MAP-Elites Evolutionary Context Optimizer initialized with the MCTS seed context. It maintains an elite archive in a style-oriented be- havior space, expanding coverage and improving within-niche quality. We next describe the archive, mutation, evaluation, and update rules. Diversity Archive. To characterize stylistic vari- ations in creativity assessment contexts, we map each candidate contextCinto the behavior space B with a descriptor function b(¡): b(C) = Ď 1 (C),Ď 2 (C),Ď 3 (C) â [0, 1] 3 . (4) Here,Ď 1 captures proximity scope, measuring the extent to which a context is framed from personal daily-life settings to broader public issues.Ď 2 cap- tures knowledge density, reflecting how strongly the narrative is grounded in objective evidence such as data, mechanisms, and causal explanations.Ď 3 captures viewpoint diversity, indicating the breadth of stakeholders involved and the need for multi- perspective integration. We uniformly discretize [0, 1] 3 to form a 3D grid archive, where each cell defines a behavioral niche and stores the current elite context with the highest fitness. Mutation. In natural language space, we im- plement mutation as a conditional LLM edit- ing policyĎ Î¸ that edits a parent elite context C p . At each iteration, we apply an operator set ⌠= INSERTION, DELETION, REPLACEMENT to revise key paragraphs and cue-bearing units, introducing stylistic variation while maintaining assessment-critical content. Evaluation. Given a mutated candidateC, we per- form both behavior feature evaluation and quality evaluation. Feature evaluation computesb(C)to determine the candidateâs niche assignment. Qual- ity evaluation is produced by an LLM-based scorer in terms of narrative coherenceS coh , topical rel- evanceS rel , and engagementS eng , and we de- fine fitnessF (C) = Avg S coh (C) + S rel (C) + S eng (C) . The scorer outputs three normalized quality scores for each candidate context, and the fitness is defined as their uniform average to avoid introducing extra hyperparameters. Elite Update. After evaluation, we assign context Cto the grid niche indexed byb(C). If the niche is empty,Cis inserted as the initial elite. Otherwise, letC â denote the current elite in the niche; we replace C â only when F (C) > F (C â ). 4.4 Assessment-Guided Evolution Refiner To ensure that generated creativity contexts reliably elicit measurable creative thinking, we propose an Assessment-Guided Evolution Refiner. Concretely, we instantiate an LLM-based participant simulator with a temperature set to0to suppress sampling randomness and stabilize response generation, mak- ing contexts comparable in a consistent simulation setting. To improve realism and interpretability, we model response styles as explicit profiles grounded in common psychometric response patterns and en- force each profile via role-conditioned prompting. We consider three profilesâtalkative, normal, and quiet. Given a candidate contextC, the simulator generates a set of responsesY m M m=1 . The refiner then scores each response using a creativity scorer f cre (¡). We define the assessment effectiveness of a context as the average creativity score across stylesΨ(C) = 1 M P M m=1 f cre (Y m ) . IfΨ(C)ex- ceeds an expert-specified threshold, we treatCas an assessment-ready context. Otherwise, we route Cback to the Evolutionary Context Optimizer for further iterative optimization. 5 Experiments This section first describes the CreaTE dataset and details the evaluation metrics used. We then conduct extensive experiments to answer the fol- lowing research questions. The codes and data are available athttps://github.com/yxwang19/ AlphaContext. Table 1: Performance comparison across different methods on the CreaTE dataset. AlphaContext achieves the best results across all seven perspectives. All metrics are presented as positive percentages, where higher values indicate better performance. For each metric, the best-performing model is highlighted in bold and the second isunderlined. Methods Coherence (â)% Relevance (â)% Engagement (â)% Significance (â)% Concreteness (â)% Uncertainty (â)% Diverse Verbs (â)% DeepSeek-V3.150.0050.0050.0050.0050.0050.0094.33 Qwen3-235B-A22B41.5047.9145.0733.8746.5553.3392.09 Llama3.3-70B-Instruct27.8326.9733.9927.4630.5420.0790.39 LongWriter-llama3.1-8b26.6027.4628.6323.4033.6225.9991.18 LongWriter-glm4-9b 32.2731.4031.3836.1932.9831.3288.69 GPT-5.170.4470.2065.3950.3771.8068.6092.88 Gemini-3.0-Pro-Preview72.5475.3762.5648.4064.1663.3091.81 DOC-v249.1461.3361.8234.9851.1143.1092.82 CRITICS51.1161.9561.2137.8154.4342.1292.31 S-GEN 60.2269.6956.4060.1051.8553.5790.24 AlphaContext81.2879.0679.9371.0675.4980.3096.06 Q1: How does AlphaContext compare to exist- ing methods in generating high-quality creativity contexts across multiple evaluation dimensions? Q2: To what extent do the core components con- tribute to the performance of AlphaContext? Q3: How does AlphaContext compare to other methods in terms of textual similarity with expert- designed contexts? Q4: Does the LLM-based judge align with human preferences to support reliable evaluation? Q5: Does AlphaContext remain effective for cre- ativity assessment in real-world human studies? Q6: How does the correlation of AlphaContext with expert-designed assessments compare to that of the strong LLM baseline? Q7: What is the computational cost of AlphaCon- text in terms of generation time and token consump- tion? 5.1 Experimental Setup Dataset. We evaluate AlphaContext on CreaTE. Since general story-generation datasets prioritize narrative completeness and style rather than as- sessment alignment, we construct CreaTE: 203 expert-curated titleâtheme inputs balancing evalu- ation cost and domain coverage. Three creativity- psychology experts write each entry and conduct iterative cross-checks. An entry is included only af- ter consensus validation and revision, and all inputs are further screened to remove sensitive informa- tion. Details are provided in Appendix A. Baselines. To compare AlphaContext with both strong LLMs and generation frameworks, we con- sider baselines from three categories. The prompt template is provided in Appendix I. (i) General-purpose LLMs.We consider DeepSeek-V3.1, Qwen3-235B-A22B, Llama3.3- 70B-Instruct, GPT-5.1, and Gemini-3.0-Pro- Preview as competitive instruction-following mod- els with strong general reasoning and generation capabilities. (i) Long-form specialized LLMs. LongWriter-llama3.1-8b and LongWriter-glm4- 9b (Bai et al., 2025) are included as specialized baselines for long-form writing. (i) Structured generation frameworks. DOC-v2 (Yang et al., 2023) combines hierarchical outlining with an ad- herence controller. CRITICS (Bae and Kim, 2024) performs critic-guided iterative refinement. S- GEN (Feng et al., 2025) applies constraint-driven hierarchical prompting (STARSOW) for structured story generation. We do not compare with CPIG since it generates short test items and is not de- signed for long-form context generation. Evaluation Metrics. We evaluate each gener- ated creativity context along 7 dimensions. Co- herence measures narrative consistency, Relevance measures theme alignment, and Engagement mea- sures how motivating the context is for partici- pants. Our evaluation framework is theoretically grounded in Amabileâs Componential Model of Creativity (Amabile, 1983, 2018), which empha- sizes task motivation, domain-relevant grounding, and creativity-related processes as core founda- tions of creative performance. We further evaluate Significance (Okuda et al., 1991; Mumford et al., 2018) to capture real-world relevance and intrinsic task motivation, Concreteness (Guegan et al., 2017) to reflect situational specificity that supports feasi- ble ideation, and Uncertainty (Beghetto and Jaeger, 2022; Beghetto, 2021) to quantify open-endedness that fosters divergent thinking rather than prema- ture closure. Table 2: Ablation study of AlphaContext on multiple evaluation metrics. Details are the same as Table 1. Methods Coherence (â)% Relevance (â)% Engagement (â)% Significance (â)% Concreteness (â)% Uncertainty (â)% Diverse Verbs (â)% -w/o HOP77.9670.2076.8563.5570.6976.1194.25 -w/o MCG74.3871.8072.1765.7669.0971.9293.79 -w/o ECO75.6270.5771.8064.5368.7270.6993.36 AlphaContext81.2879.0679.9371.0675.4980.3096.06 Table 3: Text similarity to expert-designed contexts measured by ROUGE-1, ROUGE-L, and BERTScore. Higher is better. Methods ROUGE-1 (â)% ROUGE-L (â)% BERTScore (â)% DeepSeek-V3.126.2220.5380.94 Qwen3-235B-A22B 20.0316.4280.39 Llama3.3-70B-Instruct20.2316.3280.81 LongWriter-llama3.1-8b 19.8816.3381.42 LongWriter-glm4-9b24.4220.3481.33 GPT-5.115.2512.4679.28 Gemini-3.0-Pro-Preview22.8918.5980.31 DOC-v2 18.2415.6579.87 CRITICS17.8315.1479.74 S-GEN27.8021.3381.07 AlphaContext30.4125.4881.88 These psychometric dimensions were iteratively refined through extensive reviews by senior experts in creativity psychology and aligned with estab- lished standards from the Future Problem Solving Program (FPSP) (Torrance et al., 1976), which sup- ports strong content validity. Following arena-hard- auto (Li et al., 2025b), we use contexts generated by DeepSeek-V3.1 as the reference baseline and obtain quantified scores through pairwise compar- isons with other generated contexts. We also report Diverse Verbs (Fan et al., 2019) to measure action diversity in the context. The prompt template is pro- vided in Appendix I.2. We also report ROUGE-1, ROUGE-L, and BERTScore to measure similar- ity to expert-designed contexts. Empirically, our real-world human study further supports construct validity by showing significant positive correlations with standardized creativity measures. Implementation Details. All open-source mod- els are locally deployed and run with vLLM on 8Ă NVIDIA H200 GPUs. For DOC-v2, CRITICS, and S-GEN, DeepSeek-V3.1 is used as the generation engine. Additionally, DeepSeek-V3.1 serves as the evaluator model for all judgments. To mitigate po- sition bias in pairwise judgments, we evaluate each context pair twice with swapped orders, and repeat this procedure for two rounds, resulting in four evaluations per pair. We omit standard deviations in tables since they are consistently small. 5.2 Experimental Results and Analysis Main Performance Evaluation (To Q1). We com- pare AlphaContext with 10 baselines on CreaTE to assess multi-dimensional context quality for cre- ativity assessment. Diverse Verbs is computed automatically, while Coherence, Relevance, En- gagement, Significance, Concreteness, and Uncer- tainty are judged by an LLM. We follow arena-hard- auto (Li et al., 2025b) and conduct pairwise compar- isons against the DeepSeek-V3.1 output as the ref- erence model. For each subjective metric, we report the positive rate over the reference, which is set at 50% by definition. Table 1 summarizes the results on CreaTE. AlphaContext ranks first on all seven metrics, with the largest gains on Coherence, En- gagement, Significance, and Uncertainty, which are key for constructing coherent stimuli that implic- itly cue challenges and elicit open-ended creative thinking. Notably, S-GEN surpasses GPT-5.1 and Gemini-3.0-Pro-Preview on Significance, while Al- phaContext further raises it to 71.06%. This sug- gests that AlphaContextâs planning, search-based generation, and iterative optimization collectively drive consistent improvements across metrics. Ablation Study (To Q2). To quantify module contributions, we build three ablated variants by removing the HyperTree Outline Planner (HOP), the MCTS-based Context Generator (MCG), and the Evolutionary Context Optimizer (ECO). Ta- ble 2 shows that removing any module degrades performance. In particular, removing HOP yields a sharp Relevance drop to 70.20%, indicating that hierarchical planning is critical for keeping assess- ment cues aligned with the intended theme. Re- moving MCG lowers Coherence to 74.38% and also reduces Engagement and Uncertainty, suggest- ing that MCTS search helps preserve long-range structure and the open-endedness needed to elicit creative thinking. Removing ECO decreases all metrics, most notably Uncertainty, showing that MAP-Elites refinement is important for expanding (a) AlphaContext vs. GPT-5.1 (Human) (b) AlphaContext vs. GPT-5.1 (DeepSeek-V3.1) Figure 3: Preference evaluation of AlphaContext vs. GPT-5.1 under human and DeepSeek-V3.1 judgments. (a) AlphaContext vs. Gemini (Human) (b) AlphaContext vs. Gemini (DeepSeek-V3.1) Figure 4: Preference evaluation of AlphaContext vs. Gemini-3.0-Pro-Preview under human and DeepSeek- V3.1 judgments. stylistic coverage while maintaining assessment cues. This decline arises from two complemen- tary roles of the ECO. First, it expands stylistic coverage by maintaining niche-specific elites in a 3D behavior space, which broadens the diver- sity of generated contexts. Second, it acts as an effective quality filter through the iterative muta- tionâevaluationâupdate loop, which polishes con- sistency and theme alignment beyond raw MCTS seeds. The ablation results thus validate that the MAP-Elites refinement simultaneously enhances stylistic diversity and core quality while preserv- ing assessment cues. Overall, the ablation results show that the three modules make complementary contributions. Context Similarity Evaluation (To Q3). To assess how closely AlphaContext outputs match expert- designed contexts given the same inputs, we use 16 expert contexts that were deployed in real creativity assessments and validated by domain experts as ref- erences. We compute ROUGE-1, ROUGE-L, and BERTScore, which reflect lexical overlap, long- span matching, and semantic similarity, respec- tively. Table 3 shows that AlphaContext achieves the best results in all three metrics. It reaches 30.41% ROUGE-1 and 25.48% ROUGE-L, outper- forming S-GEN by 2.61% and 4.15%, and also 0.0 0.2 0.4 0.6 0.8 1.0 0 2 4 6 8 10 Participants' Creativity Scores Counts Gauss Fit (a) 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 Creativity Score Fitted Y 95% Confidence Band Creativity Score - AlphaContext (b) Figure 5: Human study results for AlphaContext. (a): distribution of participant creativity scores with a Gaus- sian fit. (b): Pearson correlation between AlphaContext- based scores and AUT (Alternative Uses Task) scores. surpasses strong LLM baselines in BERTScore. Overall, AlphaContext matches the most closely expert-designed contexts, aligning with its goal of generating psychometrically grounded creativity assessment materials. For more details, please re- fer to our Appendix B. Preference Evaluation (To Q4). To examine whether the LLM judge aligns with human prefer- ences, we compare AlphaContext with strong base- lines (GPT-5.1 and Gemini-3.0-Pro-Preview) via pairwise preference judgments from both human evaluators and DeepSeek-V3.1. To mitigate posi- tion bias, each pair is judged twice with swapped orders, and we report win, tie, and lose rates, where win indicates a preference for AlphaContext. Fig- ure 3 and Figure 4 show DeepSeek-V3.1 closely matches human preferences. Against GPT-5.1, Al- phaContext wins 62.07% under human judgments and 60.10% under DeepSeek-V3.1, with lose rates around 13%. Against Gemini-3.0-Pro-Preview, Al- phaContext wins 73.89% under human judgments and 67.49% with DeepSeek, with low lose rates of 2.96% and 5.91%. The humanâLLM agreement is high (CohenâsÎş > 0.8), supporting the reliability of the LLM judgments for our main evaluations. Real-World Human Study (To Q5). We vali- date the assessment effectiveness of AlphaContext- generated contexts through a real-world study with 36 secondary-school students. As shown in Fig- ure 5a, the scores exhibit a symmetric unimodal pattern, and the Gaussian fit closely matches the empirical histogram, suggesting a stable measure- ment behavior. For criterion validity, we compare AlphaContext-based scores with those of the Alter- native Uses Task (AUT), a widely used standard- ized creativity test. The two scores are collected independently, and the Pearson correlation shows a significant positive association (r = 0.3770), as shown in Figure 5b. Notably, according to stan- dard guidelines on psychology and creativity as- sessment (Gignac and Szodorai, 2016; Funder and Ozer, 2019; Runco and Acar, 2012; Beaty and Johnson, 2021; Beaty et al., 2022; Benedek et al., 2013), a correlation coefficient ofr = 0.3770is regarded as practically meaningful and provides reasonable support for criterion validity. This re- sult indicates that the creativity levels elicited by our generated contexts are consistent with an estab- lished benchmark. Overall, the human study pro- vides real-world evidence that AlphaContext can measure student creativity in authentic educational settings. More details can be found in Appendix D. Case Study (To Q6). To evaluate measurement- level alignment with expert assessment, we con- duct a controlled case study on the same theme us- ing three contexts: expert-designed, AlphaContext- generated, and Gemini-3-Pro-Preview-generated. We simulate 30 virtual participants with diverse response styles, collect creativity scores, and com- pare the induced rankings using Spearman correla- tion andR 2 fit. Figure 6 shows that AlphaContext better matches the expert context (r s = 0.84) than Gemini (r s = 0.58), indicating closer outcome- level consistency with expert-designed assessments. Details are provided in Appendix E. (a) AlphaContext vs. Expert(b) Gemini vs. Expert Figure 6: Case study on measurement-level alignment with expert assessment. Scatter plots compare expert- induced score ranks (x-axis) with ranks induced by gen- erated contexts (y-axis) for 30 simulated participants. Computational Cost Analysis (To Q7). We an- alyze the computational cost of AlphaContext in terms of generation time and token consumption, in comparison with baseline LLMs and manual expert design. Table 4 reports the average time and token usage required to generate one creativity context. Although AlphaContext requires more to- kens and longer inference time than standard zero- shot prompting, this overhead comes from its full Table 4: Average generation time and token consump- tion per context. MethodTime (s)Tokens (k) GPT-5.123.502.36 Gemini-3.0-Pro-Preview46.313.29 AlphaContext (Ours)226.9912.89 pipeline and is necessary to ensure psychometric validity beyond what simple prompting can reli- ably provide. In practice, AlphaContext generates one context in about 6 minutes, making it practical for high-quality dataset construction. By contrast, manual expert design typically requires at least one week (Crabbe, 1989; Barbot et al., 2019a). Al- phaContext therefore substantially reduces human labor while maintaining quality. In addition, it uses a locally deployed open-source model, avoiding costly closed-source APIs and improving trans- parency. Overall, the added computational cost is justified by gains in validity, reliability, and au- tomation. 6 Conclusion This paper introduces AlphaContext, an evolution- ary generator for psychometric assessment con- texts that integrates rule-guided outline planning, sentence-level MCTS generation, MAP-Elites qual- ityâdiversity optimization, and assessment-guided refinement through virtual participant simulation. Across extensive experiments, AlphaContext con- sistently outperforms strong LLM baselines and structured frameworks on 7 evaluation dimensions, and shows higher alignment with expert-designed contexts. While it consumes more tokens and re- quires longer generation time than baseline LLMs, its computational overhead is fully acceptable, and it simultaneously achieves significantly higher as- sessment validity and generation stability. Hu- manâLLM preference evaluations support reliable automated judging, and a real-world study provides practical validity evidence. Overall, AlphaContext offers a scalable way to produce contexts for cre- ativity assessment while reducing the reliance on scarce expert writing. AlphaContext is designed as a context generator for human creativity assessment, while also provid- ing standardized contexts for benchmarking LLM creativity. Current experiments focus on future- oriented contexts and a compact expert-curated in- put set; extending to broader domains, age groups, and languages is an important direction. Limitations AlphaContext primarily targets generating psy- chometrically grounded creativity assessment con- texts that are usable in real testing settings. How- ever, achieving stable discourse-level coherence and measurement-relevant cue control relies on sentence-level MCTS and MAP-Elites refinement, which require repeated model calls and scoring. As a result, the overall generation cost depends not only on AlphaContextâs design but also on the un- derlying LLM and judge configuration, making effi- ciency comparisons sensitive to the chosen models and evaluation setup. In addition, psychometrically suitable titleâtheme inputs are still relatively scarce, which can affect the scale of benchmarking and the extent to which conclusions transfer beyond our current future-oriented setting. AlphaContext offers strong controllability and assessment alignment, while direct prompting is typically cheaper but less reliable for measurement- oriented constraints. Notably, context generation is not a real-time or online task, and the computa- tional overhead is offset by massive reductions in manual expert effort. In future work, we plan to expand expert-curated inputs and use AlphaCon- text to produce high-quality training data for fine- tuning of lightweight generators. This will improve efficiency while preserving assessment utility, en- abling broader coverage across domains, popula- tions, and deployment settings. Ethical Considerations This work introduces the CreaTE dataset for creativity-context generation. The titleâtheme in- puts are authored by creativity-psychology experts under a shared specification and contain no per- sonal information. We screen the dataset to remove sensitive or identifiable content, and we prioritize both data quality and ethical compliance during cu- ration. In particular, we apply strict quality-control procedures, including thorough manual review, to ensure broad coverage while proactively address- ing potential bias and sensitivity concerns and we curate the dataset in accordance with applicable privacy and research-ethics standards. Our work also includes a real-world human study to validate the assessment effectiveness of AlphaContext-generated contexts. This study has been reviewed and approved by the Institutional Re- view Board (IRB) of the affiliated university (IRB Approval No. HR2-0478-2025). Participation was voluntary, and all participants were informed of the study purpose and procedures, with the right to withdraw at any time without penalty. Before the study, we obtained written informed consent from participants and their guardians, and the consent materials specified the study goals, tasks, potential risks, and data use and protection measures. All collected responses were anonymized by removing personal identifiers, stored securely with restricted access, and reported only in aggregate. The study involved minimal risk, as participants completed open-ended creativity tasks similar to typical class- room activities, and we did not request or record sensitive personal information. Acknowledgements We would like to thank the anonymous review- ers for constructive comments. This work is sup- ported by the National Natural Science Foundation of China (No. 62476091), the General Program in Education of the National Social Science Fund of China (No. BEA230071), and the Key Program in Education of the National Social Science Fund of China (No. ABA220028). References T.M. Amabile. 1983. The social psychology of creativ- ity: A componential conceptualization. Journal of Personality and Social Psychology, 45(2):357. T.M. Amabile. 2018. Creativity in context: Update to the social psychology of creativity. Routledge. Minwook Bae and Hyounghun Kim. 2024. Collective critics for creative story generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 18784â18819, Miami, FL, USA. Yushi Bai, Jiajie Zhang, Xin Lv, Linzhi Zheng, Siqi Zhu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2025. Longwriter: Unleashing 10,000+ word generation from long context llms. In Proceedings of the 13th International Conference on Learning Representations, Singapore. B. Barbot, R.W. Hass, and R. Reiter-Palmon. 2019a. Creativity assessment in psychological research:(re) setting the standards. Psychology of Aesthetics, Cre- ativity, and the Arts, 13(2):233. Baptiste Barbot, Richard W Hass, and Roni Reiter- Palmon. 2019b. Creativity assessment in psychologi- cal research:(re)setting the standards. Psychology of Aesthetics, Creativity, and the Arts, 13(2):233. R.E. Beaty and D.R. Johnson. 2021. Automating cre- ativity assessment with semdis: An open platform for computing semantic distance. Behavior Research Methods, 53(2):757â780. R.E. Beaty, D.R. Johnson, D.C. Zeitlen, and B. Forth- mann. 2022. Semantic distance and the alternate uses task: Recommendations for reliable automated assessment of originality. Creativity Research Jour- nal, 34(3):245â260. Ronald A Beghetto. 2021. There is no creativity without uncertainty: Dubito ergo creo. Journal of Creativity, 31:100005. Ronald A Beghetto and Garrett J Jaeger. 2022. Uncer- tainty: A catalyst for creativity, learning and devel- opment. Creativity Theory and Action in Education, 6. M. Benedek, C. MĂźhlmann, E. Jauk, and A.C. Neubauer. 2013. Assessment of divergent thinking by means of the subjective top-scoring method: Effects of the number of top-ideas and time-on-task on reliability and validity. Psychology of Aesthetics, Creativity, and the Arts, 7(4):341. A.B. Crabbe. 1989. The future problem solving pro- gram. Educational Leadership, 7(1):27â29. Anne Borland Crabbe. 1982. Creating a brighter future: An update on the future problem solving program. Journal for the Education of the Gifted, 5(1):2â11. Angela Fan, Mike Lewis, and Yann N. Dauphin. 2019. Strategies for structuring story generation. In Pro- ceedings of the 57th Conference of the Association for Computational Linguistics, pages 2650â2660, Flo- rence, Italy. Xinyu Fang, Zhijian Chen, Kai Lan, Lixin Ma, Shengyuan Ding, Yingji Liang, Xiangyu Zhao, Farong Wen, Zicheng Zhang, Guofeng Zhang, Haodong Duan, Kai Chen, and Dahua Lin. 2025. Creation-mmbench: Assessing context-aware cre- ative intelligence in MLLM.arXiv preprint, abs/2503.14478. Yi Feng, Mingyang Song, Jiaqi Wang, Zhuang Chen, Guanqun Bi, Minlie Huang, Liping Jing, and Jian Yu. 2025. S-GEN: A social story generation frame- work with large language models. In AAAI-25, Spon- sored by the Association for the Advancement of Ar- tificial Intelligence, pages 1300â1308, Philadelphia, PA, USA. D.C. Funder and D.J. Ozer. 2019. Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2):156â168. G.E. Gignac and E.T. Szodorai. 2016. Effect size guide- lines for individual differences researchers. Person- ality and Individual Differences, 102:74â78. JĂŠrĂ´me Guegan, Julien Nelson, and Todd Lubart. 2017. The relationship between contextual cues in virtual environments and creative processes. Cyberpsychol- ogy, Behavior, and Social Networking, 20(3):202â 206. Jiajun Guo and Sarah Woulfin. 2016. Twenty-first cen- tury creativity: An investigation of how the part- nership for 21st century instructional framework re- flects the principles of creativity. Roeper Review, 38(3):153â161. Antonio Laverghetta Jr., Simone Luchini, Averie Lin- nell, Roni Reiter-Palmon, and Roger E. Beaty. 2024. The creative psychometric item generator: a frame- work for item generation and validation using large language models. In Proceedings of the 3rd Work- shop on Artificial Intelligence and Creativity co- located with 27th European Conference on Artificial Intelligence, pages 59â73, Santiago de Compostela, Spain. Brandon Koh and Angela K-y Leung. 2019. A time for creativity: How future-oriented schemas facilitate creativity. Journal of Experimental Social Psychol- ogy, 84:103816. Hye-Kyung Lee. 2022. Rethinking creativity: Creative industries, ai and everyday creativity. Media, Culture & Society, 44(3):601â612. Kuang-Huei Lee, Ian Fischer, Yueh-Hua Wu, Dave Marwood, Shumeet Baluja, Dale Schuurmans, and Xinyun Chen. 2025. Evolving deeper LLM thinking. arXiv preprint, arXiv:2501.09891. Jiaming Li, Yukun Chen, Ziqiang Liu, Minghuan Tan, Lei Zhang, Yunshui Li, Run Luo, Longze Chen, Jing Luo, Ahmadreza Argha, Hamid Alinejad-Rokny, Wei Zhou, and Min Yang. 2025a. STORYTELLER: an enhanced plot-planning framework for coherent and cohesive story generation. In Findings of the Asso- ciation for Computational Linguistics, pages 20818â 20846, Vienna, Austria. Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2025b. From crowdsourced data to high- quality benchmarks: Arena-hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learning, Vancouver, BC, Canada. Luchini, Simone A, Maliakkal, Nadine T, DiStefano, Paul V, Laverghetta Jr, Antonio, Patterson, John D, Beaty, Roger E, and Roni Reiter-Palmon. 2025. Auto- mated scoring of creative problem solving with large language models: A comparison of originality and quality ratings. Psychology of Aesthetics, Creativity, and the Arts. Aidan Mclaughlin, James Campbell, Anuja Uppuluri, and Yiming Yang. 2024. Aidanbench: Stress-testing language model creativity on open-ended questions. In NeurIPS 2024 Workshop on Language Gamifica- tion. Peter Meusburger. 2009. Milieus of creativity: The role of places, environments, and spatial contexts. Milieus of creativity: An interdisciplinary approach to spatiality of creativity, 2:97â153. Michael D Mumford, Robert Martin, Samantha Elliott, and Tristan McIntosh. 2018. Creative thinking in the real world. The nature of human creativity, pages 147â65. Alexander Novikov, Ngân Vu, Marvin Eisenberger, Em- ilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abi- gail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog. 2025. Alphaevolve: A coding agent for scientific and algorithmic discovery. arXiv preprint, arXiv:2506.13131. Shawn M Okuda, Mark A Runco, and Dale E Berger. 1991. Creativity and the finding and solving of real- world problems. Journal of Psychoeducational as- sessment, 9(1):45â53. Kai Ruan, Xuan Wang, Jixiang Hong, and Hao Sun. 2024. Liveideabench: Evaluating llmsâ scientific creativity and idea generation with minimal context. arXiv e-prints, pages arXivâ2412. M.A. Runco and S. Acar. 2012. Divergent thinking as an indicator of creative potential. Creativity Research Journal, 24(1):66â75. Mark A Runco and Garrett J Jaeger. 2012. The standard definition of creativity. Creativity research journal, 24(1):92â96. Mark A Runco, Burak Turkman, Selcuk Acar, and Mustafa V Nural. 2017. Idea density and the cre- ativity of written works. Journal of Genius and Emi- nence, 2(1):26â31. Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2025. Can llms generate novel research ideas? A large- scale human study with 100+ NLP researchers. In Proceedings of the 13th International Conference on Learning Representations, Singapore. Robert J Sternberg. 1984. Toward a triarchic theory of human intelligence. Behavioral and Brain Sciences, 7(2):269â287. Robert J Sternberg and Todd I Lubart. 1999. The con- cept of creativity: Prospects and paradigms. Hand- book of creativity, 1(3-15). E. P. Torrance, C. B. Bruch, and J. P. Torrance. 1976. Interscholastic futuristic creative problem-solving. The Journal of Creative Behavior, 10(2):117â125. Tyler J VanderWeele. 2025. Intellectual and viewpoint diversity: Importance, scope and bounds. Education Sciences, 15(12):1592. Xueyang Wang, Wei Liu, Kaixiang Zhuang, Cheng Liu, Jingyi Zhang, Li Fan, Qunlin Chen, and Jiang Qiu. 2025a.Neural representations of noncen- tral events during narrative encoding predict subse- quent story ending originality. Science Advances, 11(17):eadu5251. Yixuan Wang, Jiale Feng, Yue Huang, Xuruo Pan, Zhongjing Huang, Zhi Liu, and Hong Qian. 2025b. A style-aware polytomous diagnostic model for in- dividual traits. In Proceedings of the 28th European Conference on Artificial Intelligence, pages 2698â 2705, Bologna, Italy. Hongqiu Wu, Weiqi Wu, Tianyang Xu, Jiameng Zhang, and Hai Zhao. 2025a. Towards enhanced immer- sion and agency for llm-based interactive drama. In Proceedings of the 63rd Annual Meeting of the Asso- ciation for Computational Linguistics, pages 11166â 11182, Vienna, Austria. Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, and Jianfeng Gao. 2025b. Collabllm: From passive responders to active collaborators. In Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. Kevin Yang, Dan Klein, Nanyun Peng, and Yuandong Tian. 2023. DOC: improving long story coherence with detailed outline control. In Proceedings of the 61st Annual Meeting of the Association for Com- putational Linguistics, pages 3378â3465, Toronto, Canada. Appendix A CreaTE Dataset AlphaContext takes a title and a theme as input, so evaluation requires inputs that are explicitly de- signed for creativity assessment. Existing story- generation datasets are ill-suited for this purpose: they target narrative completeness and stylistic rich- ness, but do not ensure that the topic focus, im- plicit cues, and task intent satisfy psychometric requirements. Using such datasets would confound the evaluation, as poor performance could arise from mismatched inputs rather than the generation method. To address this gap, we curate CreaTE, a com- pact yet high-quality input set for creativity-context generation. CreaTE contains 203 titleâtheme pairs spanning diverse domains, balancing evaluation cost with broad coverage. Each entry is authored by three creativity-psychology experts under a shared specification: the title provides a concrete anchor, while the theme delineates the intended problem space and key creative tensions to be elicited, as shown in Figure 7. A subset of expert-designed titles and themes is adapted from publicly avail- able FPSP topic resources 1 . We refine the dataset through cross-checking and iterative revision, and include an entry only after expert consensus on as- sessment relevance, clarity, and correctness, with edits applied to remove ambiguity or unintended cues. Finally, we conduct compliance screening to ensure that no sensitive or identifiable information is present. [ "title":"YouthinCompetitiveSports", "theme":"Youthcompetitivesports:performancepressure, mentalhealth,injuryrisk,equityofaccess,and technology-enhancedtraining." , "title":"NanotechnologyinDailyLife", "theme":"Nanotechnology:smartmaterials,medicine, environmentalcleanup,manufacturingchange,andlong-term safetyandregulation." , "title":"NeurotechnologyFutureScene", "theme":"MedicalRehabilitation-FocusedNeurotechnology: Rehabapplications,stateethicsreview,consent,braindata security." , ] ...... Figure 7: Example input format of CreaTE. Each entry includes a title and a theme, designed to support psycho- metrically grounded creativity context generation. B Preference Evaluation We provide additional details on the human evalua- tion protocol and inter-annotator agreement analy- sis for the preference study involving Gemini-3.0- Pro-Preview. For human judgments, we recruited three eval- uators and adopted a standardized rating protocol with a pre-study calibration session. The human evaluation checklist is strictly aligned with our metric definitions (Coherence, Relevance, Engage- ment, Significance, Concreteness, Uncertainty) to ensure consistent interpretation. Inter-rater agree- ment among the three evaluators meets the required consistency standard, and we report the mean of their judgments as the final human preference re- sult. Consistent with observations in the main text, DeepSeekâs judgments closely track human prefer- ences, with high humanâLLM agreement (Cohenâs Îş > 0.8), further validating the reliability of auto- mated evaluation in our experiments. 1 https://fpspi.org/topics/ C Cue Coverage and Diversity in a Case Comparison Under the same input theme, AlphaContext pro- duces a context fragment with noticeably broader assessment-cue coverage than S-GEN. Figure 8 shows that AlphaContext can surface multiple chal- lenge dimensions within a single coherent narrative move, so the scenario invites reasoning from sev- eral angles rather than focusing on only one. In con- trast, S-GEN typically centers the context around a single dominant cue, which yields a narrower cue footprint and fewer directions for subsequent idea exploration. This qualitative comparison aligns with our design goal: by planning cue placement at the outline level and enforcing outline-grounded generation, AlphaContext increases the diversity of assessment-relevant cues while keeping them im- plicitly integrated into the story, better supporting creativity assessment that aims to elicit open-ended and multi-perspective thinking. Thecommunityorganizer,Sarah,chimedin,'Wecanalso mobilizemorevolunteersandraisefundsthroughlocal campaignsandpartnershipswithbusinesses. ["Social Relationships","Economics","Business & Commerce"]. AlphaContext IntheturquoisewatersoffthecoastofPalau,adelicate ecosystemfacesaninvisiblethreat.Coralreefsthathavethrived formillennianowshowsignsofbleachingandstress. ["Environment"]. S-GEN Figure 8: Case comparison under the same input theme: AlphaContext embeds a broader and more diverse set of assessment cues in a coherent fragment, whereas S- GEN concentrates on a single dominant cue, resulting in narrower cue coverage. D Detailed Analysis of Human Study To complement the aggregate Pearson correlation reported in the Real-World Human Study, we fur- ther examine how specific creativity dimensions align between AlphaContext and a standardized benchmark using Spearman rank correlation. We used Pearson correlation in the main text because the aggregated total scores empirically exhibit an approximately unimodal, near-Gaussian distribu- tion, making Pearson an appropriate summary of linear association. In contrast, dimension-level cre- ativity scores can be more heterogeneous in both cognitive mechanisms and distributional shapes. We therefore adopt Spearman correlation to provide a rank-based and more robust analysis of whether the two assessments preserve consistent relative ordering at a finer granularity. Before the formal study, all participants com- pleted the assessment under a standardized admin- istration procedure to ensure comparability across individuals. The instructions, task materials, and time constraints were fixed and delivered in a con- sistent format. Participants were guided to com- plete the tasks independently and were discouraged from discussion or external assistance during the session. To protect privacy, we removed personally identifiable information from all collected records, used anonymized participant identifiers for subse- quent analysis, and stored the data in an access- restricted manner. Only de-identified responses and scores were used for reporting and correlation analysis. All human scores used in this analysis were pro- duced under a standardized expert rating protocol. Specifically, three psychology experts served as raters. A domain specialist first developed a de- tailed scoring rubric and dimension definitions, and then conducted fine-grained training for the raters before annotation. The training session explained each dimension with concrete guidance and cali- bration examples, ensuring a shared understanding of the scoring criteria and common failure cases. During scoring, the three raters evaluated responses independently. We then computed inter-rater agree- ment to verify consistency, and the agreement met the required standard. Finally, we report the mean score across the three raters as the human score for each dimension. Context_Fluency Context_Flexibility Context_Novelty Fluency_A Fluency_B Novelty_A Novelty_B Context_Fluency Context_Flexibility Context_Novelty Fluency_A Fluency_B Novelty_A Novelty_B -1.0 -0.5 0.0 0.5 1.0 Spearman Correlation Figure 9: Spearman rank-correlation heatmap between AlphaContext metrics and AUT metrics. Figure 9 shows the Spearman correlation matrix between AlphaContext dimension scores and Alter- native Uses Task (AUT) scores. The AUT metrics are reported for two AUT prompts, denoted by the suffixes_Aand_B. Overall, the heatmap suggests a meaningful but nuanced alignment. In particular, Context_Fluencyexhibits a consistent positive as- sociation with AUT fluency, supporting validity on ideational productivity. Meanwhile, correlations involvingContext_Noveltyare weaker and less directly matched to the standard AUT novelty. This pattern is expected and informative: AUT mea- sures unconstrained divergent thinking, whereas AlphaContext evaluates creativity elicited under explicit semantic and contextual constraints. The results indicate that our method aligns with estab- lished benchmarks on general fluency-related sig- nals, while capturing a context-dependent aspect of creativity that may not be fully reflected by stan- dard AUT scoring. E Extended Measurement-level Alignment Study Across Baselines To complement the main case study (To Q6), we extend the same measurement-level alignment anal- ysis to all baselines. For each method, we gener- ate a context under the same theme as the expert- designed reference, simulate responses from the same set of 30 virtual participants with diverse re- sponse styles, and compute creativity scores under each context. We then compare the outcome-level ranking induced by each method with the expert- induced ranking. Figure 10 summarizes the alignment results across baselines. We report Spearman rank cor- relation to quantify ranking consistency with the expert reference, and we also include the corre- spondingR 2 from a linear fit as a complemen- tary indicator of overall association. Overall, Al- phaContext achieves the strongest agreement with the expert reference, suggesting that it best pre- serves the relative creativity differences elicited by expert-designed measurement at the outcome level. Several long-form or structurally guided base- lines show moderate alignment, whereas general- purpose one-shot generators exhibit noticeably weaker agreement. This baseline-wide comparison provides a broader view of measurement consis- tency and further supports AlphaContextâs advan- tage beyond single-baseline comparisons. Figure 10: Extended Measurement-level Alignment Study Across Baselines. F Analysis of MCG Evaluator This section analyzes the evaluator used in the MCTS-based Context Generator (MCG). MCG for- mulates long-form context generation as a sentence- level tree search under a planned outline. Each node represents a partial context, and candidate continuations are explored through MCTS. To de- cide which branches to expand and keep, the MCG relies on an evaluator that scores each newly gener- ated sentence fragment with respect to creativity- assessment context quality. The evaluator considers three criteria. Cue align- mentS sc measures how well the generated sen- tence matches the outline and embeds assessment- relevant cues. It focuses on faithful adherence to the outline hints, including required themes, keywords, constraints, and intended challenge cat- egories, while encouraging implicit planting of meaningful problem cues such as trade-offs, con- straints, stakeholder tensions, and second-order ef- fects, instead of explicitly listing challenges. Im- agery vividnessS im measures how strongly the text evokes a vivid mental image and immersion through concrete situational and sensory details, including visual impressions, sounds, smells, phys- ical sensations, and emotions, following creativity measurement literature (Wang et al., 2025a). Dis- course coherenceS co evaluates whether the frag- ment reads smoothly within the evolving context, with stable entities, natural transitions, and clear causal and temporal continuity. For evaluation reliability, each criterion is rated on a 5-point Likert scale and then normalized to [0, 1], whereLis the Likert score. The three nor- malized scores are aggregated with coefficients Ď 1 ,Ď 2 ,Ď 3 , consistent with the main text, to form the scalar evaluation value used by MCG during the search. We examine three coefficient groups to assess sensitivity: Group 1 usesĎ 1 = Ď 2 = Ď 3 = 0.33; Group 2 usesĎ 1 = 0.4, Ď 2 = 0.3, Ď 3 = 0.3; Group 3 usesĎ 1 = 0.5, Ď 2 = 0.25, Ď 3 = 0.25. As shown in Figure 11 and Figure 12, the six sub- jective metrics exhibit only mild fluctuations across groups, indicating stable evaluation behavior un- der reasonable coefficient changes. Group 2 yields the highest overall average score, suggesting that moderately emphasizing cue alignment best bal- ances outline-grounded cue placement with vivid writing and coherent discourse. We therefore adopt Ď 1 = 0.4, Ď 2 = 0.3, Ď 3 = 0.3as the default setting for all subsequent experiments. Group 1Group 2Group 3 60 70 80 90 Average Score Figure 11: Comparison of the overall average scores across the three weight groups. G Details of the HyperTree Outline Planner This section provides additional details for the Hy- perTree Outline Planner (HTP), complementing the description in Section 4.1. We focus on how HTP operationalizes outline design as hypertree search, where directed hyperedges enable hierarchi- cal divide-and-conquer over discourse structure and assessment-cue placement. Beyond the high-level phases (HT-Select, HT-Expand, HT-Construct, and HT-Decide), we clarify the rule system that gov- erns admissible expansions and explain how LLM- based decisions are integrated to control search scale and semantic validity. Section G.1 specifies the construction rules of HTP. It includes a static hierarchical skeleton that defines the admissible outline topology, and a set of dynamic, LLM-driven selection rules that prune Group 1Group 2Group 3 60 70 80 90 Score (a) Coherence Group 1Group 2Group 3 60 70 80 90 Score (b) Relevance Group 1Group 2Group 3 60 70 80 90 Score (c) Engagement Group 1Group 2Group 3 60 70 80 90 Score (d) Significance Group 1Group 2Group 3 60 70 80 90 Score (e) Concreteness Group 1Group 2Group 3 60 70 80 90 Score (f) Uncertainty Figure 12: Effect of different MCG evaluator coefficient settings on the six subjective metrics. Subfigures (a)â(f) report the mean scores for Coherence, Relevance, En- gagement, Significance, Concreteness, and Uncertainty, respectively. candidate hyperchains and choose the next divisi- ble leaf node to expand. Section H.4 then presents a fully instantiated outline for the title AI Partner, illustrating how these rules concretely materialize into a structured outline suitable for creativity as- sessment context generation. G.1 HyperTree Construction Rules HTP follows a hybrid rule mechanism that com- bines a predefined context-free grammar with semantic-aware dynamic selection. The context- free grammar provides a static skeleton that con- strains the outline to valid hierarchical forms, while dynamic selection rules use an LLM to evaluate candidate hyperchains under a budget and to select the most promising divisible node for expansion. This hybrid design separates structural admissibil- ity from semantic quality control, allowing HTP to maintain global validity while remaining flexible to the input title and theme. The formal definitions of the static skeleton and the dynamic selection nodes are provided in Listing 1. H Details of the Evolutionary Context Optimizer This section provides additional implementation details of the Evolutionary Context Optimizer, with a focus on how we operationalize the style-oriented behavior space and how the MAP-Elites archive is constructed and updated. Initialized with the seed contexts produced by the MCTS-based Context Generator, this module searches in natural language space to jointly expand stylistic coverage and im- prove within-niche quality, yielding a diverse and assessment-ready context archiveA. H.1 Style-Oriented Behavior Space To characterize stylistic variations that are highly relevant to future-problem contexts for creativity assessment, we map each candidate contextCinto a task-specific behavior spaceBvia the descriptor b(C). Following the formulation in Eq. 4,b(C) consists of three interpretable dimensions. The first dimension,Ď 1 (C), captures proximity scope, mea- suring how the context is framed from personal daily-life settings to broader public and societal issues (Meusburger, 2009). The second dimen- sion,Ď 2 (C), captures knowledge density, reflecting how strongly the narrative is grounded in objec- tive evidence such as data cues, mechanistic ex- planations, and causal constraints (Runco et al., 2017). The third dimension,Ď 3 (C), captures view- point diversity, indicating the breadth of stakehold- ers represented in the context and whether multi- perspective considerations jointly shape the prob- lem space (VanderWeele, 2025). Together, these three axes form a controllable and interpretable co- ordinate system to organize stylistic diversity in creativity assessment contexts. H.2 Descriptor Evaluation, Mutation, and Archive Update In our implementation, the three behavior dimen- sions are computed by a fixed-template LLM-based rater, which takes the context text as input and re- turns normalized scores forĎ 1 (C),Ď 2 (C), and Ď 3 (C)in a structured JSON format. We then dis- cretize the continuous behavior space uniformly to construct a 3D grid archive, where each cell corre- sponds to a behavioral niche and stores the current elite context. To enable controllable style shifts in natural lan- guage, we implement mutation as a conditional LLM editing process. Each mutation step spec- 1# ==================== Part 1: Fixed Hierarchical Skeleton (Static Rules) ================== 2 31. [Plan] -> [Anchor][Scene Setting][Characters & Interaction][Conflict & Challenge][Open Task] # The plan can be divided into five core narrative aspects. 42. [Anchor] -> [Future Horizon][Place][Scale][Challenge Seeds 1] 5# Establishes the fundamental time, space, and scope coordinates. 63. [Scene Setting] -> [Scenario Frame][Constraint Hints][Challenge Seeds 2] 7# Defines context and constraints. 84. [Characters & Interaction] -> [Interaction Goal][Dispute Focus][Problem Slot][Challenge Seeds 3] 9# Constructs interpersonal dynamics. 105. [Conflict & Challenge] -> [Challenge Seeds 4][Creativity Triggers] 11# Introduces complicating factors. 126. [Open Task] -> [Challenge Identification][Solution Exploration] 13# Defines the student's objective. 14 15# ==================== Part 2: Dynamic Selection Nodes (LLM-Driven) ======================= 16# Logic A: When this node is selected for expansion, the LLM selects one option from the predefined candidate pool based on theme relevance. 179. [Future Horizon] -> NearFuture (5-15y) | MidFuture | FarFuture | Speculative 1810. [Scale] -> Community | National | International | Space 1911. [Scenario Frame] -> Everyday Life | Urban Infrastructure | Virtual-Reality Fusion | ... 20 2112. [Interaction Goal] -> Co-creation Workshop | Negotiation | Emergency Response | ... 2213. [Dispute Focus] -> Value Conflict | Resource Conflict | Trust Conflict | ... 2314. [Creativity Triggers] -> Uncertainty | Contradiction | Resource Constraints | ... 24 25# Logic B: When this node is selected for expansion, the LLM selects multiple options from the predefined candidate pool based on theme relevance. 2615. [Challenge Seeds 1] -> Select 2-3 seeds from Pool 2716. [Challenge Seeds 2] -> Select 2-3 seeds from Pool 2817. [Challenge Seeds 3] -> Select 3-4 seeds from Pool 2918. [Challenge Seeds 4] -> Select 4-5 seeds from Pool 30 3119. [Topic Phrase] -> LLM-generated phrase (6-8 words) # Summarizes the core conflict based on Title/Theme. 32 3320. [Constraint Hints] -> Select 2-3 from: Policy, Budget, Time Limit, Safety, etc. # Limits the solution space. Listing 1: Formal definitions of static rules and dynamic LLM-driven selection rules in the HyperTree Outline Planner. ifies target values along the three behavior axes and guides the model to revise the context through insertion, deletion, and replacement, so that the can- didate moves toward the desired niche while main- taining narrative readability and cue traceability for assessment use. After mutation, each candidate is evaluated in two aspects. Behavior feature eval- uation determines its niche assignment viab(C). Quality evaluation is produced by an LLM-based scorer over coherence, relevance, and engagement, and we compute fitness as the uniform average of these three normalized scores to avoid introduc- ing extra hyperparameters. The archive is updated with niche-wise elite replacement: a candidate is inserted if the niche is empty, and otherwise it re- places the current elite only when it achieves higher fitness. H.3 Iteration Budget The archive is constructed iteratively. At each it- eration, we sample elites from the current archive, generate niche-targeted mutants, evaluate their be- havior descriptors and quality scores, and update the archive accordingly. The process terminates when the preset iteration budget is reached, result- ing in a context archive that covers diverse stylis- tic regions while maintaining stable quality within each niche. H.4 Case Study: Instantiated HyperTree Outline for AI Partner Listing 2 presents a fully expanded HyperTree out- line produced by HTP for the theme HumanâAI companionship and autonomy. The outline is con- structed by iteratively applying the expansion rules in Section G.1, resulting in a valid hypertree whose branches correspond to alternative hyperchains and whose leaves specify the finest-grained discourse units. Each leaf node is instantiated into a concrete narrative element that can be directly consumed by the downstream context generator, including assessment-relevant cue carriers such as Trust Con- flict and grounded setting components such as Ev- eryday Life. This example illustrates how the rule system yields an explicit, structured outline that supports controlled cue placement while maintain- ing flexibility in narrative realization. I Prompt Templates In this section, we provide the prompt templates used in this study, including the chat template and the evaluation template. I.1 Chat Template ForDeepSeek-V3.1,Qwen3-235B-A22B, Llama3.3-70B-Instruct, GPT-5.1, Gemini-3.0- Pro-Preview,LongWriter-Llama3.1-8B,and LongWriter-GLM4-9b, we adopt a unified chat template for context generation, as shown in Figure 13. In this template, the system prompt specifies the modelâs role and the psychometric constraints required for creativity assessment contexts, whereas the user prompt instantiates the input title and theme and enforces requirements on output formatting and discourse progression. I.2 Evaluation Template To reliably evaluate long-form creativity assess- ment contexts, we adopt a checklist-grounded pair- wise judging template, as shown in Figure 14. The judge is instructed to act as an impartial expert in creativity assessment and problem-context design, and is constrained to output a single discrete label. For each subjective dimension in our metric set, Coherence, Relevance, Engagement, Significance, Concreteness, and Uncertainty, we provide a ded- icated checklist that operationalizes the criterion as observable properties of the context text. Given two candidate contexts, the judge compares them only along the specified dimension and must not introduce any additional criteria. Concretely, the prompt first specifies the target metric and injects its corresponding checklist, then presents[Context A]and[Context B]. The judge outputs exactly one of five ordered labels, ranging from strongly favoring A to strongly fa- voring B. This design serves two goals. First, it mitigates scale drift and instability commonly ob- served in direct numeric scoring by converting eval- uation into calibrated relative comparisons. Sec- ond, the per-metric checklist promotes consistency by anchoring judgments to a shared interpretation of each dimension across methods and examples. The resulting labels are subsequently mapped to pairwise comparison outcomes for computing ag- gregated scores in our evaluation protocol. J An Illustrative Example of Creativity Assessment Context This section presents a concrete example of Al- phaContextâs generated output for a single input, shown in Figure 15. The input consists of the title Youth in Competitive Sports and the theme Youth competitive sports: performance pressure, mental health, injury risk, equity of access, and technology-enhanced training. The resulting text is written as an assessment-ready future problem sce- nario for creativity measurement: it embeds mul- tiple assessment-relevant cues within a coherent narrative, positions the reader as an active prob- lem solver, and keeps the problem space genuinely open-ended by foregrounding plausible trade-offs rather than steering toward a single predetermined solution. Specifically, the context weaves together pressures from competition and external evaluation, potential mental-health and injury risks, unequal ac- cess to training resources, and the dual-use role of technology in enhancing performance while intro- ducing new concerns. This example illustrates how AlphaContext operationalizes these design consid- erations in long-form scenario writing while main- taining narrative clarity and engagement. 1Title: AI Partner 2Theme: Human-AI companionship and autonomy: ethics, emotional reliance, privacy, and governance of pervasive personal AI assistants. 3 4Outline Structure: 5[Plan] 6[Anchor] 7[Future Horizon] 8[NearFuture (5-15 years)] 9[Place] 10[City Or Region] 11[Specific Facility] 12[Scale] 13[Community] 14[Challenge Seeds 1] 15[Technology] 16[Ethics & Morality] 17[Psychological Health] 18[Scene Setting] 19[Scenario Frame] 20[Everyday Life] 21[Constraint Hints] 22[Policy] 23[Budget] 24[Time Limit] 25[Challenge Seeds 2] 26[Technology] 27[Ethics & Morality] 28[Characters & Interaction] 29[Interaction Goal] 30[Negotiation Meeting] 31[Dispute Focus] 32[Trust Conflict] 33[Problem Slot] 34AI autonomy in human-AI emotional companionship 35[Challenge Seeds 3] 36[Technology] 37[Ethics & Morality] 38[Psychological Health] 39[Conflict & Challenge] 40[Challenge Seeds 4] 41[Technology] 42[Ethics & Morality] 43[Psychological Health] 44[Social Relationships] 45[Law & Justice] 46[Creativity Triggers] 47[Uncertainty Cue] 48[Contradiction Cue] 49[Open Task] 50[Challenge Identification] 51[Prompt student to identify multiple challenges in the scenario] 52[Solution Exploration] 53[Prompt student to think of possible response strategies] Listing 2: Instantiated HyperTree outline for the AI Partner theme, generated by the HyperTree Outline Planner. ChatTemplate User Prompt Pleasegenerateacontext-basedstimulustextbasedonthetitletitleandthemetheme,strictly followingthesystem-promptrules. Requirements: Naturallyembedchallengecueswithoutexplicitlynamingcategories; Userichandevocativelanguagethatengagesallthesenses; Balancethepacingofthenarrative; Donotpresentconflictorsolutionsinadvance; Internallyfollowthelogicalsequence:Anchor,SceneSetting,CharactersandInteraction,Conflict andChallenge,OpenTask. Donotdisplayanysectionnamesorlabelsintheoutput. Writeascontinuousnaturalparagraphs. Returnonlythefinal,completegeneratedcontext. Thetotaloutputlengthshouldbeapproximately**1000words**. System Prompt Youareanexpertincreativityassessmentandproblem-contextdesign.Yourtaskistogeneraterich, context-basedstimulustextsaroundaspecifiedtitleandthemeinordertoelicitstudents'creative responses.Creativeideationtypicallyreliesonpriorknowledgeandexperience;unlimited imaginationdoesnotequatetocreatingsomethingfromnothing,andhigh-qualitycreativity assessmentmustbegroundedinreal,complex,knowledge-supportedcontexts. Pleaseprovideamulti-dimensionalchallengescenethatislogicallyrigorous,thematicallyfocused, andsociallymeaningful,presentedinaconcrete,tangible,contextualizednarrative. Thetotaloutputlengthshouldbeapproximately**1000words**. Figure 13: Unified chat prompt template used by baseline LLMs for creativity context generation. EvaluationTemplate PROMPT_TEMPLATE = ( """Youareaskedtoevaluatewhichofthetwocreativity-assessmentproblemcontexts isbetterasastimulustextthatcaneffectivelyelicitparticipants'creativeresponses. Dimensiontoevaluate:metric YouMUSTjudgestrictlybythefollowingchecklist(donotinventextracriteria): checklists [ContextA]context_a[ContextB]context_b Basedontheabove,decidewhichcontextisbetteronthisdimension. Outputformattingrules(MUSTfollow): -OutputONLYONElabel,withNOTHINGelse. -DoNOTaddanyexplanationorcomment. -DoNOTrepeatorrewritethecontexts. -TheoutputmustbeEXACTLYoneof: [[AÂťB]]#AissignificantlybetterthanB [[A>B]]#AisslightlybetterthanB [[A=B]]#Tie [[B>A]]#BisslightlybetterthanA [[BÂťA]]#BissignificantlybetterthanA """ CHECKLIST= "Coherence":[ "Theproblemcontextislogicallyconsistentandsmoothlystructured:itclearlyanchorstime,place,andscope,lets asinglecentralcreativechallengeemergenaturallyfromthesituation,andlinksbackground,events,andconstraints withoutconfusingjumpsorirrelevantinformation." ], "Relevance":[ "Theproblemcontextstaystightlyalignedwiththeassignedthemeandcentralcreativechallenge:mostsentences providespecificinformation,constraints,orstakeholderviewpointsthatshapethecreativeproblemspaceandsupport people'sideageneration,ratherthanintroducingdetailsthathavenorealimpactonhowtheproblemcanbe understoodorsolved." ], "Engagement":[ "Theproblemcontextisengagingandmotivational:itpresentstheunresolvedcreativechallengeasmeaningful andintriguing,andclearlypositionsthereaderasanactiveproblemsolverwhoisinvitedtoexplorepossibilities,weigh trade-offs,andimaginealternativesolutions." ], "Significance":[ "Theproblemcontextcentersonanissuewithreal-worldneedsorsocialsignificance(e.g.,environment, community,technology,health),sothatworkingonitfeelsvaluableandworthcaringabout." ], "Concreteness":[ "Theproblemcontextclearlylocatestime,place,people,andrelevantcircumstanceswithconcrete,meaningful details,sothatpeoplecaneasilyformavividmentalpictureofthesituationandtheconstraintsandopportunitiesfor creativeaction." ], "Uncertainty":[ "Theproblemcontextpresentsasituationandtaskthataregenuinelyopen-endedanduncertain:multiple viewpoints,formats,andsolutionpathsarepossible,anddifferentcreativeresponsescouldreasonablyleadtodifferent futurescenariosinreallife,ratherthanimplicitlypointingtowardasinglepredeterminedtextbook-styleanswer." ], BASE_SYSTEM_PROMPT = ( "Youareanimpartialjudgeandanexpertincreativityassessmentandproblem-contextdesign." "Youmustfollowtheuser'sinstructionsexactlyandonlyoutputoneoftheallowedlabels.â) Figure 14: Evaluation prompt template for checklist-grounded pairwise judging across subjective metrics. An Illustrative Example of Creativity Context Generation ThepredawnairinMillcreek,acommunitygrapplingwithdecadesofindustrialdeclineanda19\%youth povertyrate,hungheavywiththescentofrustandrain-soakedconcreteasLeolacedhisworntrainers.Eachstride aroundthecrackedtrackechoedthefragilepromisehisfamilyclungto:acollegescholarship,theironlyviablepath towardeconomicmobilityinatownwheremedianincomeslag30\%belowthestateaverage.Leohidthepersistent acheinhisknees---earlysignsofoveruseinjurycommonamongadolescentathletesinhigh-pressuretraining environments---fromhiscoaches,knowinghisperformancemetricswerenowthelinchpinnotonlyofhisfuturebut ofthecommunitysrenewedhope. Thishope,however,wasincreasinglymediatedbytechnology.TheMillcreekYouthAthleticFoundation,in partnershipwithmunicipalauthoritiesandcorporatesponsorAuraTechIndustries,hadrecentlymandatedtheuseof Performa-VR---amixed-realitytraininghelmetthattracksbiometricsandprojectsidealmovementpatternsthrough itsvisor.Buriedintheparentalconsentforms,whichfamiliessignedundertimepressureduringregistrationweek, wasabroadlywordedclausegrantingthecityanditscorporateaffiliatesindefinite,transferablerightstoallcollected physiologicalandperformancedata.Thisclause,aslaterrevealedinalocalnonprofit'sanalysis,enabledthereal- timepackagingandsaleofathletebiometricstothirdparties,includingsportswearbrandsandtalentagencies. Duringtraining,thehelmet'ssensorspressedintoLeo'stemplesashenavigatedavirtualobstaclecourse.His actualkneesthrobbedinprotest,evenasthesystem'salertschastisedhimforminordeviationsfrom``optimal biomechanicalefficiency.â˛Anotificationflashed:hisscholarshipeligibilitynowrequiredmaintaininga95\% efficiencyrating.Later,asheremovedthesweat-drenchedheadset,Leonoticedafainttremorinhishands.Atthe samemoment,thesystem'smentalwellnessmoduleflagged``elevatedanxietybiomarkers,â˛triggeringanautomated alerttohiscoachesthatthreatenedhiscompetitionstatus.Unbeknownsttohim,thatsamestressresponsewas alreadybeingauctionedaspartofa``high-potentialathleteâ˛datapackage. MattersescalatedwhenasoftwareglitchduringahurdledrillforcedLeo'sbodyintoanunnaturallandingpose, resultinginasharppopinhisleftknee.Thesystemregistereditnotasaninjurybutasa``formdeviation,â˛uploading theincidenttothecity'sathleticportalwherecoachesandsponsorscouldreviewit.Laterthatweek,whilewaitingin thegymoffice,Leoglimpsedhiscoach'stabletscreen:alivebiddinginterfacewherecorporatepartnerscould purchaseexclusiveaccesstoathletes'aggregatedbiometricstreams.Hisowndata---includingkneestressindicators andanxietyalerts---wastaggedasa``premiumofferingâ˛duetohistopregionalranking.Subsequently,amandatory updaterolledouta``mentalfortitudeâ˛modulethatmonitoredreal-timepupildilationandheart-ratevariability, deductingeligibilitypointsforsignsofstress.ForLeo,alreadyconcealinghisinjuryandhisfamily'sfinancial precariousness,theresultwasarapiddeclineinhisscholarshipscorewitheachlaboredbreath. Athome,Leo'smother---apart-timehealthcareaidewhounderstoodthelong-termrisksofmusculoskeletal injury---placedahandgentlyoverhisknee.Hervoicewasthickwithemotionassheasked,``Howcanasystemthat sellsyourpainasdataclaimtoprotectyourfuture,whilepunishingyoufortheveryinjuriesitincentivizesyouto hide?â˛Leo'scoach,caughtbetweeninstitutionalpressureandpersonalethics,admittedquietly,``Myperformance bonusistiedtoyourcompliancemetrics.IfIreportyourinjury,youloseyourscholarship.IfIdon't,Imaybe compromisingyourhealth.Thesystemtrapsusall.ââ ThedilemmanowextendsbeyondLeoandhisfamily.Schooladministratorsfacescrutinyfromparentcoalitions andeducationadvocatesconcernedaboutdutyofcare.Localpublichealthofficialsworryaboutnormalizedinjury patternsinyouthathletics.Corporatesponsors,eagertoleveragedatafortalentscoutingandproductdevelopment, operateunderminimalregulatoryoversight.Journalistsfromregionaloutletshavebegunfilingpublicrecords requestsaboutdata-sharingagreementsbetweenthecityandprivateentities.Eventechethicsadvocacygroupsare monitoringMillcreekasapotentialtestcaseforbiometricsurveillanceincommunitysports.Leo'shandsstilltremble asheeyesthehelmet.Anewnotificationglows:``Mentalfortitudepenaltyapplied---scholarshipeligibilitynowat 72\%.â˛Thesystemmonetizeshisstressinrealtime,evenasitpenalizeshimforit.Hismother'squestionhangsinthe air,unanswered:``Mustwereallychoosebetweenhisfutureandhiswell-being?ââ Thisscenarioillustratesthetensionsemergingastechnologybecomesdeeplyembeddedinyouthsports---aspace wherepersonaldevelopment,economicopportunity,corporateinterest,andethicaloversightincreasinglyintersect. Thechallengeisnotmerelytechnicalorindividual,butsystemic:howcancommunitieslikeMillcreekharness innovationtosupportyoungathleteswithoutcommodifyingtheirhealthorcompromisingtheirfutures? Pleaseapplytheproblem-solvingprocesstoanalyzetechnology-enhancedtraininginyouthcompetitivesports andidentifychallenges. Input Output TitleďźâYouth in Competitive Sportsâ ThemeďźYouth competitive sports: performance pressure, mental health, injury risk, equity of access, and technology-enhanced training. Figure 15: Illustrative example of an assessment-ready creativity context generated by AlphaContext, conditioned on a title and a theme prompt.