Paper deep dive
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
Mete Ismayilzada, Renqing Cuomao, Daniil Yurshevich, Anna Sotnikova, Lonneke van der Plas, Antoine Bosselut
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/10/2026, 2:09:05 AM
Summary
CresOWLve is a new bilingual benchmark designed to evaluate the creative problem-solving capabilities of Large Language Models (LLMs) using real-world knowledge. Unlike existing benchmarks that rely on contrived brainteasers, CresOWLve utilizes 2,061 human-validated puzzles from the 'What? Where? When?' intellectual game. The study demonstrates that while models excel at factual retrieval, they struggle with the non-obvious creative connections required for complex problem-solving. Results indicate that 'thinking' models significantly outperform non-thinking models, suggesting that extended reasoning is critical for creative tasks.
Entities (4)
Relation Signals (2)
CresOWLve → derivedfrom → What? Where? When?
confidence 100% · CRESOWLVE is constructed from questions drawn from the renowned Russian intellectual game “What? Where? When?”
CresOWLve → evaluates → LLM
confidence 100% · we introduce CRESOWLVE, a benchmark for evaluating creative problem-solving using puzzles grounded in real-world knowledge.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insights that connect seemingly unrelated pieces of information. However, most existing benchmarks for large language models (LLMs) evaluate only specific components of this process. Moreover, many creativity-oriented benchmarks rely on artificially constructed brainteasers or contrived scenarios that do not reflect how creative problem-solving occurs in real-world settings. To address this gap, we introduce CresOWLve, a benchmark for evaluating creative problem-solving using puzzles grounded in real-world knowledge. Problems in CresOWLve require employing multiple creative thinking strategies, retrieving facts from diverse domains, and creatively combining them to arrive at a solution. Evaluating several frontier non-thinking and thinking LLMs, we show that CresOWLve remains highly challenging. Our analysis reveals a consistent performance gap: models perform substantially better on factual questions than on creative ones (up to a -17% drop). While models can often retrieve the relevant knowledge, they struggle to form the non-obvious creative connections required to integrate this information and arrive at the correct answer.
Tags
Links
- Source: https://arxiv.org/abs/2604.03374v1
- Canonical: https://arxiv.org/abs/2604.03374v1
Trouble viewing inline? Open PDF directly →
Full Text
89,924 characters extracted from source content.
Expand or collapse full text
Preprint. Under review. CRESOWLVE: Benchmarking Creative Problem-Solving Over Real-World Knowledge Mete Ismayilzada 1,2 Renqing Cuomao *1 Daniil Yurshevich *1 Anna Sotnikova 1 Lonneke van der Plas 2 Antoine Bosselut 1 1 EPFL, 2 Università della Svizzera italiana (USI) Abstract Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and com- monsense knowledge, to discover insights that connect seemingly unre- lated pieces of information. However, most existing benchmarks for large language models (LLMs) evaluate only specific components of this pro- cess. Moreover, many creativity-oriented benchmarks rely on artificially constructed brainteasers or contrived scenarios that do not reflect how creative problem-solving occurs in real-world settings. To address this gap, we introduce CRESOWLVE, a benchmark for evaluating creative problem- solving using puzzles grounded in real-world knowledge. Problems in CRESOWLVE require employing multiple creative thinking strategies, re- trieving facts from diverse domains, and creatively combining them to arrive at a solution. Evaluating several frontier non-thinking and thinking LLMs, we show that CRESOWLVE remains highly challenging. Our anal- ysis reveals a consistent performance gap: models perform substantially better on factual questions than on creative ones (up to−17% drop). While models can often retrieve the relevant knowledge, they struggle to form the non-obvious creative connections required to integrate the knowledge and arrive at the correct answer. 1 Introduction Creative problem-solving is a central component of human intelligence and plays a crucial role in scientific discovery, innovation, and everyday reasoning (Duncker & Lees, 1948). Solving problems creatively often requires a combination of several cognitive processes (Ismayilzada et al., 2024). Vertical thinking refers to systematic, step-by-step reasoning that derives conclusions through logical progression, while lateral thinking involves restructuring a problem and identifying non-obvious connections between seemingly unrelated ideas (De Bono & Zimbalist, 1970; Waks, 1997). Similarly, creativity research distinguishes between divergent thinking, which emphasizes generating multiple possible ideas or solutions, and convergent thinking, which focuses on identifying a single insightful solution by integrating different pieces of information (Guilford, 1967). Creative problem solving also relies on additional abilities such as abstraction and analogy making (Hofstadter et al., 2001), which allow individuals to transfer knowledge across domains, as well as commonsense reasoning (Davis & Marcus, 2015), which provides background knowledge about how the world works. Despite rapid advances in large language models (LLMs) (Zhao et al., 2023), most existing benchmarks evaluate only specific aspects of creative problem solving. Some benchmarks emphasize vertical reasoning through tasks such as factual reasoning (Hendrycks et al., 2020; Romanou et al., 2024), mathematical/logical inference (Hendrycks et al., 2021; Srivastava et al., 2022), and commonsense reasoning (Sakaguchi et al., 2021; Talmor et al., 2019; Lin * Equal contribution 1 arXiv:2604.03374v1 [cs.CL] 3 Apr 2026 Preprint. Under review. The Turkic poet Saif Sarai in the poem “Sukhail and Guldursun” metaphorically describes a girl circling around her illustrious beloved. Whom, according to the translator, did Sarai precede by about one and a half centuries? [Nicolaus] Copernicus 14th century poet Heliocentrism metaphor discovered 14 + 1.5 = 15.5th century lived in Earth revolving around Sun Figure 1: An example from CRESOWLVE annotated with the real-world knowledge and creative thinking strategy. et al., 2021), while others focus on lateral thinking using brainteasers and situational puzzles (Han et al., 2025; Chen et al., 2024; Huang et al., 2024; Jiang et al., 2023). Other works have proposed benchmarks to measure convergent/divergent thinking using psychometric tests (Stevenson et al., 2022; Góes et al., 2023; Bellemare-Pepin et al., 2024) or puzzles and games (Tian et al., 2024; Alavi Naeini et al., 2023). Although these benchmarks capture important individual capabilities, real-world creative problem-solving typically requires a combination of these skills. Furthermore, many existing evaluations rely on artificially constructed brainteasers (e.g. “What type of cheese is made backwards?”) or contrived situational scenarios (e.g. “Two men are beating each other up and both of them suddenly fall to the floor. How?”). In practice, however, creative insights often emerge from retrieving relevant knowledge across diverse domains and reasoning about it in novel ways. As a result, these benchmarks provide only a partial picture of the real-world creative problem-solving abilities of models. To address this gap, we introduce CRESOWLVE, a benchmark designed to evaluate cre- ative problem-solving in LLMs through single-answer puzzles that require the combination of multiple cognitive abilities, such as vertical/lateral, convergent/divergent thinking, analogy-making, and are grounded in real-world knowledge (Appendix Table 3 provides a full comparison to past relevant benchmarks and Figure 1 illustrates an example puzzle from our benchmark). CRESOWLVE is constructed from questions drawn from the renowned Russian intellectual game “What? Where? When?” 1 , in which expert human participants solve carefully crafted problems that require not only broad world knowledge but also creative insight to combine disparate facts in non-obvious ways. To ensure accessibility and relevance, we design a multi-stage benchmark construction pipeline that filters unsuitable and non-creative questions and translates the remaining puzzles into English with manual validation. The resulting dataset provides a diverse and high-quality benchmark for eval- uating creative problem-solving grounded in real-world knowledge 2 . We summarize our main contributions as follows: •We present CRESOWLVE, a bilingual benchmark for creative problem-solving grounded in real-world knowledge and solvable by human experts. CRESOWLVE spans a diverse range of knowledge and creative domains, varies in difficulty, 1 https://en.wikipedia.org/wiki/What%3F_Where%3F_When%3F 2 We release the benchmark at https://huggingface.co/datasets/mismayil/cresowlve 2 Preprint. Under review. requires multiple creative thinking strategies and is manually validated to ensure quality. •We evaluate several frontier open-weight and proprietary LLMs varying in their size and reasoning mode, on CRESOWLVE and demonstrate that, despite recent advances, the benchmark remains highly challenging, revealing substantial gaps in model performance on creative reasoning over real-world knowledge. Notably, thinking models substantially outperform their non-thinking counterparts, demon- strating that extended reasoning is a key enabler for creative problem-solving. •We conduct a comprehensive analysis of model performance across multiple di- mensions, including question difficulty, domains of knowledge and creativity, and show that models consistently underperform on creative questions relative to fac- tual ones. Our error analysis also shows that while models can often retrieve the relevant knowledge, they frequently fail to make the creative connections among facts necessary to arrive at the correct answer. 2 Related Work Creativity Evaluation Evaluating creativity is inherently challenging due to its subjec- tive nature, and traditional assessments often rely on human judgment using techniques such as the Consensual Assessment Technique (Ismayilzada et al., 2024; Amabile, 1983) or psychometric tests (Stevenson et al., 2022; Góes et al., 2023; Guilford, 1967; Mednick, 1962). To complement these approaches, several computational metrics and corresponding benchmarks have been proposed that capture different dimensions of creativity, including novelty (Zhang et al., 2025; Lu et al., 2024; Organisciak et al., 2023; Johnson et al., 2023), diversity (Padmakumar & He, 2023), surprise (Bunescu & Uduehi, 2022; Karampiperis et al., 2014; Itti & Baldi, 2009), and quality (Franceschelli & Musolesi, 2025; 2022). Other works have introduced multi-task creativity evaluation datasets (Ismayilzada et al., 2025; Hou et al., 2025; Xue et al., 2025) where problems don’t have a single correct answer, but rather many subjective answers, and the evaluation focuses on measuring different dimensions of output creativity. In creative problem-solving, however, there are often single (or a few) correct answers, and the emphasis is on the type of creative thinking process involved to get to the final answer. Our work similarly targets creative problem solving with a single correct answer, hence our analysis focuses more on the evaluation of different creative thinking abilities employed by the model to solve our benchmark puzzles. Creative Problem-Solving Benchmarks While most works have focused on measuring vertical thinking capabilities of LLMs through tasks such as factual reasoning (Hendrycks et al., 2020; Romanou et al., 2024), mathematical/logical inference (Hendrycks et al., 2021; Srivastava et al., 2022), and commonsense reasoning (Sakaguchi et al., 2021; Talmor et al., 2019; Bisk et al., 2020; Ismayilzada et al., 2023; Lin et al., 2021), several benchmarks have also been proposed to evaluate other aspects of problem-solving that requires more creativity (Ismayilzada et al., 2024). More specifically, past work has evaluated lateral thinking using mostly brainteasers and situational puzzles (Han et al., 2025; Chen et al., 2024; Huang et al., 2024; Jiang et al., 2023; Kraaijveld et al., 2025; Todd et al., 2024), convergent/divergent thinking using psychometric tests (Stevenson et al., 2022; Góes et al., 2023; Bellemare-Pepin et al., 2024) or puzzles and games (Tian et al., 2024; Alavi Naeini et al., 2023; Wadhwa et al., 2026) and abstraction/analogy-making through synthetic puzzles (Ahrabian et al., 2024; Chollet, 2019; Moskvichev et al., 2023; Lewis & Mitchell, 2024). Other works have focused on specific domains such as coding (Lu et al., 2025), and mathematics (Ye et al., 2025; Sun et al., 2025). Our benchmark on the other hand requires models to employ multiple creative thinking abilities and reason creatively over real-world knowledge drawn from diverse domains. We also note that while the questions from the “What?Where?When” game has been used for LLM evaluation in the past by Lifar et al. (2024), our work significantly differs in several aspects: 1) Lifar et al. (2024) only considers questions with more factuality and shorter reasoning chains while we focus on more creative questions 2) Lifar et al. (2024) does not perform any data filtering to remove unanswerable or Russian-culture specific questions 3) Lifar et al. (2024) evaluates only one LLM (namely,LlaMa3-405B) and provides 3 Preprint. Under review. limited analysis while we benchmark several frontier LLMs and provide extensive analysis on question difficulty, reasoning types and error categories. 4) Lifar et al. (2024) considers Russian-only evaluation, while we prepare and evaluate models on a human-validated English version of the dataset as well. 3CRESOWLVE 3.1 Benchmark Construction Data Collection We collect 3, 789 questions from the public database atdb.chgk.info, which contains questions from the popular Russian intellectual game “What?Where?When?” spanning over 50 years. Questions in this game have been manually crafted by humans and often require a combination of skills such as logical thinking, intuition, and creative insight. Each question is annotated with a short answer, an answer explanation, and a difficulty rating (1 to 5) manually assigned by the game organizers. To ensure diversity in benchmark difficulty, we collect≈700 questions per difficulty rating. We perform several filtering and annotation stages usingGPT-4o(Hurst et al., 2024) due to its strong performance (He et al., 2024; Gilardi et al., 2023). Prompts can be found in Appendix Tables 5, 7, 6, 8, 9 and 11. Example questions for each stage can be found in Appendix Table 4. Data Filtering Since the original game is in Russian and also in-person interactive, some questions rely heavily on linguistic and cultural knowledge specific to Russia, and some require physical inspection of external material, such as images or handouts. Hence, these questions are either extremely hard to answer or unanswerable completely. Therefore, we perform several filtering steps to ensure the validity and wide accessibility of the questions. We first filter out unanswerable questions by annotating each question on whether they require physical external materials to be solved. This step removes 295 samples. Next, we filter out questions that depend heavily on knowing specific facts rooted deeply in the Russian language and culture, rendering them extremely hard to answer or untranslatable into English. This step removes 799 samples. TranslationOur goal in this work is to design a benchmark that is relevant and accessible for measuring creative problem-solving in frontier LLMs. However, since most LLMs are more proficient reasoners in English than in other languages, we also prepare an English version of our benchmark by translating each question-answer pair with all its metadata. Human Validation While LLMs have recently become remarkably effective in data an- notation tasks, they can still exhibit hallucination and reasoning failures (Tan et al., 2024). Therefore, we conduct a final human validation on the entire benchmark (including the factual reasoning questions) to ensure high quality. More specifically, three authors of this paper reviewed all remaining questions from the last filtering step to validate that they are answerable, not Russian-specific, and that the translations are correct. This step further removed 282 samples from the benchmark, leaving 2, 413 samples for final evaluation. Creative vs. Factual Reasoning As noted by Lifar et al. (2024), some questions can be answered with purely factual retrieval and reasoning involving no to minimal creative leap- of-thought. To further distinguish this dimension, we automatically categorize questions as either factual or creative, where creative questions require the model to combine and apply knowledge in ways that go beyond direct retrieval. This annotation yields 352 factual and 2, 061 creative samples. We later use this factual subset to analyze the performance on creative vs. factual reasoning questions. 3.2 Final Benchmark Our final benchmark contains 2, 061 samples (creative subset) both in Russian and English (referred to as CRESOWLVE-Ru and CRESOWLVE-En respectively). Data statistics about the benchmark can be found in Appendix Table 2. 4 Preprint. Under review. Knowledge Domains As noted earlier, one of the challenges of this benchmark is to creatively reason over real-world knowledge across domains. To visualize the domain diversity, we automatically annotate each benchmark sample with the knowledge domains required to solve it. The initial annotation yields 541 domains, which are too fine-grained and contain substantial overlap. We then consolidate them into broader subject domains, resulting in 34 coarse-grained, largely non-overlapping categories. Figure 2a illustrates the coarse domain breakdown. We can see that while the benchmark covers a wide range of topics, notably, Literature, History, Film & Media Studies, and Languages & Linguistics are dominating subjects. We also note that each question is annotated with multiple topics, and most questions involve at least two to four topics (See Appendix Figure 6a). Visual Arts 190 (9.2%) Literature 1065 (51.7%) Music 189 (9.2%) Military 71 (3.4%) Languages & Linguistics 375 (18.2%) Religious Studies 271 (13.1%) Political Science 180 (8.7%) History 845 (41.0%) Sports 247 (12.0%) Art & Culture 38 (1.8%) Anthropology 247 (12.0%) Biology 215 (10.4%) Physics 77 (3.7%) Human Geography 291 (14.1%) Film & Media Studies 402 (19.5%) Design & Arch. 63 (3.1%) Medicine & Health 62 (3.0%) Astronomy 77 (3.7%) Education 19 (0.9%) Psychology 127 (6.2%) Law & Criminology 39 (1.9%) Sociology 121 (5.9%) Communication 20 (1.0%) Economics 56 (2.7%) Performing Arts 156 (7.6%) Engineering & Technology 186 (9.0%) Business Studies 77 (3.7%) Chemistry 46 (2.2%) Home & Daily Life 171 (8.3%) Philosophy 64 (3.1%) Other Sciences 36 (1.7%) Mathematics 58 (2.8%) Earth & Environment 80 (3.9%) Archaeology 12 (0.6%) Total 2,061 (a) Distribution of knowledge domains. metaphor 207 (8.6%) lateral thinking 1740 (72.1%) analogy 830 (34.4%) abstraction 357 (14.8%) pun 225 (9.3%) simile 8 (0.3%) joke 226 (9.4%) neologism 36 (1.5%) sarcasm 34 (1.4%) commonsense reasoning 143 (5.9%) idiom 79 (3.3%) poem 83 (3.4%) divergent thinking 27 (1.1%) proverb 31 (1.3%) compositionality 13 (0.5%) other 17 (0.7%) (b) Distribution of creative domains. english 1186 (57.5%) french 210 (10.2%) hebrew 21 (1.0%) greek 134 (6.5%) german 160 (7.8%) dutch 24 (1.2%) spanish 75 (3.6%) russian 796 (38.6%) japanese 54 (2.6%) latin 85 (4.1%) american 75 (3.6%) swedish 21 (1.0%) arabic 26 (1.3%) italian 119 (5.8%) ukrainian 19 (0.9%) chinese 21 (1.0%) polish 34 (1.6%) scottish 13 (0.6%) portuguese 14 (0.7%) norwegian 17 (0.8%) danish 16 (0.8%) egyptian 12 (0.6%) czech 13 (0.6%) roman 23 (1.1%) indian 16 (0.8%) other 326 (15.8%) (c) Distribution of cultures/demographics. Figure 2: Diversity of real-world knowledge, creative language, and cultures. Creative Language & Thinking In addition to the knowledge domains, our benchmark questions also require reasoning about creative language and employing multiple creative thinking strategies such as lateral thinking, abstraction, and analogy-making. To quantify them, we perform an automatic annotation and report the breakdown in Figure 2b. We see that most of the questions require lateral thinking, an ability to identify non-obvious associa- tions, and a substantial number of questions involve abstraction and analogy-making. Many questions also require reasoning about jokes, puns, or metaphors, as well as commonsense knowledge, and most questions involve two creative domains (Appendix Figure 6b). Cultures & Demographics While the original game is in the Russian language, its ques- tions often involve knowledge about entities and people from other cultures, too. To quantify the diversity of cultures and demographics in our benchmark, we automatically annotate each sample in the benchmark with the cultures involved in solving the question. We report the breakdown in Figure 2c. We note that while English and Russian cultures dominate, our benchmark also contains a substantial number of questions on other cultures, such as French, German, Italian, and Greek. Similar to knowledge and creative domains, most questions often involve knowledge about more than one culture (Appendix Figure 6c). 5 Preprint. Under review. CRESOWLVE-EnCRESOWLVE-Ru ModelThinking EffortExact MatchLLM JudgeExact MatchLLM Judge OLMO-2-32B-INSTRUCT-2.095.290.441.31 C4AI-COMMAND-A-4.469.073.747.08 LLAMA-3.3-70B-INSTRUCT-5.0510.143.647.04 MISTRAL-LARGE-3-675B-INSTRUCT-11.2617.6115.4822.22 QWEN3-235B-A22B-INSTRUCT-13.0520.8213.5420.14 GPT-4.1-MINI-7.3313.784.858.83 GPT-4.1-13.4424.9917.0326.06 QWEN3.5-397B-A17Badaptive19.3631.2526.1535.71 QWEN3-235B-A22B-THINKINGadaptive13.4920.8213.3421.11 DEEPSEEK-V3.2adaptive15.3325.4710.4826.44 GLM-5adaptive18.9237.1728.8744.25 GPT-5.4none9.4120.9615.3826.54 GPT-5.4medium17.8153.4726.8868.22 GEMINI-3-FLASHminimal24.9941.1039.9854.34 GEMINI-3-FLASHmedium35.0852.1651.8267.25 GEMINI-3.1-PROlow46.4367.5464.1080.49 GEMINI-3.1-PROmedium49.1072.9765.7983.60 GEMINI-3.1-PROhigh51.2976.0367.7885.74 Non-Thinking Models Thinking Models Table 1: Overall performance results. 4 Experimental Setup Models We assess a broad set of LLMs on our benchmark, varying in reasoning mode, architecture, size, and training paradigm. Since our benchmark requires complex reason- ing, we particularly distinguish between non-thinking models, which generate responses directly without explicit intermediate reasoning (unless instructed to via Chain-of-Thought prompting (Wei et al., 2022)) and thinking models, which are trained to explicitly think with a certain amount of effort before producing a final answer. We consider the following non-thinking models:GPT-4.1-mini,GPT-4.1(Achiam et al., 2023),OLMo-2-32B-Instruct (OLMo et al., 2024),Qwen3-235B-A22B-Instruct(Yang et al., 2025),C4AI-Command-A(Cohere et al., 2025),Mistral-Large-3-675B-Instruct(Mistral, 2025) andLlama-3.3-70B-Instruct (Grattafiori et al., 2024) and following thinking models:Qwen3.5-397B-A17B(Qwen Team, 2026),Qwen3-235B-A22B-Thinking(Yang et al., 2025),DeepSeek-V3.2(Liu et al., 2025), GLM-5(Zeng et al., 2026),GPT-5.4(OpenAI, 2026),Gemini-3-Flash(Google, 2025), and Gemini-3.1-Pro (Google, 2026). Task & Evaluation Methods We formulate our problem as an open-ended question answering task. Since our benchmark requires substantial reasoning over real-world knowl- edge, we evaluate non-thinking models using Chain-of-thought prompting and thinking models using a standard prompt with varying levels of thinking effort. We evaluate all models using their default recommended decoding setup. Evaluation prompts can be found in Appendix Table 12. Evaluation MetricsWe measure the model performance using two metrics: Exact Match Accuracy, where we match the reference answer with the model response after light nor- malization (i.e., lowercase, remove punctuation, and unicode normalization for Russian), and LLM-as-a-judge Accuracy, where we employGPT-4oto judge the correctness of the model response given the reference answer (Gu et al., 2024). To ensure high-quality LLM judgment, we instruct it to ignore typos, articles, or formatting differences and provide it with additional context, including the human answer explanations and other acceptable answers, if any. LLM-as-a-judge prompt can be found in Appendix Table 13. 5 Results Overall Performance Table 1 reports overall performance on CRESOWLVE-En and CRE- SOWLVE-Ru. Results vary widely across models and metrics, ranging from below 10% 6 Preprint. Under review. Very SimpleSimpleMediumHardVery Hard Difficulty 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy CresOWLve-En Very SimpleSimpleMediumHardVery Hard Difficulty CresOWLve-Ru Model Llama-3.3-70B-Instruct Mistral-Large-3-675B-Instruct GPT-4.1 Qwen3-235B-A22B-Thinking (adaptive) DeepSeek-V3.2 (adaptive) GPT-5.4 (medium) Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) Figure 3: LLM Judge results by difficulty level (Exact Match, Appendix Figure 13). to above 80%. Non-thinking models all remain under 30% accuracy, withGPT-4.1being the strongest in this group. Thinking models generally outperform their non-thinking counterparts, and greater thinking effort consistently yields higher accuracy. This effect is particularly pronounced forGPT-5.4, whose performance nearly doubles when moving from no thinking to medium effort. Among thinking models, closed-source models lead overall; theGeminiseries in particular achieves top performance even at minimal thinking effort, surpassing all open-weight models. Regarding cross-lingual transfer, we observe mixed trends. Most non-thinking models suffer a notable performance drop on the Russian benchmark (up to−5% LLM Judge and −3% Exact Match), withQwen3-235B-A22B-InstructandGPT-4.1being the only exceptions. Thinking models, however, show the opposite trend: closed-source thinking models in particular achieve substantially higher scores on the Russian benchmark (up to+15% LLM Judge and+18% Exact Match). This surprising advantage on CRESOWLVE-Ru may reflect stronger multilingual reasoning capabilities at higher levels of thinking, or potential data contamination (for which we find some evidence in Appendix Section A). Regardless, the consistently low performance of open-weight (and some proprietary) models across both languages confirms that CRESOWLVE poses a genuine challenge and serves as a robust testbed for benchmarking creative problem-solving in LLMs. Performance by difficulty level Figure 3 illustrates model performance stratified by difficulty level for both CRESOWLVE-En and CRESOWLVE-Ru under LLM Judge evaluation. As expected, accuracy decreases monotonically with difficulty for all models, confirming that the benchmark difficulty levels capture the underlying complexity of the questions. Across both benchmarks, the performance degradation from Very Simple to Very Hard is substantial — even the strongest thinking models lose over 20% accuracy. This highlights a key challenge posed by CRESOWLVE: while average accuracy figures may appear moderate, the benchmark contains a non-trivial proportion of questions that remain genuinely difficult even for state-of-the-art models, underscoring the headroom for future progress. Performance by domain, creativity, and cultures Detailed performance breakdowns by knowledge and creative domains, and culture groups are provided in Appendix Figures 7, 8, 9 and 10, 11, 12 for the CRESOWLVE-En and CRESOWLVE-Ru benchmarks respectively. Across benchmarks, performance is largely uniform across all three stratifications, though several notable trends emerge. Across domains, most models tend to perform better on Earth & Environmental Sciences questions, withQwen3-235B-A22B-Thinkingshowing a particular advantage on Astronomy andDeepSeek-V3.2on Mathematics. Regarding creative language concepts, models generally perform worse on questions involving poems and metaphors, suggesting that figurative and poetic language remains a particular challenge. On the culture dimension,Gemini-3.1-Properforms notably well on questions involving American cultural references, GPT-4.1 on Japanese, and DeepSeek-V3.2 on Latin. 7 Preprint. Under review. factualcreative Category 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy CresOWLve-En factualcreative Category CresOWLve-Ru Model Llama-3.3-70B-Instruct Mistral-Large-3-675B-Instruct GPT-4.1 Qwen3-235B-A22B-Thinking (adaptive) DeepSeek-V3.2 (adaptive) GPT-5.4 (medium) Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) Figure 4: LLM Judge results by reasoning category (Exact Match, Appendix Figure 14). Human Performance We do not conduct a formal human evaluation, as the puzzles in CRESOWLVE are specifically designed to be solved by expert players of the original game and indeed, all questions in our benchmark have been correctly answered in past game episodes. This provides an implicit upper bound: the benchmark is solvable, but requires a combination of broad knowledge and creative reasoning. Interestingly, human experts in this setting operate under conditions that are in some respects more constrained than LLMs: they have been exposed to far less data and cannot rely on exhaustive factual recall. Instead, they compensate through intuition and creativity, using partial knowledge to make imaginative leaps toward the correct answer. LLMs, by contrast, have been trained on virtually all publicly available text and should in principle, have access to all the factual knowledge required to solve every question. The fact that they nonetheless fall short seems to highlight that the bottleneck is not knowledge retrieval, but the creative reasoning needed to connect that knowledge, precisely the capability that CRESOWLVE is designed to probe. 6 Analysis Creative vs. Factual ReasoningTo investigate whether models struggle more with creative reasoning than factual retrieval, we evaluate on the factual subset of CRESOWLVE as discussed in Section 3.1. Figure 4 reveals a consistent and substantial performance drop from factual to creative questions across all models and languages. On CRESOWLVE- En, performance drop ranges from−6.08% forGemini-3.1-Pro (high)to−17.21% for Mistral-Large-3-675B-Instruct. On CRESOWLVE-Ru, the same trend holds, though the absolute drop is smaller forGemini-3.1-Promodels (−2.14% and−2.73%). For a fair comparison, we match the sample sizes across both categories and perform bootstrap sampling with 1000 iterations. In addition, we balance the number of samples across difficulty levels for both creative and factual questions (Appendix Figure 15). These results confirm that creative questions pose a fundamentally harder challenge than factual ones. Difficulty vs. ComplexityTo better understand the sources of question difficulty in CRE- SOWLVE, we investigate whether difficulty correlates with measurable complexity features of the questions. Specifically, we consider three proxy measures of question complexity: the number of domains involved, the maximum semantic distance between domains for a given question, and the number of atomic facts required to resolve it. The latter is esti- mated by promptingGPT-4oto decompose each question into a list of constituent factual sub-questions (prompt provided in Appendix Table 10). Despite the intuitive appeal of these features as predictors of difficulty, we find no significant correlation between any of them and either the assigned difficulty level or model performance (Appendix Figures 16, 17). These findings strongly suggest that the difficulty of questions in CRESOWLVE doesn’t stem from the number of involved domains, the challenge of bridging semantically distant 8 Preprint. Under review. Missing creative connection Clue misinterpretation Incorrect concept anchoring Wrong reference Overthinking Wrong hypothesis Hallucination Error Category GPT-4.1 GPT-5.4 (medium) DeepSeek-V3.2 (adaptive) Gemini-3-Flash (medium) Gemini-3.1-Pro (high) Model 47.4%27.5%13.1%6.9%2.0%2.3%0.5% 42.4%18.2%25.4%5.8%3.5%4.4%0.2% 48.2%22.1%13.5%8.2%2.6%4.8%0.5% 17.8%32.4%6.8%32.2%4.4%4.8%1.5% 29.4%32.7%16.5%6.6%8.3%5.4%0.5% Error Category Distribution by Model 0.0 0.1 0.2 0.3 0.4 Percentage Figure 5: Error category distribution for best performing models. domains, or the volume of facts one needs to recall, but rather from the need to connect knowledge in creative ways. Error Analysis We conduct a manual analysis of errors produced by Gemini-3-Flash (medium). From a total of 1,087 error examples, we randomly sample 150 instances with the model’s reasoning traces and obtain annotations from three authors of the paper, who are also domain experts. Following an iterative error analysis process, we identify seven error categories that are most prominent in our data: •Missing creative connection: The model retrieves relevant knowledge, but fails to recognize the intended associative or metaphorical link between clues, preventing it from making the key conceptual leap needed to reach the correct answer. •Overthinking: The model identifies the correct concept during reasoning, but later replaces it with another answer due to reinterpretation or over-generalization. •Hallucination or unsupported fabrication: The model invents facts, explanations, or source details that are not supported by the question or reliable knowledge. •Incorrect concept anchoring: The model locks onto an incorrect concept early and builds reasoning around it, either through associative drift or by reinterpreting the clues to fit that concept. •Wrong hypothesis: The model infers an incorrect rule or shared property from the clues and applies it consistently to produce an answer. •Wrong reference: The model retrieves and reasons about an incorrect work, person, event, or source that superficially matches the clues but is not the intended reference. •Clue misinterpretation: The model identifies the general source, context, or line of reasoning, but selects the wrong specific element required by the clue. For each category, we provide representative examples in Appendix B. We then use these definitions to automatically label all wrong examples at scale. The resulting error distribu- tion for the best performing models is shown in Figure 5. The most frequent error categories across models are missing creative connection and clue misinterpretation, further confirming our hypothesis about the difficulty of questions. A notable exception isGemini-3.1-Flash, which shows significantly fewer missing creative connection errors but a substantially higher rate of wrong reference errors.GPT-5.4stands out for its disproportionately high rate of incorrect concept anchoring. Finally, low hallucination and wrong reference rates across most models confirm that failures stem not from missing knowledge but from an inability to make the creative connections between facts necessary to reach the correct answer. 9 Preprint. Under review. 7 Conclusion We introduced CRESOWLVE, a bilingual benchmark for evaluating creative problem-solving in LLMs, constructed from real-world puzzles sourced from the What? Where? When? intellectual game. Our evaluation of several frontier LLMs reveals that the benchmark remains highly challenging, with even the strong thinking models falling considerably short on creative questions. Analysis shows that this difficulty does not stem from surface-level complexity features, but from the need to forge creative connections between knowledge from different domains. We hope CRESOWLVE serves as a challenging testbed to drive progress in creative reasoning. Ethics Statement This work introduces CRESOWLVE, a benchmark for evaluating creative problem-solving in large language models using puzzles derived from the intellectual game “What? Where? When?”. While the goal is to advance the evaluation of creative reasoning, several ethical considerations arise. The benchmark is constructed from publicly available questions created by human authors over several decades. These questions reflect the intellectual contributions of their original creators, and we do not claim authorship of the underlying content. The dataset is used solely for research purposes, and proper attribution should be maintained where applicable. Because the questions are designed to test non-obvious associative and creative connections, they may reflect the cultural context, assumptions, and potential biases of their authors. In particular, as the source material originates in Russian, some questions may encode culturally specific knowledge or perspectives. To mitigate this, we annotate questions based on their regional content; however, one should account for potential biases. Given that the source material is publicly available and widely distributed, there is a risk that some models may have been exposed to similar questions during training. As discussed in our analysis, this potential data contamination may affect performance and should be considered when interpreting results. Finally, although the benchmark is designed to measure creative problem-solving, creativity is inherently difficult to define and evaluate. Our categorization into factual and creative questions, as well as the proposed error taxonomy, involves subjective judgments and may not capture all aspects of creative reasoning. Acknowledgements We thank the members of the EPFL NLP for their feedback on the project and the pa- per manuscript. We gratefully acknowledge the support of the Swiss National Science Foundation (No. 215390), the European Research Council (Starting grant no. 101222478, RESPECT-LM), the AI2050 program at Schmidt Sciences (Grant #G-25-69783), Sony Group Corporation, and the Swiss National Supercomputing Center (CSCS) in the form of an infrastructure engineering and development project. LP also gratefully acknowledges the support of the Swiss National Science Foundation (grant 205121_207437: C - LING). References Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. Kian Ahrabian, Zhivar Sourati, Kexuan Sun, Jiarui Zhang, Yifan Jiang, Fred Morstatter, and Jay Pujara. The curious case of nonverbal abstract reasoning with multi-modal large language models. arXiv preprint arXiv:2401.12117, 2024. Saeid Alavi Naeini, Raeid Saqur, Mozhgan Saeidi, John Giorgi, and Babak Taati. Large language models are fixated by red herrings: Exploring creative problem solving and einstellung effect using the only connect wall dataset. Advances in Neural Information Processing Systems, 36:5631–5652, 2023. 10 Preprint. Under review. Teresa M Amabile. The social psychology of creativity: a componential conceptualization. Journal of personality and social psychology, 45(2):357, 1983. Antoine Bellemare-Pepin, François Lespinasse, Philipp Thölke, Yann Harel, Kory Mathew- son, Jay A Olson, Yoshua Bengio, and Karim Jerbi. Divergent creativity in humans and large language models. arXiv preprint arXiv:2405.13012, 2024. Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, p. 7432–7439, 2020. Razvan C. Bunescu and Oseremen O. Uduehi. Distribution-based measures of surprise for creative language: Experiments with humor and metaphor. In Debanjan Ghosh, Beata Beigman Klebanov, Smaranda Muresan, Anna Feldman, Soujanya Poria, and Tuhin Chakrabarty (eds.), Proceedings of the 3rd Workshop on Figurative Language Processing (FLP), p. 68–78, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.flp-1.10. URLhttps://aclanthology. org/2022.flp-1.10/. Qi Chen, Bowen Zhang, Gang Wang, and Qi Wu. Weak-eval-strong: Evaluating and eliciting lateral thinking of llms with situation puzzles. Advances in Neural Information Processing Systems, 37:79642–79665, 2024. François Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019. Team Cohere, Arash Ahmadian, Marwan Ahmed, Jay Alammar, Milad Alizadeh, Yazeed Alnumay, Sophia Althammer, Arkady Arkhangorodsky, Viraat Aryabumi, Dennis Au- miller, et al. Command a: An enterprise-ready large language model. arXiv preprint arXiv:2504.00698, 2025. Ernest Davis and Gary Marcus. Commonsense reasoning and commonsense knowledge in artificial intelligence. Communications of the ACM, 58(9):92–103, 2015. Edward De Bono and Efrem Zimbalist. Lateral thinking. Penguin London, 1970. K. Duncker and L.S. Lees. On Problem-solving. Psychological monographs. American Psychological Ass., 1948. URL https://books.google.ch/books?id=g888tAEACAAJ. Giorgio Franceschelli and Mirco Musolesi. Deepcreativity: measuring creativity with deep learning techniques. Intelligenza Artificiale, 16(2):151–163, 2022. Giorgio Franceschelli and Mirco Musolesi. Thinking outside the (gray) box: A context- based score for assessing value and originality in neural text generation. arXiv preprint arXiv:2502.13207, 2025. Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. Chatgpt outperforms crowd work- ers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30): e2305016120, 2023. Luis Fabricio Góes, Marco Volpe, Piotr Sawicki, Marek Grses, and Jacob Watson. Pushing gpt’s creativity to its limits: Alternative uses and torrance tests. 2023. Shahriar Golchin and Mihai Surdeanu. Data contamination quiz: A tool to detect and estimate contamination in large language models. Transactions of the Association for Compu- tational Linguistics, 13:809–830, 2025. Team Google. A new era of intelligence with gemini 3, 2025. URLhttps://blog.google/ products-and-platforms/products/gemini/gemini-3/#note-from-ceo. Team Google. Gemini 3.1 pro: A smarter model for your most complex tasks, 2026. URLhttps://blog.google/innovation-and-ai/models-and-research/gemini-models/ gemini-3-1-pro/. 11 Preprint. Under review. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge. The Innovation, 2024. Joy Paul Guilford. The nature of human intelligence. 1967. Simeng Han, Howard Dai, Stephen Xia, Grant Zhang, Chen Liu, Lichang Chen, Hoang Huy Nguyen, Hongyuan Mei, Jiayuan Mao, and R Thomas McCoy. Creativity or brute force? using brainteasers as a window into the problem-solving abilities of large language models. arXiv preprint arXiv:2505.10844, 2025. Zeyu He, Chieh-Yang Huang, Chien-Kuang Cornelia Ding, Shaurya Rohatgi, and Ting- Hao Kenneth Huang. If in a crowdsourced data annotation pipeline, a gpt-4. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400703300. doi: 10.1145/ 3613904.3642834. URL https://doi.org/10.1145/3613904.3642834. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021. Douglas R Hofstadter et al. Analogy as the core of cognition. The analogical mind: Perspectives from cognitive science, p. 499–538, 2001. Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, et al. Creativityprism: A holistic benchmark for large language model creativity. arXiv preprint arXiv:2510.20091, 2025. Shulin Huang, Shirong Ma, Yinghui Li, Mengzuo Huang, Wuhe Zou, Weidong Zhang, and Haitao Zheng. Lateval: An interactive llms evaluation benchmark with incomplete information from lateral thinking puzzles. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), p. 10186–10197, 2024. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, and Antoine Bosselut. Crow: Benchmarking commonsense reasoning in real-world tasks.arXiv preprint arXiv:2310.15239, 2023. Mete Ismayilzada, Debjit Paul, Antoine Bosselut, and Lonneke van der Plas. Creativity in ai: Progresses and challenges. arXiv preprint arXiv:2410.17218, 2024. Mete Ismayilzada, Antonio Laverghetta, Simone A Luchini, RN Patel, Antoine Bosselut, Lonneke Van Der Plas, and Roger E Beaty. Creative preference optimization. Findings of the Association for Computational Linguistics: EMNLP 2025, p. 9580–9609, 2025. Laurent Itti and Pierre Baldi. Bayesian surprise attracts human attention. Vision research, 49 (10):1295–1306, 2009. Yifan Jiang, Filip Ilievski, Kaixin Ma, and Zhivar Sourati. Brainteaser: Lateral thinking puzzles for large language models. arXiv preprint arXiv:2310.05057, 2023. 12 Preprint. Under review. Dan R Johnson, James C Kaufman, Brendan S Baker, John D Patterson, Baptiste Barbot, Adam E Green, Janet van Hell, Evan Kennedy, Grace F Sullivan, Christa L Taylor, et al. Divergent semantic integration (dsi): Extracting creativity from narratives with distribu- tional semantic modeling. Behavior Research Methods, 55(7):3726–3759, 2023. Pythagoras Karampiperis, Antonis Koukourikos, and Evangelia Koliopoulou. Towards machines for measuring creativity: The use of computational tools in storytelling activities. In 2014 IEEE 14th International Conference on Advanced Learning Technologies, p. 508–512. IEEE, 2014. Koen Kraaijveld, Yifan Jiang, Kaixin Ma, and Filip Ilievski. Columbus: Evaluating cogni- tive lateral understanding through multiple-choice rebuses. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, p. 4410–4418, 2025. Martha Lewis and Melanie Mitchell. Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models. arXiv preprint arXiv:2402.08955, 2024. Mikhail Lifar, Bogdan Protsenko, Daniil Kupriianenko, Nazar Chubkov, Kulaev Kirill Dmitrievich, Alexander Guda, and Irina Piontkovskaya. Llama meets cheburashka: impact of cultural background for llm quiz reasoning. In Language Gamification-NeurIPS 2024 Workshop, 2024. Bill Yuchen Lin, Ziyi Wu, Yichi Yang, Dong-Ho Lee, and Xiang Ren. Riddlesense: Reasoning about riddle questions featuring linguistic creativity and commonsense knowledge. arXiv preprint arXiv:2101.00376, 2021. Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556, 2025. Ximing Lu, Melanie Sclar, Skyler Hallinan, Niloofar Mireshghallah, Jiacheng Liu, Seungju Han, Allyson Ettinger, Liwei Jiang, Khyathi Chandu, Nouha Dziri, et al. Ai as humanity’s salieri: Quantifying linguistic creativity of language models via systematic attribution of machine text against web text. arXiv preprint arXiv:2410.04265, 2024. Yining Lu, Dixuan Wang, Tianjian Li, Dongwei Jiang, Sanjeev Khudanpur, Meng Jiang, and Daniel Khashabi. Benchmarking language model creativity: A case study on code generation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 2776–2794, 2025. Sarnoff Mednick. The associative basis of the creative process. Psychological review, 69(3): 220, 1962. Team Mistral. Introducing mistral 3, 2025. URL https://mistral.ai/news/mistral-3. Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. The conceptarc bench- mark: Evaluating understanding and generalization in the arc domain. arXiv preprint arXiv:2305.07141, 2023. Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, et al. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656, 2024. Team OpenAI.Introducing gpt-5.4, 2026.URLhttps://openai.com/index/ introducing-gpt-5-4/. Peter Organisciak, Selcuk Acar, Denis Dumas, and Kelly Berthiaume. Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity, 49:101356, 2023. Vishakh Padmakumar and He He. Does writing with language models reduce content diversity? arXiv preprint arXiv:2309.05196, 2023. 13 Preprint. Under review. Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URLhttps: //qwen.ai/blog?id=qwen3.5. Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A Haggag, Alfonso Amayuelas, et al. Include: Evaluating multilingual language understanding with regional knowledge. arXiv preprint arXiv:2411.19799, 2024. Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021. Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, AR Brown, A Santoro, A Gupta, A Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arxiv 2022. arXiv preprint arXiv:2206.04615, 10, 2022. Claire Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han van der Maas. Putting gpt-3’s creativity to the (alternative uses) test. arXiv preprint arXiv:2206.08932, 2022. Yiyou Sun, Shawn Hu, Georgia Zhou, Ken Zheng, Hannaneh Hajishirzi, Nouha Dziri, and Dawn Song. Omega: Can llms reason outside the box in math? evaluating exploratory, compositional, and transformative generalization. arXiv preprint arXiv:2506.18880, 2025. Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4149–4158, 2019. Zhen Tan, Dawei Li, Song Wang, Alimohammad Beigi, Bohan Jiang, Amrita Bhattacharjee, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. Large language models for data annotation and synthesis: A survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 930–957, 2024. Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas L Griffiths, and Faeze Brahman. Macgyver: Are large language models creative problem solvers? In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 5303–5324, 2024. Graham Todd, Tim Merino, Sam Earle, and Julian Togelius. Missed connections: Lateral thinking puzzles for large language models. In 2024 IEEE Conference on Games (CoG), p. 1–8. IEEE, 2024. Manya Wadhwa, Tiasa Singha Roy, Harvey Lederman, Junyi Jessy Li, and Greg Durrett. Create: Testing llms for associative creativity, 2026. URLhttps://arxiv.org/abs/2603. 09970. Shlomo Waks. Lateral thinking and technology education. Journal of Science Education and Technology, 6(4):245–255, 1997. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang, Kaoyan Lu, Shuai Lyu, Yifan Zhu, Ping Zong Junpeng Ding, Xinyu Liu, Qunlin Chen, et al. Crebench: Human-aligned creativity evaluation from idea to process to product. arXiv preprint arXiv:2511.13626, 2025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. 14 Preprint. Under review. Junyi Ye, Jingyi Gu, Xinyun Zhao, Wenpeng Yin, and Guiling Wang. Assessing the creativity of llms in proposing novel solutions to mathematical problems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, p. 25687–25696, 2025. Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chengxing Xie, Cunxiang Wang, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Yiming Zhang, Harshita Diddee, Susan Holm, Hanchen Liu, Xinyue Liu, Vinay Samuel, Barry Wang, and Daphne Ippolito. Noveltybench: Evaluating language models for humanlike diversity. arXiv preprint arXiv:2504.05228, 2025. Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2):1–124, 2023. 15 Preprint. Under review. A Data Contamination Analysis To investigate whether the performance advantage of thinking models on CRESOWLVE-Ru over CRESOWLVE-En could be attributed to memorization and potential data contamination, we conduct a targeted contamination analysis. To isolate other confounding factors, we restrict our analysis to the 580 questions thatGemini-3.1-Proanswered correctly in Russian but not in English. We focus on this model as the best-performing one in our evaluation, and deliberately use low thinking effort to ensure the model relies more on memorization than active reasoning. Following Golchin & Surdeanu (2025), we first generate three unique word-level perturbations for each question usingGPT-4o(prompt in Table X). We then run two contamination quizzes sharing the same prompt structure (prompt Table X) but with different answer choices. In the first quiz, we present the model with a multiple-choice question where all options are perturbed versions of the original question and an additional “none of the above” option is included; crucially, no correct answer is present. This allows us to identify positional biases by analyzing how frequently the model selects each option. We identify non-preferred positions as those chosen less frequently than random chance (less thank/4=145 withk =580 and 4 options), which yields options B and C as non- preferred(full distribution is ’A’: 187, ’B’: 60, ’C’: 47, ’D’: 273 with an additional 13 invalid answers). Notably, the model correctly prefers option D (“none of the above”) most of the time. In the second quiz, we insert the correct answer into one of these non-preferred positions (randomly chosen between B and C) for each question and repeat this quiz three times. The model achieves 72%, 73%, and 74% accuracy across the three runs (mean 73%), all substantially above the random baseline of 25%. This result strongly suggests that Gemini-3.1-Prohas likely been exposed to these questions during training, pointing to high-level data contamination as a plausible explanation for its cross-lingual performance advantage. Difficulty#questions#explanations Avg. #question tokens (En / Ru) Avg. # answer tokens (En / Ru) Very Simple (1)41737537.79± 12.84 / 29.65± 9.902.12± 1.51 / 1.82± 1.09 Simple (2)45342638.64± 12.84 / 30.79± 10.132.20± 1.73 / 1.87± 1.30 Medium (3)43641338.11± 12.13 / 30.34± 10.082.06± 1.38 / 1.87± 1.10 Hard (4)37135538.61± 13.43 / 30.60± 10.352.13± 1.88 / 1.85± 1.39 Very Hard (5)38435937.92± 13.05 / 30.27± 10.452.12± 1.46 / 1.89± 1.22 Total2061192838.22± 12.84 / 30.33± 10.182.13± 1.60 / 1.86± 1.22 Table 2: Final benchmark statistics by difficulty level. 123456 Number of domains 0 200 400 600 800 1000 1200 Number of questions 43 (2.1%) 383 (18.6%) 1202 (58.3%) 409 (19.8%) 22 (1.1%) 2 (0.1%) (a) Knowledge domains. 1234 Number of creative language constructs 0 250 500 750 1000 1250 1500 1750 Number of questions 152 (7.4%) 1824 (88.5%) 84 (4.1%) 1 (0.0%) (b) Creativity concepts. 1234568 Number of cultures/languages 0 200 400 600 800 1000 Number of questions 818 (39.7%) 1078 (52.3%) 137 (6.6%) 20 (1.0%) 5 (0.2%) 2 (0.1%) 1 (0.0%) (c) Cultures/demographics. Figure 6: Distribution of number of knowledge domains, creative language constructs and cultures/languages per question. 16 Preprint. Under review. Features BenchmarksConv.Div.Lat.AnalogyCom.SenseReal- World KG Domains BRAINTEASER (Jiang et al., 2023) """%"%Various RiddleSense (Lin et al., 2021)"%%%""Various Briangle Han et al. (2025)%""%%%Math/Logic SPLAT (Chen et al., 2024)%""%"%Daily-Life LatEval (Huang et al., 2024)%""%"%Daily-Life Columbus (Kraaijveld et al., 2025) """%%%Logic Connections (Todd et al., 2024) %""%"%Various DAT (Bellemare-Pepin et al., 2024) %"%%%%Various McGyver (Tian et al., 2024)""%%""Daily-Life OnlyConnect (Alavi Naeini et al., 2023) "%%%""Various CREATE (Wadhwa et al., 2026) ""%%%"Various (Lewis & Mitchell, 2024)"%%"%%Logic NEOCODER (Lu et al., 2025)""%%%%Coding CreativeMath (Ye et al., 2025)""%%%%Math OMEGA (Sun et al., 2025)"%%%%%Math CRESOWLVE (Ours)""""""Various Table 3: Comparison of different creative problem-solving benchmarks. Our benchmark no- tably requires employing several creative thinking strategies and involves various domains. (Conv.=Convergent, Div.=Divergent, Lat.=Lateral, Com.Sense=Commonsense reasoning) 17 Preprint. Under review. B Error Examples Missing creative connection Question In the first half of the 20th century, many American farmers had to move—some due to economic difficulties, others due to dust storms. For example, SHE moved to another country in 1911. Name HER name. Gold answer Dorothy. Model output In the first half of the 20th century, a significant migration of American farmers occurred, often driven by economic opportunities and agricultural developments. One of the most notable periods of this migration was between 1896 and 1914, when hundreds of thousands of American farmers moved to the Canadian Prairies (Alberta, Saskatchewan, and Manitoba) to take advantage of cheap, fertile land. This period is historically recognized as the “Great Migration” of American farmers to Canada. In 1911 alone, over 100,000 Americans emigrated to Canada. The placeholder SHE in the puzzle refers to a personification of the American farmer or a symbolic figure of the United States. Following the logic that SHE represents the collective identity of these migrants or a specific personification often used for ships or entities, the puzzle likely refers to the ship “The American Farmer.” However, a more direct interpretation of the riddle, considering the symbolic “SHE” often used for ships that transported these families, or the name of a famous individual, may point towards a specific ship or a personified noun. Yet, in the context of lateral thinking puzzles with these clues, the answer often points to a vessel or a specific entity named “The American Farmer ” that “moved” (in maritime terms) to another country (destination) during that era. Looking at ship records, the S American Farmer was a famous ship, though it was renamed later. Alternatively, some lateral thinking puzzles use “SHE” to refer to the “Mona Lisa”, which was stolen in 1911 and “moved” to Italy (another country). However, the connection to American farmers in the first sentence is typically a contextual setup for migration patterns of that era. Re-evaluating the direct clue, “SHE” is an example of an American farmer. In 1911, the vast number of farmers moving to Canada included many individuals, but the “SHE” most famously associated with the concept of the American farmer (in a figurative sense through art or iconography) or simply the ship named American Farmer fits the phrasing. In maritime culture, a ship named American Farmer would be referred to as “SHE.” <Answer>The American Farmer</Answer> Predicted answer: The American Farmer 18 Preprint. Under review. Overthinking Question Pyotr Vail writes about one of the properties of cinema, that essentially D’Artagnan is no better than Rochefort, but IT convinces the viewer of D’Artagnan’s rightness. “Barbarossa” was also IT. Name IT with one word. Gold answer Close-up. Model output To solve this lateral thinking puzzle, we look to the writings of Pyotr Vail, a renowned Russian-American interviewer and essayist. In his cultural commentary, Vail noted that in cinema, the protagonist (like D’Artagnan in The Three Musketeers) is often no more morally upright than the antagonist (Rochefort); however, cinema uses specific techniques to make the viewer empathize with the hero. The specific property Vail cites is the “close-up” (or in Russian, “крупныйплан”). By focusing the camera closely on the hero, cinema creates an intimate connection that “convinces the viewer of the hero’s rightness.” The second clue refers to “Barbarossa.” In a historical and military context, “Bar- barossa” was the name of a famous plan (the German invasion of the Soviet Union, known as Operation Barbarossa). In Russian, the word “план” (plan) can refer both to a military plan and to a cinematic shot scale (as inкрупныйплан). Based on this ambiguity, the model concludes that the correct answer is plan. <Answer>Plan</Answer> Predicted answer: Plan Hallucination or unsupported fabrication Question Upon learning about George Bush’s upcoming visit to Israel, a representative of “Al-Qaeda” called him one of THEM and urged to meet Bush with weapons in hand. Name the writer who made THEM his main characters. Gold answer Henryk Sienkiewicz. Model output In the context of the puzzle, “Al-Qaeda” is interpreted as a pun referring to a group of alcoholics (phonetically similar to “Al-Kashi” in Russian jokes). The model assumes that George Bush, as a recovering alcoholic, is being described as one of “THEM.” It further interprets “weapons” as bottles and concludes that the question refers to literary works centered on alcoholics. Based on this reasoning, the model identifies Venedikt Yerofeyev, known for Moscow-Petushki, as the correct answer. <Answer>Venedikt Yerofeyev</Answer> Predicted answer: Venedikt Yerofeyev 19 Preprint. Under review. Hallucination or unsupported fabrication Question In Paris, Thomas Jefferson was presented with a gift made by French craftsmen from African materials. The gift has survived to this day, although not completely: two elements are missing. Perhaps they were not lost but deliberately removed to avoid putting the official representative of the young republican state in an awkward position. Name these two elements as precisely as possible. Gold answer The white and black kings. Model output The model interprets the gift as a decorative object made with African materials such as ebony and ivory. It hypothesizes that the missing elements were symbols of slavery (chains and shackles), which may have been removed to avoid political awkwardness given Jefferson’s role and the ideals of liberty. <Answer>chains and shackles</Answer> Predicted answer: chains and shackles Incorrect concept anchoring Question Michael Chabon compared a river glistening with metal, whose waves crash between the embankments, to an X. The internet project “X,” conceived by Boris Berezovsky, was intended to report on violations of civil liberties in Russia. Name X. Gold answer Bell. Model output The model interprets the metaphor as referring to a zipper (“molniya” in Russian), and connects this to a supposed internet project named “Molniya” associated with Boris Berezovsky. <Answer>Molniya</Answer> Predicted answer: Molniya 20 Preprint. Under review. Incorrect concept anchoring Question In his descriptions, the author of the novel mentions the radiance of a rainbow, often adorning the head, as well as the crown of a weeping willow. Name this author. Gold answer Herman Melville. Model output The model interprets the clues as referring to the butterfly Apatura iris and asso- ciates this imagery with Vladimir Nabokov, known for his interest in butterflies and descriptive prose. <Answer>Vladimir Nabokov</Answer> Predicted answer: Vladimir Nabokov Incorrect concept anchoring Question In a humorous translation of the first part of the film “The Lord of the Rings,” Gandalf says to Frodo: “I love your Hobbiton — the nature here is beautiful.” In Gandalf’s line, we replaced two words with one. Write the replaced words. Gold answer New Zealand. Model output The model refers to a humorous Russian translation and interprets “Hobbiton” as a placeholder for the phrase “Tikhiy Omut” (“still waters”), concluding that the replaced words are “Still waters.” <Answer>Still waters</Answer> Predicted answer: Still waters 21 Preprint. Under review. Wrong hypothesis Question There are currently four such countries on the world map, although there used to be many more. Among them are Egypt, France, and Spain. Name the fourth country. Gold answer Panama. Model output The model identifies a shared property among Egypt, France, and Spain as countries bordering the Mediterranean Sea and having an additional coastline. It constructs a geographical explanation involving countries with dual coastlines and concludes that Morocco fits this pattern. <Answer>Morocco</Answer> Predicted answer: Morocco Wrong reference Question An alien from Zelazny’s novel gnaws off pieces of a diamond brick due to the peculiarities of its physiology. What did a character from an 1892 story involuntarily use for similar purposes? Gold answer Blue carbuncle. Model output The model interprets the question through a biological analogy involving contin- uously growing teeth, drawing a parallel between the alien and animals such as hares. It references a 19th-century story featuring a frightened hare and concludes that the character used his own teeth involuntarily. <Answer>His own teeth</Answer> Predicted answer: His own teeth 22 Preprint. Under review. Clue misinterpretation Question Sculptor Henry Moore equated women with nature, so DEFINITIONS in his works could symbolize ravines or depressions eroded by the forces of wind and water. What two words did we replace with the word “DEFINITIONS”? Gold answer Curvy lines. Model output The model focuses on Henry Moore’s use of negative space and interprets the description as referring to hollow spaces or holes within sculptures, drawing an analogy to natural formations such as caves or valleys. <Answer>hollow spaces</Answer> Predicted answer: hollow spaces Literature (1065) History (845) Film & Media Studies (402) Languages & Linguistics (375) Human Geography (291) Religious Studies (271) Sports (247) Anthropology (247) Biology (215) Visual Arts (190) Music (189) Engineering & Technology (186) Political Science (180) Home & Daily Life (171) Performing Arts (156) Psychology (127) Sociology (121) Earth & Environment (80) Physics (77) Astronomy (77) Business Studies (77) Military (71) Philosophy (64) Design & Arch. (63) Medicine & Health (62) Mathematics (58) Economics (56) Chemistry (46) Law & Criminology (39) Art & Culture (38) Other Sciences (36) Domain Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.680.770.740.780.740.700.700.770.710.720.780.770.750.730.780.760.760.790.770.790.730.720.730.730.770.690.750.720.690.660.83 0.730.800.750.790.740.750.730.790.750.770.800.810.770.730.820.820.780.890.770.840.780.790.730.760.770.720.770.760.770.710.94 0.200.280.250.300.250.280.240.290.250.200.210.320.260.250.280.200.260.330.300.310.300.250.300.240.240.280.300.220.180.240.39 0.170.250.170.250.250.240.190.210.170.170.200.240.240.190.170.180.160.290.210.380.260.140.230.220.210.310.250.170.180.130.36 0.230.280.220.300.280.260.230.290.260.240.290.260.280.300.270.250.260.350.250.270.170.250.250.250.210.380.250.260.260.130.42 0.170.220.180.210.230.240.220.190.210.190.190.270.270.190.220.210.230.330.210.230.290.230.300.240.290.160.300.220.210.320.25 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 7: Performance by domains on CRESOWLVE-En. lateral thinking (1740) analogy (830) abstraction (357) joke (226) pun (225) metaphor (207) commonsense reasoning (143) poem (83) idiom (79) neologism (36) sarcasm (34) proverb (31) Concept Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.730.760.710.750.720.610.790.580.720.750.680.81 0.770.780.760.770.740.690.800.640.780.750.680.77 0.250.260.260.220.240.180.340.220.200.310.260.32 0.200.210.250.150.220.150.270.170.190.220.240.29 0.250.270.310.210.240.200.290.220.230.310.290.29 0.200.220.210.200.180.160.230.160.280.190.240.29 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 8: Performance by creativity concepts on CRESOWLVE-En. 23 Preprint. Under review. Pipeline Step IDQuestion (Russian)Question (English)Answer Step 1: Filtering unanswerable questions (removed 295 samples) Requires physical materials 7e149044a7Статья в “Нью-Йорк Таймс”, по- священая иде возврата к золо- тому стандарту в целях укреп- ления долара, сопровождалась изображением, фрагмент кото- рого мы вам раздали. Воспроиз- ведите то, что мы закрыли чер- ным прямоугольником. An article in “The New York Times” about the idea of returning to the gold standard to strengthen the dollar was accompanied by an im- age, a fragment of which we dis- tributed to you. Reproduce what we covered with a black rectangle. In God We Trust Step 2: Filtering Russian-specific questions (removed 799 samples) Russian lan- guage/ cul- ture specific fbce8ea84bВ произведени Евгения Лукина ОН в ответ на вопрос героя при- зывает отрицать эпидемию. Не спрашиваем, какие четыре слова мы заменили словами “отрицать эпидемию”. Назовите ЕГО. In the work of Yevgeny Lukin, IT, in response to the hero’s question, urges to deny the epidemic. We do not ask which four words we replaced with “deny the epidemic.” Name IT. Ворон / Raven (Pun on Poe’s Raven: “Не верь в мор!” — “Don’t believe in plague!”) Step 3: Translation (all remaining 2,695 samples translated with GPT-4o) Translatedc97e6fd16dПо предположению американ- ских ученых, складки, которые образуются на мокрых пальцах, выполняют ту же функцию, что и ИКС. После смерти отца ИК- СОМ был провозглашен Ричард. Назовите его фамилию. According to American scientists, the wrinkles that form on wet fin- gers serve the same function as X. After the death of his father, Richard was proclaimed X. Name his surname. Кромвель/ Cromwell (X = Lord Protector; wrinkles protect grip.) Step 4: Human validation (removed 282 samples) Still Russian- specific 506e7ac07eСтарейший американский клуб, основаный в 1903 году, назы- вается “Белые медведи Кони- Айленда”. Сезон у них длится с ноября по апрель. Каким словом мы называем участников подоб- ных клубов? The oldest American club, founded in 1903, is called the “Coney Island Polar Bears.” Their season lasts from November to April. What word do we use to refer to mem- bers of such clubs? Моржи / Wal- ruses (Russian-specific term for winter swimmers.) Requires ex- ternal materi- als 17cad876d2Перед вами фрагмент альтерна- тивного постера к известному фильму. Назовите этот фильм. In front of you is a fragment of an alternative poster for a famous film. Name this film. “Кофе и сига- реты” / “Coffee and Cigarettes” Multiple questions e089848e4fБлиц. 1. На рекламной этикет- ке художника Баранова есть над- пись “С кружкой пивца и рабо- тается!” и изображение человека. Назовите этого человека. [...] Blitz. 1. On the advertising label by artist Baranov, there is an inscrip- tion “With a mug of beer, work is dear!” and an image of a person. Name this person. [...] 1. Ленин / Lenin 2. Ленин / Lenin 3. Ленин / Lenin Outdated info c95beab3bМихаил Прохоров, отвечая на неудобный вопрос, сказал, что ЕВА у него появилась в 17 лет. Назовите имя и фамилию нынеш- ней американской ЕВЫ. Mikhail Prokhorov, answering an awkward question, said that EVA appeared in his life at the age of 17. Name the first and last name of the current American EVA. Мишель Оба- ма / Michelle Obama (Answer is outdated.) Step 5: Creative vs. factual annotation (352 factual, 2,061 creative) Factual8947f691acВольф Шнайдер пишет, что один кристал ЕГО в 1970-х годах ли- шил работы 45 тысяч швейцар- цев. Назовите ЕГО. Wolf Schneider writes that one crystal of IT in the 1970s deprived 45,000 Swiss people of their jobs. Name IT. Кварц / Quartz (Quartzwatchesre- placed Swiss mechanical watches.) Creative83a712ed75В романе о военом времени ге- рой смотрит на небо, где инверси- оные следы самолетов и разры- вы зенитных снарядов образуют неприятную, по его мнению, кар- тину. Какое слово мы заменили на “картину”? In a wartime novel, the hero looks at the sky, where contrails of air- planes and bursts of anti-aircraft shells form an unpleasant, in his opinion, picture. What word did we replace with “picture”? Мелодию/ Melody (Contrails are staff lines; bursts are notes.) Table 4: Examples of questions affected at each stage of the data processing pipeline. For each step, we show a representative question in Russian (original) and English (translated), along with its answer and the reason for filtering or annotation. 24 Preprint. Under review. External Material Filtering Prompt You are a strict annotator. Your task is to decide whether the given puzzle EXPLICITLY instructs the participant to check external materials to answer the question. Answer "yes" ONLY if the puzzle requires explicit external resource lookup, such as: – A handout (“раздаточный материал”)or statements like “на розданой вам фото- графи”, “смотрите раздатку”, “не озвучивать текст раздатки”, “в раздатке”, or similar – References to a hidden or closed element such as: “мы закрыли символ”, “мы закрыли букву”, “мы закрыли часть текста”, “прочерк”,, when the missing content is NOT provided within the puzzle text. – Instructions requiring external lookup such as “зайдите по сылке и посмотрите курс”, “вставьте число из источника” etc. Answer "no" if: – The puzzle does not EXPLICITLY instruct the participant to check external materials. Output ONLY ’yes’ or ’no’ in lowercase, with no additional text. Puzzle: question Additional notes: comment Answer: Table 5: Zero-shot prompt used to filter questions requiring physical external materials (See §3.1) for details. english (1186) russian (796) french (210) german (160) greek (134) italian (119) latin (85) spanish (75) american (75) japanese (54) polish (34) Culture Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.750.670.740.780.770.710.810.710.830.810.71 0.790.720.760.810.790.780.850.720.910.760.68 0.270.180.250.230.260.290.310.330.350.410.21 0.220.130.260.190.240.180.340.280.280.310.21 0.280.210.230.260.250.230.440.290.290.350.24 0.220.170.210.220.280.210.280.200.250.280.21 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 9: Performance by cultures on CRESOWLVE-En. 25 Preprint. Under review. Translation Prompt You are a professional Russian to English translator. Your task is to translate Russian puzzles into English with absolute fidelity. Translate EXACTLY, preserving: – all logical clues – all named entities – sentence order and structure – rhetorical devices – ambiguity – references and style – the puzzle’s original difficulty Do NOT: – paraphrase – simplify or summarize – interpret hidden meanings – add explanations – rephrase stylistically If a segment is unclear, translate it literally. You are given the puzzle question, answer, comment and notes in Russian enclosed in <Question>, <Answer>, <Comment>, and <Notes> tags. Translate each section into English, preserving the structure and meaning, and output the translated puzzle in the same format with <Question>, <Answer>, <Comment>, and <Notes> tags. Output ONLY the English translation, nothing else. Russian Puzzle: <Question>question</Question> <Answer>answer</Answer> <Comment>comment</Comment> <Notes>notes</Notes> English Translation: Table 6: Zero-shot prompt used for translating puzzles from Russian to English (§3.1). Literature (1065) History (845) Film & Media Studies (402) Languages & Linguistics (375) Human Geography (291) Religious Studies (271) Sports (247) Anthropology (247) Biology (215) Visual Arts (190) Music (189) Engineering & Technology (186) Political Science (180) Home & Daily Life (171) Performing Arts (156) Psychology (127) Sociology (121) Earth & Environment (80) Physics (77) Astronomy (77) Business Studies (77) Military (71) Philosophy (64) Design & Arch. (63) Medicine & Health (62) Mathematics (58) Economics (56) Chemistry (46) Law & Criminology (39) Art & Culture (38) Other Sciences (36) Domain Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.810.840.850.850.810.820.800.890.830.820.870.890.840.850.840.860.810.900.870.860.840.850.840.870.890.900.930.890.920.840.94 0.840.870.860.860.860.830.840.910.870.830.880.910.870.870.870.880.800.930.880.870.830.900.880.900.920.880.880.830.950.920.94 0.230.280.260.300.240.250.260.260.250.220.250.330.290.290.240.270.230.360.350.300.210.200.340.240.340.330.390.260.230.290.44 0.180.230.180.220.210.250.180.240.210.170.220.270.240.180.230.220.140.330.220.260.230.140.220.220.270.340.270.280.280.130.36 0.240.300.230.270.260.260.240.320.290.220.280.330.300.300.220.300.210.360.270.230.220.240.280.300.320.310.340.240.150.130.31 0.240.290.230.260.270.280.240.280.270.220.260.310.320.290.270.270.260.360.320.300.230.270.410.380.390.190.390.220.180.390.42 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 10: Performance by domains on CRESOWLVE-Ru. 26 Preprint. Under review. Russian-Specific Filtering Prompt (Instruction) You are a strict annotator. You are given a puzzle in Russian, along with its answer, comment, and notes and translation of the puzzle into English. Your task is to decide whether the puzzle can not be solved in English because solving the puzzle REQUIRES knowledge specific to the Russian LANGUAGE. Answer ’yes’ ONLY if the solution depends on Russian-specific linguistic features such as: – idioms, sayings, or set expressions that break in translation – wordplay, puns, or jokes that work ONLY in Russian – Russian-specific homonyms or near-homonyms – phonetic or rhyming clues that exist ONLY in Russian – letter-based tricks involving the Russian alphabet or spelling – translating the puzzle into another language would change or destroy the meaning, making the puzzle unsolvable or fundamentally different. Answer ’no’ if: – the puzzle depends on general knowledge, culture, history, geography, literature, or context, even if these are Russian. – the puzzle uses Russian names/entities but no Russian linguistic tricks – the translated puzzle still is solvable without losing the key idea. Given the puzzle question, answer, comment, notes in Russian and the English translation of the puzzle, output your reasoning in English within <Reasoning>...</Reasoning> tags and your answer within <Answer>...</Answer> tags. Russian-Specific Filtering Prompt (Few-Shot Example Format) Russian puzzle: Question: question Answer: answer Comment: comment Notes: notes English translation: Question: question_en Answer: answer_en Comment: comment_en Notes: notes_en <Reasoning>reasoning</Reasoning> <Answer>shot_answer</Answer> Table 7: Few-shot prompt used to filter Russian-specific questions (See §3.1 for details). lateral thinking (1740) analogy (830) abstraction (357) joke (226) pun (225) metaphor (207) commonsense reasoning (143) poem (83) idiom (79) neologism (36) sarcasm (34) proverb (31) Concept Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.840.850.830.810.820.780.850.720.860.830.790.97 0.860.870.860.850.830.810.900.780.890.780.820.97 0.260.280.300.240.290.190.290.180.180.190.180.42 0.210.230.220.190.200.160.280.120.190.140.180.29 0.260.280.300.230.220.200.330.200.250.250.210.39 0.260.280.300.270.280.230.230.180.270.220.380.45 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 11: Performance by creativity concepts on CRESOWLVE-Ru. 27 Preprint. Under review. Creative vs. Factual Reasoning Annotation Prompt (Instruction) You are an expert in annotating type of reasoning involved in solving creative thinking puzzles. Given a creative thinking puzzle question, its answer, comments about the puzzle and other acceptable answers (if any), do TWO tasks: 1) Decide whether solving the puzzle mainly requires SIMPLE FACTUAL reasoning or CREATIVE ASSOCIATION of distant pieces of knowledge/facts. – Answer ’factual’ when the puzzle is answered by simply retrieving specific facts, dates, names, or straightforward commonsense reasoning, with no need to make a non-obvious, creative leaps between facts. – Answer ’creative’ when the solver in addition to retrieving the facts, must also make a non-obvious, creative connection, see a hidden twist, reinterpret words, or combine clues in an indirect way. 2) If your answer to the previous task is ’CREATIVE’, then choose ALL relevant creativity concepts involved in solving the puzzle from the following list (use EXACT spelling, lowercase): poem, metaphor, idiom, proverb, joke, pun, simile, sarcasm, hyperbole, neologism, analogy, abstraction, lateral thinking, divergent thinking, commonsense reasoning, compositionality. – Choose ALL that apply (multi-label). – If none of the specific labels fit, suggest a list of new labels that are relevant for the puzzle. Carefully read the given puzzle and output: – One line with the reasoning type: <Answer>factual</Answer> OR <Answer>creative</Answer>. – If you chose ’creative’, on separate lines output ONE OR MORE <concept> tags, each containing either exactly one label from the allowed list or a new label if none fits. – If you chose ’factual’, do NOT output any <concept> tags. Creative vs. Factual Reasoning Annotation Prompt (Few-Shot Example Format) Example index: Puzzle: question Answer: answer Comment: comment Acceptable Answers: notes Annotation: shot_answer Table 8: Few-shot prompt used for creative vs. factual reasoning annotation ( See §3.1 for details). 28 Preprint. Under review. Domain Annotation Prompt (Instruction) You are an expert in annotating creative thinking puzzles. Given a creative thinking puzzle, its answer, comments about the puzzle and other acceptable answers (if any), identify a list of domains that are involved in solving the puzzle. A domain is a specific area of knowledge, expertise, or human activity such as physics, literature, sports, music, etc. Wrap each domain on a separate line in <domain>...</domain> tags. Domain Annotation Prompt (Few-Shot Example Format) Example index: Puzzle: question Answer: answer Comment: comment Acceptable Answers: notes Domains: domains Table 9: Few-shot prompt used for knowledge domain annotation (§3.2). Knowledge Annotation Prompt (Instruction) You are an expert in annotating creative thinking puzzles. Given a creative thinking puzzle, identify a list of knowledge/facts that are explicitly required to answer the puzzle, and write them in the form of independent questions. Don’t solve the problem. Don’t include the answer itself. Wrap each question on a separate line in <knowledge>...</knowledge> tags. Knowledge Annotation Prompt (Few-Shot Example Format) Example index: Puzzle: question Answer: answer Comment: comment Acceptable Answers: notes Knowledge: knowledge Table 10: Few-shot prompt used for knowledge decomposition annotation (See §6). english (1186) russian (796) french (210) german (160) greek (134) italian (119) latin (85) spanish (75) american (75) japanese (54) polish (34) Culture Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) GPT-4.1 Qwen3-235B-A22B-Thinking DeepSeek-V3.2 GPT-5.4 (medium) Model 0.850.780.840.860.860.830.870.760.920.910.88 0.870.810.880.870.890.850.910.810.910.890.85 0.270.200.200.280.290.330.340.310.350.370.29 0.230.150.180.210.240.190.280.250.280.310.21 0.280.210.270.300.290.300.400.350.330.370.29 0.260.220.270.290.310.290.320.310.280.350.21 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Figure 12: Performance by cultures on CRESOWLVE-Ru. 29 Preprint. Under review. Culture/Demographics Annotation Prompt (Instruction) You are an expert in annotating creative thinking puzzles. Given a creative thinking puzzle, its answer, comments about the puzzle and other acceptable answers (if any), identify the languages/cultures that are involved in solving the puzzle. Output only your final answer and separate them with commas. Culture/Demographics Annotation Prompt (Few-Shot Example Format) Example index: Puzzle: question Answer: answer Comment: comment Acceptable Answers: notes Cultures/Languages: cultures Table 11: Few-shot prompt used for culture and demographics annotation (§3.2). Very SimpleSimpleMediumHardVery Hard Difficulty 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy CresOWLve-En Very SimpleSimpleMediumHardVery Hard Difficulty CresOWLve-Ru Model Llama-3.3-70B-Instruct Mistral-Large-3-675B-Instruct GPT-4.1 Qwen3-235B-A22B-Thinking (adaptive) DeepSeek-V3.2 (adaptive) GPT-5.4 (medium) Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) Figure 13: Exact Match Performance by difficulty. factualcreative Category 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy CresOWLve-En factualcreative Category CresOWLve-Ru Model Llama-3.3-70B-Instruct Mistral-Large-3-675B-Instruct GPT-4.1 Qwen3-235B-A22B-Thinking (adaptive) DeepSeek-V3.2 (adaptive) GPT-5.4 (medium) Gemini-3.1-Pro (medium) Gemini-3.1-Pro (high) Figure 14: Exact Match Performance by reasoning category. 30 Preprint. Under review. Chain-of-Thought Prompt (English) You are an expert at solving creative thinking puzzles. You are given a puzzle question that requires creative reasoning over real-world knowledge. Note that the question might contain several placeholder words (often in all capital case such as X, Y, THIS, FIRST, SECOND, ALPHA, BETA, HE, SHE, IT, HIM, HIS, HER, THEY, THEM etc.) that substitute for specific entities, objects, or concepts. These placeholders are crucial for solving the puzzle, and their meaning can only be inferred through careful reasoning about the question. Additionally, some placeholders might be gendered (e.g., HE/HIM/HIS vs SHE/HER), but do not assume that the gendered pronouns necessarily refer to human characters; they could refer to any entities, and their gender might be different. Given a puzzle question, think step by step and write your reasoning within <Reasoning>...</Reasoning> tags. Then provide only the final answer within <Answer>...</Answer> tags. Puzzle: question Chain-of-Thought Prompt (Russian) You are an expert at solving creative thinking puzzles. You are given a puzzle question in Russian that requires creative reasoning over real-world knowledge. Note that the question might contain several placeholder words (often in all capital case such as ИКС, ИГРЕК, ЭТО, ПЕРВЫЙ, ВТОРОЙ, АЛЬФА, БЕТА, ОН, ОНА, ЕГО, Е, ОНИ etc.) that substitute for specific entities, objects, or concepts. These placeholders are crucial for solving the puzzle, and their meaning can only be inferred through careful reasoning about the question. Given a puzzle question, think step by step and write your reasoning in Russian within <Reasoning>...</Reasoning> tags. Then provide only the final answer in Russian (unless specified otherwise) within <Answer>...</Answer> tags. Puzzle: question Thinking Model Prompt (English) You are an expert at solving creative thinking puzzles. You are given a puzzle question that requires creative reasoning over real-world knowledge. Note that the question might contain several placeholder words (often in all capital case such as X, Y, THIS, FIRST, SECOND, ALPHA, BETA, HE, SHE, IT, HIM, HIS, HER, THEY, THEM etc.) that substitute for specific entities, objects, or concepts. These placeholders are crucial for solving the puzzle, and their meaning can only be inferred through careful reasoning about the question. Additionally, some placeholders might be gendered (e.g., HE/HIM/HIS vs SHE/HER), but do not assume that the gendered pronouns necessarily refer to human characters; they could refer to any entities, and their gender might be different. Please provide your final answer within <Answer>...</Answer> tags. Puzzle: question Thinking Model Prompt (Russian) You are an expert at solving creative thinking puzzles. You are given a puzzle question in Russian that requires creative reasoning over real-world knowledge. Note that the question might contain several placeholder words (often in all capital case such as ИКС, ИГРЕК, ЭТО, ПЕРВЫЙ, ВТОРОЙ, АЛЬФА, БЕТА, ОН, ОНА, ЕГО, Е, ОНИ etc.) that substitute for specific entities, objects, or concepts. These placeholders are crucial for solving the puzzle, and their meaning can only be inferred through careful reasoning about the question. Please provide your final answer in Russian (unless specified otherwise) within <Answer>...</Answer> tags. Puzzle: question Table 12: Evaluation prompts for non-thinking (CoT) and thinking (reasoning) models in English and Russian (§4). 31 Preprint. Under review. LLM-as-a-Judge Prompt You are an expert in judging model responses. Given a question, a reference answer, a model answer, and additional context such as comments and other acceptable answers, decide whether the model’s answer is correct based on the reference answer and context. Ignore minor typos, articles, capitalization, and formatting. If the factual core is the same, answer yes. Also ignore the number of words constraint in the question if the model answer is semantically correct but does not meet the word count requirement. Answer only: Yes or No. Question: question Reference Answer: answer Additional Context: Comments: comment Acceptable Answers: notes Model Answer: prediction Table 13: Zero-shot prompt used for LLM-as-a-judge evaluation (§4). Very SimpleSimpleMediumHardVery Hard Difficulty 0 20 40 60 80 Count Distribution of Difficulty Types for Creative Samples Very SimpleSimpleMediumHardVery Hard Difficulty 0 10 20 30 40 50 60 70 80 Count Distribution of Difficulty Types for Factual Samples Figure 15: Distribution of difficulty levels for creative and factual questions. 12345 Difficulty Level 1 2 3 4 5 6 Number of Domains Number of Domains vs Difficulty 12345 Difficulty Level 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Max Domain Semantic Distance Max Domain Semantic Distance vs Difficulty 12345 Difficulty Level 0 2 4 6 8 Number of Facts Number of Facts vs Difficulty Figure 16: Correlations between question difficulty and complexity features. 32 Preprint. Under review. 1.00.50.00.51.0 Pearson correlation (r) Gemini-3-Flash (medium) GPT-5.4 (medium) Gemini-3.1-Pro (medium) Qwen3.5-397B-A17B (adaptive) DeepSeek-V3.2 (adaptive) GPT-4.1 Qwen3-235B-A22B-Thinking (adaptive) Llama-3.3-70B-Instruct Model Correlation: Accuracy vs Number of Domains 1.00.50.00.51.0 Pearson correlation (r) GPT-5.4 (medium) Gemini-3-Flash (medium) Qwen3-235B-A22B-Thinking (adaptive) DeepSeek-V3.2 (adaptive) GPT-4.1 Qwen3.5-397B-A17B (adaptive) Gemini-3.1-Pro (medium) Llama-3.3-70B-Instruct Model Correlation: Accuracy vs Max Domain Semantic Distance 1.00.50.00.51.0 Pearson correlation (r) Gemini-3.1-Pro (medium) GPT-5.4 (medium) GPT-4.1 Gemini-3-Flash (medium) Llama-3.3-70B-Instruct DeepSeek-V3.2 (adaptive) Qwen3-235B-A22B-Thinking (adaptive) Qwen3.5-397B-A17B (adaptive) Model Correlation: Accuracy vs Number of Knowledge Facts Figure 17: Correlations between model performance and complexity features. 33