Paper deep dive
Are LLMs becoming similarly creative? Evidence from three years of models
Nirav Patel, Josiah Crossman, Eva Aggarwal, Emily Wenger
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/21/2026, 3:11:35 AM
Summary
This paper analyzes the evolution of creativity in Large Language Models (LLMs) over a three-year period (2023-2026). Using the Alternate Uses Task (AUT) and Infinity-Chat100 prompts, the authors measure semantic diversity via sentence-embedding similarity. The study finds a statistically significant decrease in cross-model output diversity over time, indicating that LLMs are becoming more homogeneous in their creative outputs. This convergence raises concerns about the long-term impact on human agency and creativity in human-AI co-creative work.
Entities (9)
Relation Signals (6)
Nirav Patel ā affiliatedwith ā Duke University
confidence 99% Ā· Nirav Patel ... Affiliation: Department of Computer Science, Duke University
Infinity-Chat100 ā contains ā open-ended user queries
confidence 95% Ā· Infinity-Chat100, a real-world collection of open-ended user queries
Large Language Models ā exhibits ā decreasing diversity
confidence 95% Ā· Our findings show a statistically significant decrease in model output diversity over time
Alternate Uses Task ā usedtoassess ā Divergent Thinking
confidence 95% Ā· The AUT is a classic divergent-thinking assessment from cognitive psychology
Large Language Models ā convergesin ā creative substance
confidence 90% Ā· suggesting that LLM outputs may be converging in creative substance across models
LLMs ā maydiminish ā Human Agency
confidence 85% Ā· LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.
Tags
Links
- Source: https://arxiv.org/abs/2608.19437v1
- Canonical: https://arxiv.org/abs/2608.19437v1
Trouble viewing inline? Open PDF directly ā
Full Text
36,297 characters extracted from source content.
Expand or collapse full text
Are LLMs becoming similarly creative? Evidence from three years of models Nirav Patel Thanks: Corresponding author: nirav.patel@duke.edu Affiliation: Department of Computer Science, Duke University Josiah Crossman Affiliation: Department of Computer Science, Duke University Eva Aggarwal Affiliation: Department of Computer Science, Duke University Emily Wenger Affiliation: Department of Computer Science, Duke University Affiliation: Department of Electrical & Computer Engineering, Duke University Abstract Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMsā role in the human creative process. 1 Introduction In a few short years, conversational AI systems powered by large language models (LLMs) have moved from experimental prototypes to widely-used everyday tools. By July 2025, ChatGPT had more than 700 million weekly active users, who collectively sent more than 2.5 billion messages per day, or roughly 29,000 messages per second 9. Anthropic and DeepSeek, other popular AI chatbot providers, similarly have tens of millions of users 4; 2. Studies show a wide variety of use cases for LLM-powered products, but creativity-relevant tasks like writing, content creation, and research are common. Approximately 70% of consumer ChatGPT use is non-work-related, with practical guidance (a category which includes ācreative ideationā) and writing accounting for over half of ChatGPT messages 9. Similarly, Anthropic estimates that nearly 20% of Claude messages fall into categories like ācontent creationā or āmultidisciplinary academic research and writing.ā 31. As LLMs mediate more creative work, it is critical to understand how these systems affect the range of ideas users encounter and how model-generated outputs may shape downstream human creativity. A growing body of research suggests that LLM outputs can appear individually creative but are often collectively homogeneousāsimilar either to other ācreativeā content produced by the same LLM or content produced by other LLMs. In one study of AI-assisted creative writing, writers with access to AI-generated ideas produced stories that were judged as more creative and better written than stories from writers working alone. However, the AI-assisted stories were also more similar to one another than stories written without AI assistance 10. Across real-world open-ended queries, models exhibit both intra-model repetition and inter-model convergence 19. Human baseline studies on standardized divergent thinking tasks again surface this pattern, finding that LLM outputs have similar or higher creativity scores than humans on these tasks but are substantially less diverse 38; 36; 5. Algorithmic monoculture has long been a subject of academic concern (e.g. 40; 22), and observed homogeneity across AI creative outputs represents yet another facet of this well-documented problem. However, it remains unclear whether this homogeneity is a temporary byproduct of a still-developing technology or an inevitableāpotentially compounding 29āfeature of statistical language models. Plausible forces point in both directions. A growing body of academic work suggests that generative models trained on overlapping data and optimized toward similar objectives will organize concepts and semantic relationships in increasingly similar ways 16; 35; 39, resulting in similar outputs. Yet, as models become more capable, they could produce more context-sensitive, stylistically distinct, and domain-specific responses, counteracting homogeneity 18. All these factors suggest an urgent need to understand the trend of homogeneity in AI-generated creative outputs, but no such analysis exists. All existing work on homogeneity captures a snapshot of specific modelsā behavior, rather than a longitudinal view of model performance on creative tasks. Nor can examination of existing benchmarks for AI models elucidate this trend. Most popular benchmarks for LLMs, such as MMLU 13, HELM 6, SWE-Bench 20, and Humanityās Last Exam 7, judge model performance against explicit criteria with verifiable answers. In contrast, studying creative outputs requires evaluating models on questions for which response correctness is less important than its novelty, diversity, or contextual sensitivity. Our contribution. This paper fills this gap by performing a preliminary temporal analysis of model responses to open-ended prompts. We sample models across popular model providers and release periods from 2023 to the present, generate responses to two complementary sets of creative prompts, and compare outputs across models and generations. Our analysis combines the Alternate Uses Tasksā 12 assessment of divergent thinking with open-ended prompts from Infinity-Chat100 19, which span six categories of real-world use, from creative content generation and alternative styles of writing, to information seeking, brainstorming and ideation. We then analyze trends in model response diversity by computing responsesā semantic similarity via embedding-based distance metrics. This study provides a framework and initial findings for tracking how LLM creativity is evolving across model generations. Key findings. We find that LLM responses to the open-ended prompts we testāAUT and Infinity-Chat100āhave become increasingly similar over time. This suggests that LLMs are becoming less creative in tasks that involve generating open-ended responses, demanding scrutiny of their long-term usefulness as creative assistants. 2 Related Work Measuring creativity in LLMs. Recent work studying creativity in LLMs adapts a family of methods derived from human psychometric assessments of creativity. These often include Torrance-style creativity dimensions such as fluency, flexibility, originality, and elaboration 33, as well as divergent-thinking tasks that measure novelty through semantic distance 42; 24. Studies of AI creative writing combine expert human judgments, Consensual Assessment-style methods, and Torrance-inspired rubrics to test whether LLM outputs are genuinely creative products rather than merely fluent text 8. Other benchmarks shift from writing to problem solving, asking whether models can produce physically feasible, unconventional solutions to everyday challenges 32. More recent work on scientific ideation similarly treats creativity as domain-specific idea generation 28. Although applying human psychometric methods to AI models is imperfect 30, comparative studies suggest that LLMs can approach, and sometimes exceed, human performance on standard creativity tasks. For example, Hubert et al found that models score better than humans on standard creativity tasks, including the Divergent Association Test (DAT) 25 generating more original and elaborate answers 15. However, a larger study of 9,198 humans and 215,542 LLM observations finds that humans are slightly more creative on average than LLMs on the DAT 37. Recent work similarly finds that LLMs exceed average human DAT performance but fall below the most creative human subgroups. Notably, there was significant repetition across model DAT responsesāāoceanā appeared in more than 90 percent of modelsā DAT response sets 5. This suggests an emerging āAI creativity paradoxā: LLMs may appear individually creative but produce less diverse outputs in aggregate. Homogeneity in LLMs. A growing body of literature documents this homogeneity across creative settings. In short-story generation, LLMs produce fluent and stylistically complex text, but score below humans on novelty, surprise, and diversity in expert and automated evaluations 17. AI-assisted writing studies find a similar tradeoff, in which LLM-generated ideas improve individual story quality but reduce collective diversity 10. Still more studies show that LLMs generate responses that are more similar to one another than human responses are, reuse plot elements, and exhibit both intra-model repetition and inter-model homogeneity across real-world open-ended prompts 38; 19; 41. While these observations of homogeneity are concerning, existing work merely provides snapshots of this phenomenon at a particular point in time. Furthermore, trends in homogeneity cannot easily be reverse-engineered from existing collections of benchmark data, which judge model performance against explicit criteria with verifiable answers (e.g. 13; 6; 20; 7). Even creativity-specific benchmarks, like CreativityPrism, LitBench, and Deep Associations, focus largely on whether LLMs can meet recognizable standards of creativity 14; 11; 27, not relationships between responses. Thus, no framework exists for tracking how the diversity or originality of LLM creative outputs is evolving. Figure 1: Methodology overview: we select open-ended prompts designed to elicit creative behavior from models; run these against different model lineages; embed model outputs in a semantic space to compute their semantic distance; and perform a regression analysis to understand changes in semantic distance over time. 3 Methodology We perform our time-based analysis of LLM output diversity via a four step process. We first select a set of open-ended prompts, run these across generations of language models, embed the responses in a semantic space to compute distances between them, and perform regression analysis to study trends in response distance. These steps are summarized in Figure 1 and described in detail below. Selecting prompts. We select two prompt sets that elicit creativity in distinct ways: the Alternate Uses Task (AUT) 12 and 100 real-world open-ended prompts from Infinity-Chat100 19. The AUT is a classic divergent-thinking assessment from cognitive psychology in which participants generate unconventional uses for common objects such as a book, shoe, or hammer. Its standardized structure isolates differences in the ideas models produce under consistent task constraints, providing a controlled setting for measuring responsesā semantic similarity. Conversely, the Infinity-Chat100 prompts provide a broader, more natural set of creative tasks. They were drawn from Infinity-Chat 19, a large dataset of real-world user interactions with language models, and span creative content generation, problem-solving, brainstorming, ideation, and other open-ended requests. These prompts permit substantial variation in both what models produce and how they approach each task. This naturally complements the more structured approach of the AUT. Evaluating prompts against model lineages. To construct the response dataset, we evaluate these prompts against 68 models released between March 2023 and July 2026, representing 12 major model providers. This covers 33 closed weight models and 34 open weight models across 27 release months (See A for model coverage details). Each model receives the complete set of AUT and Infinity-Chat100 prompts through OpenRouterās API, with temperature and top-p set to 1.0 for all generations (See A for data coverage details). These settings retain stochastic sampling without distorting the output distribution, allowing us to measure similarity among independently generated responses without narrowing models toward their highest-probability outputs. We assign each model a release date based on the month and year they were made publicly available, and order the resulting responses accordingly. Release date may not perfectly capture the combination of changes that distinguishes successive model generations, including newer training data, architectures, and post-training methods. However, we believe it serves as a consistent observable proxy for each modelās position in the broader progression of LLM development. Trends across release periods should therefore be interpreted as associations with broader model development patterns rather than as causal effects of time itself. Compute embeddings and semantic distances. Next, we embed model responses using the all-MiniLM-L6-v2 sentence-transformer model and compute semantic distances between these embeddings. The sentence-transformer represents each response as a dense vector in a semantic space, such that responses with similar meanings are represented by vectors pointing in similar directions. For the AUT, the complete set of uses generated in response to each object prompt was embedded as a single response, yielding 10 embeddings per model. For the Infinity-Chat dataset, each response was embedded separately, yielding 100 embeddings per model. Regression specification. Finally, we perform a linear regression of semantic distance between model responses over time (See A for OLS assumptions). To do this, we must first place the model pairs on a common timeline. We chronologically order the 27 (month, year) timestamps in our sample and group them into nine successive bins of three timestamps. Bins therefore span unequal calendar durations, because model releases have become more regular as more providers enter the market and existing providers iterate more rapidly. Grouping releases into equal-count bins ensures that each bin, even in the sparse early period, contributes measurable information to the linear regression. Using nine bins of three timestamps each balances a reasonable number of models per bin with appropriate separation between model releases. Following 38, we only compare responses between models from different providers (e.g. ācross-familyā), to minimize possible confounders like architecture or training data overlap. With our time bins established, we now compute distances between model responses within each bin. Formally, for each cross-family model pair within a bin, we match responses to the same prompts. Let pā1,ā¦,Ppā\1,ā¦,P\ index prompts, with P=10P=10 for the AUT and P=100P=100 for Infinity-Chat100, and let i,px_i,p and j,px_j,p denote the embedding vectors produced by models i and j for prompt p. We compute prompt-level cosine distances as diāj,p=1ācosā”(i,p,j,p),d_ij,p=1- (x_i,p,x_j,p ), where lower values indicate greater semantic similarity and higher values indicate greater divergence. We summarize each pair by its mean distance across prompts, dĀÆiāj=1Pāāp=1Pdiāj,p, d_ij= 1P _p=1^Pd_ij,p, which serves as the single observation for pair (i,j)(i,j), computed separately for each dataset. To test whether cross-family divergence changes systematically over time, we fit an ordinary least squares regression of dĀÆiāj d_ij on biājā1,ā¦,9b_ijā\1,ā¦,9\, the ordinal index of the release bin containing the pair, dĀÆiāj=α+βābiāj+εiāj, d_ij=α+β\,b_ij+ _ij, so that β β is interpretable as the average change in cross-family cosine distance per release bin. Model resampling for regression. Because model families contain different numbers of models, an unadjusted regression would weight heavily represented families more strongly. To reduce this influence, we run 1,000 iterations of family-balanced resampling. At each iteration kā1,ā¦,1000kā\1,ā¦,1000\, we sample one model from each family represented in each release bin, compute distances between cross-family response sets from the sampled models, and perform a separate regression for each response set across all bins. This yields a distribution of slope estimates β^(k)k=11000\ β^(k)\_k=1^1000. We summarize the temporal trend using the median slope across iterations, with the 2.5th and 97.5th percentiles reported as empirical resampling intervals. We additionally construct a pointwise 95% band from the corresponding percentiles of the fitted regression lines at each bin index, which we display alongside the binned means in Figure 2. These intervals quantify how sensitive the trend is to which model represents each family, addressing the confounding influence of unequal family representation at any point in time. Because our interest is not in any individual model but rather in the collective outputs of models in a given period, the model that happens to represent its family within a bin should not matter. A tight interval thus indicates that the trend is not an artifact of any particular selection of models and holds regardless of which model stands in for its family in each bin. A consistent negative slope then indicates decreasing cross-family divergence over time. Figure 2: Regression results and bootstrapped slope estimates for AUT and Infinity-Chat response distances across model generations. Our findings of consistently negative slopes across the observation period indicate semantic convergence across LLM creative outputs over time. 4 Results Key finding: For both the AUT and Infinity-Chat response sets, we observe a decline in cross-provider output distances over the observation period. This suggest decreasing diversityāor increasing homogeneityāof LLM creative outputs over time. Figure 2 shows that cross-family cosine distance declines over the observation period for both prompt sets, with the corresponding slope estimates reported in Table 1. The decline is most pronounced for the Alternate Uses Task, where mean cross-family distance falls from approximately 0.50 in the earliest release bin to below 0.40 in the most recent. Infinity-Chat exhibits the same pattern at a much gentler rate, declining from approximately 0.34 to just above 0.32. Notably, all 1,000 resampling iterations produced negative slopes for both prompt sets, indicating that the direction of the relationship is robust across the sampled combinations of models from each family. Table 1: Bootstrap OLS estimates for mean cosine distance across nine model release bins. Dataset Models (Pairs) Slope per Bin 95% CI Alternate Uses Task 68 (273) ā0.01385-0.01385 [ā0.01695,ā0.01044][-0.01695,-0.01044] Infinity-Chat 67 (268) ā0.00167-0.00167 [ā0.00267,ā0.00074][-0.00267,-0.00074] The magnitude of the AUT decline is particularly notable given the taskās purpose. The AUT explicitly tests divergent thinking by instructing models to produce uses that are as original and unexpected as possible, making it precisely the setting in which outputs would be expected to differ. The Infinity-Chat results extend this finding beyond a single, highly structured task format. Its 100 prompts elicit many forms of open-ended generation, ranging from imaginative and expressive tasks to reflective and conversational ones. A decline across this heterogeneous collection is therefore less likely to reflect the creative demands of any one task and instead suggests a broader change in the distinctiveness of modelsā creative outputs over time. Although the overall Infinity-Chat slope is shallower, the decline is most visible in the recent release bins (e.g. 2025 onward). This pattern may mark the beginning of a longer-term trend that warrants continued monitoring. 5 Discussion Our findings, though preliminary, raise concerns about the long-term usefulness of LLMs as creative partners. Even if models perform well on creative tasks, converging outputs could bound the range of possibilities LLM users are exposed to, and with it, the breadth of their own thinking. If using an LLM for creative tasks like essay writing decreases oneās brain activity 23, could using increasingly less creative LLMsāthe trend suggested by our studyāfurther worsen LLMsā effects on human creativity, as observed by this and other studies 10; 1? Future work should examine this possibility while accounting for the limitations of this initial work. Limitations. Although our regression approach is rigorous, it cannot account for hidden relationships across model families, such as distillation 26; 3, shared training data 35, or common training techniques. Second, we anecdotally observe several prompts in the InfinityChat dataset that request similar outputs (e.g. 2 prompts asking about the movie Zootopia). While this does not affect our measurement of response similarity, since we compute distances between model responses on the same prompts, it suggests the need for a more principled set of unstructured creative prompts. Moreover, while sampling a single response per model per prompt is sufficient for our aggregate trend analysis, it yields only a point estimate of each modelās output distribution. Comparing distributions drawn from repeated generations could more richly characterize modelsā output creativity. Finally, our analysis does not account for the role of the user in shaping the diversity of LLM outputs. Recent work suggests that prompting alone cannot prevent homogeneity in AI creative outputs 38, but savvier use of AI models by humans may produce better results 34; 21. Future work. Our work leaves open many avenues for interesting future work. Our aggregate analysis does not establish whether the observed creative similarity is uniform across all types of open-ended prompts, so a prompt-level decomposition may help identify the factors driving model convergence. For example, dividing prompts by task type, category, and output length could reveal whether similarity patterns differ across short-form creativity, long-form writing, or practical ideation. Our embedding-based distance measurements could be complemented by text-level and structural measures, such as Jaccard similarity, sentence structure, response organization, and other stylistic features. Further robustness checks across prompt wording, sampling parameters, response length constraints, model family, and model size would also help rule out confounders. Such directions would help clarify not only the existence of defined trends in LLMsā creative evolution, but also the underlying mechanisms driving them. References [1] B. R. Anderson, J. H. Shah, and M. Kreminski (2024) Homogenization effects of large language models on human creative ideation. In Creativity and Cognition, p. 413ā425. External Links: Link, Document Cited by: §5. [2] Anthropic (2025) Anthropic economic index report: uneven geographic and enterprise ai adoption. Note: https://w.anthropic.com/research/anthropic-economic-index-september-2025-reportAccessed: 2026-06-26 Cited by: §1. [3] Anthropic (2026) Detecting and preventing distillation attacks. Note: https://w.anthropic.com/news/detecting-and-preventing-distillation-attacks Cited by: §5. [4] E. Baptista (2026) A year on from DeepSeek shock, get set for flurry of low-cost Chinese AI models. Note: https://w.reuters.com/world/china/year-deepseek-shock-get-set-flurry-low-cost-chinese-ai-models-2026-02-12/Reuters. Accessed: 2026-06-26 Cited by: §1. [5] A. Bellemare-Pepin, F. Lespinasse, P. Thƶlke, Y. Harel, K. Mathewson, J. A. Olson, Y. Bengio, and K. Jerbi (2026) Divergent creativity in humans and large language models. Scientific Reports 16, p. 1279. External Links: Document, Link Cited by: §1, §2. [6] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. S. Chatterji, A. S. Chen, K. A. Creel, J. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. E. Gillespie, K. Goel, N. D. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. F. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. S. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. P. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. F. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. H. Roohani, C. Ruiz, J. Ryan, C. Rāe, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. P. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. TramĆØr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. A. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang (2021) On the opportunities and risks of foundation models. ArXiv. External Links: Link Cited by: §1, §2. [7] Center for AI Safety, Scale AI, and HLE Contributors Consortium (2026) A benchmark of expert-level academic questions to assess AI capabilities. Nature 649, p. 1139ā1146. External Links: Document, 2501.14249, Link Cited by: §1, §2. [8] T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, and C. Wu (2024) Art or artifice? large language models and the false promise of creativity. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, p. 1ā34. Cited by: §2. [9] A. Chatterji, T. Cunningham, D. J. Deming, Z. Hitzig, C. Ong, C. Y. Shan, and K. Wadman (2025) How people use chatgpt. Working Paper Technical Report 34255, National Bureau of Economic Research. External Links: Document, Link Cited by: §1, §1. [10] A. R. Doshi and O. P. Hauser (2024) Generative ai enhances individual creativity but reduces the collective diversity of novel content. Science advances 10 (28), p. eadn5290. Cited by: §1, §2, §5. [11] D. Fein, S. Russo, V. Xiang, K. Jolly, R. Rafailov, and N. Haber (2026) Litbench: a benchmark and dataset for reliable evaluation of creative writing. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7740ā7755. Cited by: §2. [12] J. P. Guilford, P. R. Christensen, P. R. Merrifield, and R. C. Wilson (1978) Alternate uses. Cited by: §1, §3. [13] D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2020) Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300. Cited by: §1, §2. [14] Z. J. Hou, B. A. Zhang, Y. Lu, B. K. Baghel, A. Brei, X. Lu, M. Jiang, F. Brahman, S. Chaturvedi, H. Chang, D. Khashabi, and X. L. Li (2026) CreativityPrism: a holistic evaluation framework for large language model creativity. External Links: 2510.20091, Link Cited by: §2. [15] K. F. Hubert, K. N. Awa, and D. L. Zabelina (2024) The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific reports 14 (1), p. 3440. Cited by: §2. [16] M. Huh, B. Cheung, T. Wang, and P. Isola (2024) The platonic representation hypothesis. arXiv preprint arXiv:2405.07987. Cited by: §1. [17] M. Ismayilzada, C. Stevenson, and L. van der Plas (2025) Evaluating creative short story generation in humans and large language models. External Links: 2411.02316, Link Cited by: §2. [18] S. Jain, J. Lanchantin, M. Nickel, K. Ullrich, A. Wilson, and J. Watson-Daniels (2025) LLM Output Homogenization is Task Dependent. arXiv. Note: arXiv:2509.21267 [cs] External Links: Link, Document Cited by: §1. [19] L. Jiang, Y. Chai, M. Li, M. Liu, R. Fok, N. Dziri, Y. Tsvetkov, M. Sap, A. Albalak, and Y. Choi (2025) Artificial hivemind: the open-ended homogeneity of language models (and beyond). Note: NeurIPS 2025 Datasets and Benchmarks Track oral External Links: 2510.22954, Document, Link Cited by: §1, §1, §2, §3, §3. [20] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan (2024) SWE-bench: can language models resolve real-world github issues?. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1, §2. [21] N. Jo and M. Raghavan (2026) Incentives shape how humans co-create with generative AI. arXiv. Note: arXiv:2604.03529 [cs.HC] External Links: Link, Document Cited by: §5. [22] J. Kleinberg and M. Raghavan (2021) Algorithmic monoculture and social welfare. Proceedings of the National Academy of Sciences 118 (22), p. e2018340118. External Links: Link, Document Cited by: §1. [23] N. Kosmyna, E. Hauptmann, Y. T. Yuan, J. Situ, X. Liao, A. V. Beresnitzky, I. Braunstein, and P. Maes (2025) Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv. Note: arXiv:2506.08872 [cs.AI] External Links: Link, Document Cited by: §5. [24] K. Nakajima, J. Zuiderveld, and S. Pezzelle (2026) Beyond divergent creativity: a human-based evaluation of creativity in large language models. arXiv preprint arXiv:2601.20546. Cited by: §2. [25] J. A. Olson, J. Nahas, D. Chmoulevitch, S. J. Cropper, and M. E. Webb (2021) Naming unrelated words predicts creativity. Proceedings of the National Academy of Sciences 118 (25), p. e2022340118. External Links: Link, Document Cited by: §2. [26] Open AI (2024) Model Distillation in the API. Note: https://openai.com/index/api-model-distillation/ Cited by: §5. [27] Z. Qiu and R. Hu (2025) Deep associations, high creativity: a simple yet effective metric for evaluating large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 10870ā10883. Cited by: §2. [28] K. Ruan, X. Wang, J. Hong, P. Wang, Y. Liu, and H. Sun (2026) Evaluating llmsā divergent thinking capabilities for scientific idea generation with minimal context. Nature communications. Cited by: §2. [29] I. Shumailov, Z. Shumaylov, Y. Zhao, Y. Gal, N. Papernot, and R. Anderson (2023) The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv. Note: arXiv:2305.17493 [cs] External Links: Link, Document Cited by: §1. [30] M. Stella, T. T. Hills, and Y. N. Kenett (2023) Using cognitive psychology to understand gpt-like models needs to extend beyond human biases. Proceedings of the National Academy of Sciences 120 (43), p. e2312911120. Cited by: §2. [31] A. Tamkin, M. McCain, K. Handa, E. Durmus, L. Lovitt, A. Rathi, S. Huang, A. Mountfield, J. Hong, S. Ritchie, M. Stern, B. Clarke, L. Goldberg, T. R. Sumers, J. Mueller, W. McEachen, W. Mitchell, S. Carter, J. Clark, J. Kaplan, and D. Ganguli (2024) Clio: Privacy-Preserving Insights into Real-World AI Use. arXiv. Note: arXiv:2412.13678 [cs.CY] External Links: Link, Document Cited by: §1. [32] Y. Tian, A. Ravichander, L. Qin, R. Le Bras, R. Marjieh, N. Peng, Y. Choi, T. L. Griffiths, and F. Brahman (2024) MacGyver: are large language models creative problem solvers?. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 5303ā5324. Cited by: §2. [33] E. P. Torrance (1966) Torrance test of creative thinking. Scholastic Testing Service, Inc.. Cited by: §2. [34] M. Urban, J. Lukavskỳ, C. Brom, V. Hein, F. Svacha, F. DÄchtÄrenko, and K. Urban (2025) Prompting for creative problem-solving: a process-mining study. Learning and Instruction 99, p. 102156. Cited by: §5. [35] H. Vu, G. Reeves, and E. Wenger (2026) What happens when generative ai models train recursively on each othersā outputs?. In International Conference on Learning Representations, Vol. 2026, p. 142067ā142086. Cited by: §1, §5. [36] D. Wang, D. Huang, H. Shen, and B. Uzzi (2026) A large-scale comparison of divergent creativity in humans and large language models. Nature Human Behaviour 10 (3), p. 531ā540. External Links: Document, Link Cited by: §1. [37] D. Wang, D. Huang, H. Shen, and B. Uzzi (2026) A large-scale comparison of divergent creativity in humans and large language models. Nature Human Behaviour 10 (3), p. 531ā540. Cited by: §2. [38] E. Wenger and Y. N. Kenett (2026) Large language models are homogeneously creative. PNAS Nexus 5 (3), p. pgag042. External Links: Document, Link Cited by: §1, §2, §3, §5. [39] P. West and C. Potts (2025) Base Models Beat Aligned Models at Randomness and Creativity. arXiv. Note: arXiv:2505.00047 [cs] External Links: Link, Document Cited by: §1. [40] F. Wu, E. Black, and V. Chandrasekaran (2025) Generative monoculture in large language models. In International Conference on Learning Representations, Vol. 2025, p. 33068ā33107. Cited by: §1. [41] W. Xu, N. Jojic, S. Rao, C. Brockett, and B. Dolan (2025) Echoes in ai: quantifying lack of plot diversity in llm outputs. Proceedings of the National Academy of Sciences 122 (35), p. e2504966122. Cited by: §2. [42] Y. Zhao, R. Zhang, W. Li, and L. Li (2025) Assessing and understanding creativity in large language models. Machine Intelligence Research 22 (3), p. 417ā436. Cited by: §2. Appendix A Appendix Model coverage details. Figure A1 shows all the models used in our analysis along with their public release dates. Data coverage details. Across 69 models, the AUT dataset contained 690 possible model-level prompt responses, of which 668 were observed and 22 were missing or dropped (3.19%). Infinity-Chat100 contained 6,900 possible responses, of which 6,677 were observed and 223 were missing or dropped (3.23%). Nine models had at least one missing response across the two datasets, while the remaining 60 models were complete in both. Table A1 reports the observed and missing counts for these nine models. Note that responses are missing or dropped only when an API request failed after repeated trials or when the returned output was corrupted or unusable as a direct response to the prompt. Additionally, when coverage was incomplete, the mean cosine distance for each model pair was computed only across prompts for which both models had valid responses. This preserved the available pairwise information without imputing missing outputs. Given the low overall rate of missingness and aggregation across 1,000 model-sampling iterations, missing responses were unlikely to materially influence our reported estimates. OLS assumptions. Figure A2 reports standard OLS diagnostics for both regressions. Residuals are roughly centered across fitted values and release bins, and departures from linearity roughly track the number of models in each bin while the overall relationship remains approximately linear. The Q-Q plots indicate approximately normal residuals for Infinity-Chat and modest right-tail deviation for the AUT. The most influential observations by Cookās distance are concentrated in the earliest release bins, where fewer models were available and each cross-family pair carries more weight. Our family-balanced resampling further mitigates this by ensuring no single model dominates the estimates. As such, we do not expect these deviations to affect our analysis, but rather see them as confirming the need to continue collecting data as models continue to release. Figure A1: Distribution of model releases over time across the 12 model providers and 68 versions in our dataset. Table A1: Observed and missing responses for models with at least one missing response. Each model had 10 possible AUT responses and 100 possible Infinity-Chat100 responses. AUT Infinity-Chat100 Model Observed Missing Observed Missing minimax/minimax-01 0 10 0 100 minimax/minimax-m1 6 4 0 100 mistralai/mistral-small-3.1-24b-instruct 10 0 89 11 qwen/qwen3.6-max-preview 2 8 100 0 qwen/qwen3-max 10 0 94 6 meta-llama/llama-3.2-3b-instruct 10 0 97 3 anthropic/claude-fable-5 10 0 99 1 minimax/minimax-m2.1 10 0 99 1 qwen/qwen-2.5-72b-instruct 10 0 99 1 All 69 models 668 22 6,677 223 Figure A2: OLS Linearity, Normality, Independence, and Influential Observations