Paper deep dive
Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty
Tim Schopf, Tobias Schreieder, Akiko Aizawa
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 3:17:26 AM
Summary
The paper introduces Think-Probe-Respond (TPR), a lightweight method to improve Large Language Models' (LLMs) accuracy in judging research idea novelty. It identifies a systematic miscalibration where LLMs exhibit a bias toward 'medium novelty' judgments despite generating accurate reasoning rationales. TPR mitigates this by probing latent novelty beliefs from the model's hidden states during the reasoning phase and conditioning the final response on these probed judgments, achieving a 22.30% performance improvement over baselines.
Entities (7)
Relation Signals (4)
Think-Probe-Respond â evaluatedon â RINoBench
confidence 95% · We conduct all experiments on rino... TPR improves novelty judgment performance... on the rino test set.
Think-Probe-Respond â mitigates â Novelty Miscalibration
confidence 95% · TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent 'medium novelty' bias.
Large Language Models â exhibits â Novelty Miscalibration
confidence 92% · their final novelty judgments often diverge substantially... stems from a systematic bias towards judging ideas as 'medium novel'.
Think-Probe-Respond â uses â Hidden States
confidence 90% · TPR probes latent novelty judgments from hidden states during the reasoning phase.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as "medium novel". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent "medium novelty" bias.
Tags
Links
- Source: https://arxiv.org/abs/2608.25660v1
- Canonical: https://arxiv.org/abs/2608.25660v1
Trouble viewing inline? Open PDF directly â
Full Text
109,840 characters extracted from source content.
Expand or collapse full text
LLM large language model MAE Mean Absolute Error RINoBench Research Idea Novelty Judgment Benchmark SOTA state-of-the-art TPR Think-Probe-Respond CoT Chain-of-Thought Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty Tim Schopf Affiliation: National Institute of Informatics, Tokyo, Japan Email: tim.schopf@t-online.de Tobias Schreieder Affiliation: TU Dresden & ScaDS.AI Dresden/Leipzig, Germany Email: tobias.schreieder@tu-dresden.de Akiko Aizawa Affiliation: National Institute of Informatics, Tokyo, Japan Email: aizawa@nii.ac.jp Abstract Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as âmedium novelâ. To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent âmedium noveltyâ bias. Figure 1: Overview of our tpr approach. The snowflake () indicates frozen parameters, while the flame () indicates trainable parameters. 1 Introduction Judging the novelty of research ideas is crucial to scientific progress, enabling original contributions and shaping future scientific directions. However, manual novelty judgment requires substantial expertise and a broad understanding of the relevant literature, making it time-consuming, subjective, and difficult to scale Picard et al. (2025). As scientific output continues to grow rapidly Fortunato et al. (2018), automated support for novelty judgment becomes important for helping researchers assess, refine, and compare research ideas. Recent work increasingly relies on llm to judge research idea novelty (Lu et al., 2024; Li et al., 2024a; Si et al., 2025; Su et al., 2025; Lu et al., 2026; Gottweis et al., 2026; Mostafa et al., 2026; Wu et al., 2026, inter alia). However, a critical discrepancy persists: while these models produce plausible, human-like novelty arguments Afzal et al. (2026), their final novelty judgments remain fundamentally disconnected from their own reasoning rationales and fail to align with human judgments Si et al. (2025); Schopf and FĂ€rber (2026). In this work, we investigate the underlying causes of this miscalibration and demonstrate that llm exhibit a systematic bias toward conservative âmedium noveltyâ judgments, even when their internal reasoning supports substantially different judgments. Motivated by this finding, we propose tpr (tpr), a lightweight method for mitigating novelty miscalibration. tpr probes latent novelty beliefs from hidden states during reasoning and conditions the final response on these extracted beliefs. Experimental results show that tpr improves novelty judgment performance by 22.30% on average over strong baselines while producing less biased novelty judgments. 2 Related Work Early efforts for automated research idea novelty judgment have evolved from citation- and lexical-based methods Uzzi et al. (2013); Wang et al. (2017); Amplayo et al. (2019); Wang et al. (2019); Sarica et al. (2020) to semantic embedding approaches GĂłmez-PĂ©rez et al. (2022). While these methods improve semantic matching, they largely remain limited to surface-level similarity estimation Mysore et al. (2022). More recent work adopt llm for automated novelty judgments (Wu et al., 2025; Si et al., 2025; Liu et al., 2025; Wang et al., 2025; Baek et al., 2025; Lin et al., 2025; Zhang et al., 2025; Tang et al., 2025; Li et al., 2025; Feng et al., 2025; Hou et al., 2026, inter alia). In contrast to prior work, we investigate a previously overlooked failure mode of llm: the systematic miscalibration between their generated novelty rationales and final novelty judgments. 3 Benchmark We conduct all experiments on rino Schopf and FĂ€rber (2026), the only publicly available benchmark for research idea novelty judgment with human-annotated novelty scores. It comprises 1,381 expert-authored ideas, each paired with related works, a human-annotated novelty score on a five-point Likert scale (rubric in Table 5), and an expert-written justification supporting the assigned novelty score. Given a research idea and its related work, models must predict the novelty score ranging from 1 (not novel) to 5 (highly novel) and generate a textual justification grounded in comparisons with prior work. F1F_1 MAE ALI Recall Add. Ratio Hall. Rate Model Macro 1 2 3 4 5 KA NA KA NA KA NA Gemini 2.5 Pro 15.6 0.0 6.3 28.8 42.9 0.0 1.0 0.41 64.7 60.6 104.6 96.5 8.2 1.2 Gemini 3 Pro 15.3 9.8 30.9 35.8 0.0 9.8 1.0 0.46 64.9 67.0 101.9 81.2 13.8 3.2 Claude Sonnet 4.5 15.5 0.0 14.6 43.2 19.5 0.0 0.9 0.59 80.0 70.3 150.1 14.8 8.7 1.2 Claude Opus 4.5 17.1 0.0 12.1 41.6 31.9 0.0 0.8 0.61 80.4 68.8 143.4 106.0 8.0 1.2 GPT-5 mini 16.2 0.0 5.1 40.8 34.9 0.0 0.9 0.62 77.3 67.8 144.8 104.8 7.8 1.0 GPT-5.4 15.1 0.0 3.0 48.3 24.3 0.0 0.8 0.62 75.8 67.8 117.2 77.3 4.8 0.6 Table 1: Evaluation results of novelty judgments on the rino test set. The reported metrics include F1F_1 macro averaged and for each rubric category (1-5), mae, Alignment (ALI), Recall, Additional Ratio (in %), and Hallucination Rate (in %) for Known Aspects (KA) and Novelty Aspects (NA) respectively (for more details on the metrics, see Appendix B). 4 Miscalibration in LLM Novelty Judgments Although llm novelty rationales align closely with human reasoning Afzal et al. (2026), their final novelty judgments often diverge significantly from human consensus Si et al. (2025); Schopf and FĂ€rber (2026). To investigate this miscalibration and examine why llm struggle to produce accurate novelty judgments, we prompt six sota (sota) llm (Gemini 2.5 Pro Comanici et al. (2025), Gemini 3 Pro Google (2025), Claude Sonnet 4.5 Anthropic (2025b), Claude Opus 4.5 Anthropic (2025a), GPT-5 mini Singh et al. (2026), GPT-5.4 OpenAI (2026)) to judge research idea novelty (prompt in Figure 3). As Table 1 shows, all models perform poorly on the novelty judgment task, yielding mae values around one and consistently low macro-F1F_1 scores, with the best-performing model achieving just 17.1. LLMs Avoid Extreme Novelty Judgments The dominant failure pattern is a strong middle-ground bias. Across models, predictions concentrate on novelty classes 3 and 4, while the lowest and highest categories are rarely predicted correctly. With the exception of Gemini 3 Pro, which occasionally assigns classes 1 and 5, models fail almost entirely on these extreme cases. Table 8 shows examples of this llm behavior. Latent Belief vs. Expressed Novelty Judgment This middle-ground bias contrasts with the quality of the generated justifications. Models achieve high recall with respect to human-annotated justification arguments, indicating that they often identify the same overlaps, differences, and novelty aspects as human experts. Overall, these results suggest that llm often âknowâ substantially more about the true novelty level than their generated numerical novelty scores reveal. Takeaway The miscalibration of llm as novelty judges is therefore best understood as a mismatch between latent belief and expressed judgment. The modelsâ reasoning often contains evidence for accurate novelty judgment, but their final outputs are biased towards safe middle categories. 5 Probing Latent LLM Judgments Building on the finding that llm internalize beliefs about research idea novelty that closely mirror those of human expertsâand therefore generate comparable novelty rationalesâyet are biased towards predicting medium novelty categories, we propose the tpr approach for research idea novelty judgment. tpr explicitly exploits the modelâs internal beliefs during reasoning about novelty to yield less biased and more accurate quantitative judgments. These judgments are then reused as conditioning signals to generate textual justifications that are coherent and well-aligned with the predicted numerical scores. As illustrated in Figure 1, tpr consists of three stages. (1) Think: We instruct an llm to judge the novelty of a research idea and to think step by step before producing a final response. For reasoning models that generate think tokens by default, we omit any explicit âthink step by stepâ instruction. Importantly, we provide only textual descriptions of the novelty categoriesâwithout numerical scoresâand instruct the model to evaluate novelty solely based on these descriptions without generating numerical judgments. This design encourages the model to think about the novelty of research ideas qualitatively, avoiding anchoring its internal representations to explicit numeric outcomes that could bias the reasoning process. Figure 4 shows the prompt used in this approach. (2) Probe: Given an LLM with L hidden layers, let H=h(1),âŠ,h(L)H=h^(1),âŠ,h^(L) represent the stack of hidden states. For a generated sequence of reasoning (âthinkâ) tokens T=t1,âŠ,tnT=t_1,âŠ,t_n, we terminate generation upon the production of the final think token tnt_n and extract the hidden state htn(L)h^(L)_t_n (Section 6 motivates the choice of tnt_n). This representation htn(L)h^(L)_t_n is then used as the input feature vector for a logistic regression probing classifier.11 1 While probing intermediate layers is a viable alternative, we follow prior work showing that representations from the final layer typically yield strong probing performance and offer practical advantages, as extraction of the last hidden state of a given model is more easily facilitated in common open-llm frameworks than earlier layers Maiya et al. (2025). (3) Respond: We append the textual description of the predicted novelty class to the llm-generated output and resume generation. Conditioned on both its prior reasoning and the predicted novelty judgment, the llm generates the final response, which is used as the justification of the novelty judgment. tpr is computationally efficient and lightweight. All llm parameters remain frozen and are used only at inference. Training is limited to a simple logistic regression classifier, which can be learned efficiently on CPU. 5.1 Experimental Setup We evaluate tpr against a diverse set of approaches. These include Zero-shot prompting, as used in Section 4; Few-shot prompting adapted from Shahid et al. (2025), which provides one example per novelty class; cot (cot) prompting, where the model is instructed to reason step by step before producing a novelty judgment; several prompt-based methods derived from Moose Yang et al. (2024), ResearchAgent Baek et al. (2025), AI Scientist Lu et al. (2024), and AI Researcher Si et al. (2025); as well as a FineTune approach that uses LoRA Hu et al. (2022) for llm fine-tuning (training details are provided in Appendix C). Since we require access to model parameters, we conduct experiments exclusively with open-source llm spanning multiple model families. We evaluate reasoning models that generate explicit think tokens by default, namely Qwen3 (4B, 14B, 32B) Yang et al. (2025) and GPT-OSS-20B OpenAI et al. (2025), as well as non-reasoning models for which we explicitly instruct step-by-step reasoning to elicit think tokens under tpr, including Gemma 3 (4B, 12B, 27B) Google et al. (2025) and Llama 3.1 (8B, 70B) Grattafiori et al. (2024). 5.2 Evaluation Results Figure 2: Macro-F1F_1 scores for different approaches and llm on the rino test set. Figure 2 presents macro-F1F_1 scores for different llm and approaches (see Appendix E for metric selection details). Across all evaluated models, tpr achieves the highest performance, surpassing the best competing approach by an average of 22.30%. Remarkably, this performance is achieved using only a lightweight logistic regression probing classifier, significantly surpassing the computationally expensive FineTune approach. While FineTune often outperforms prompt-based approaches, it cannot match the performance of tpr. Consistency Across Prompt-Based Approaches The effectiveness of prompt-based methods varies substantially across models, with no single prompting strategy consistently outperforming others. In contrast, tpr demonstrates robust, model-agnostic performance, delivering strong results across all llm investigated Model Size Is Not Determinative Probing llm beliefs during reasoning yields strong novelty judgments regardless of model size. Larger models do not necessarily perform better: for instance, the biggest Llama-3.1-70B ranks second-worst, whereas the smallest Gemma3-4B ranks second-best. Moreover, all open-source models using tpr surpass the novelty judgment macro-F1F_1 scores of proprietary models in Table 1. This indicates that even smaller models using tpr can outperform prompting substantially larger llm. Reasoning vs. Non-Reasoning Models tpr is effective for both reasoning and non-reasoning models, showing that think tokens encode useful information about novelty judgments, whether generated automatically or via explicit instruction. Interestingly, reasoning models exhibit slightly lower tpr performance on average than non-reasoning models. Fine-tuning reasoning models such as Qwen3 provides modest gains, yet tpr still outperforms FineTune, albeit with a smaller margin than for non-reasoning models. Overall, the largest gains appear when tpr is applied to non-reasoning models with explicit "think step-by-step" instructions. tpr vs. cot Across all llm, tpr substantially outperforms the cot approach. While cot marginally benefits non-reasoning models, its effect on reasoning models is inconsistent, and its overall performance lags behind both FineTune and tpr. These results suggest that merely instructing llm to reason is insufficient; explicitly probing the hidden representations formed during the reasoning process is crucial for achieving high-quality novelty judgments. tpr Enables Balanced Predictions Across Novelty Classes Table 2 reports class-wise F1F_1 scores for different llm using our tpr approach, while Table 1 shows the corresponding scores for proprietary LLMs under zero-shot prompting. Compared to the prompting approach, tpr produces substantially more balanced predictions across all novelty classes. In particular, prompted proprietary llm often avoid extreme novelty judgments (classes 1 and 5), whereas tpr enables models to recognize both very low and very high research idea novelties. The exception is the Qwen3 model family, which rarely assigns class 1 even under tpr, indicating a persistent tendency to avoid âno noveltyâ judgments. Overall, these results show that tpr not only improves macro-level F1F_1 performance but also encourages more uniform coverage of the full novelty spectrum, mitigating the middle-class novelty judgment bias observed in Section 4. F1F_1 Model Macro 1 2 3 4 5 Non-Reas. Gemma-3-4B 24.4 27.3 21.5 35.11 29.71 8.3 Gemma-3-12B 21.8 7.7 16.3 34.1 35.0 15.7 Gemma-3-27B 25.2 18.2 19.8 28.9 33.5 25.4 Llama-3.1-8B 23.4 14.3 26.0 36.0 31.2 9.5 Llama-3.1-70B 20.6 7.7 27.9 33.9 24.1 9.5 Reas. Qwen3-4B 20.0 0.0 24.4 31.2 28.2 16.1 Qwen3-14B 21.2 0.0 27.1 33.5 37.4 7.8 Qwen3-32B 21.1 0.0 23.5 36.5 35.8 9.8 GPT-OSS-20B 22.8 10.0 24.1 35.7 34.1 10.0 Table 2: F1F_1 scores per novelty class using tpr with different llm. An additional evaluation of the textual justifications is provided in Appendix F. 6 Probing over Time Thinking Phase Response Generation Phase Model t1t_1 t25%t_25\% t50%t_50\% t75%t_75\% tnt_n Avg. r1r_1 r25%r_25\% r50%r_50\% r75%r_75\% rnr_n Avg. Non-Reason. Gemma-3-4B 22.74 16.26 15.47 16.86 24.38 19.14 20.84 18.09 19.86 17.94 19.81 19.31 Gemma-3-12B 22.34 22.12 17.60 21.12 21.74 20.98 17.87 16.23 23.73 20.53 17.20 19.11 Gemma-3-27B 18.88 21.36 18.42 21.51 25.17 21.07 21.90 16.98 15.48 19.83 14.96 17.83 Llama-3.1-8B 17.01 19.90 19.77 19.94 23.38 20.00 17.92 19.47 17.40 17.37 19.51 18.33 Llama-3.1-70B 21.11 19.12 20.16 15.42 20.61 19.28 15.82 21.56 18.49 20.78 16.50 18.63 Reason. Qwen3-4B 19.06 19.23 20.51 21.37 19.98 20.03 16.89 19.25 17.84 18.89 17.78 18.13 Qwen3-14B 19.05 16.76 19.65 16.76 21.17 18.68 20.47 18.56 19.77 19.81 19.49 19.62 Qwen3-32B 17.51 17.77 19.76 22.26 21.13 19.69 19.32 19.30 20.20 22.01 20.07 20.18 GPT-OSS-20B 22.65 17.02 19.00 15.85 22.78 19.46 19.68 20.00 22.24 19.21 18.43 19.91 Table 3: Macro F1F_1 scores for research idea novelty judgments on the rino test set obtained by probing last-layer hidden states of various llm at different generation steps: when generating the first and last think tokens (t1t_1, tnt_n), the first and last response tokens (r1r_1, rnr_n), and intermediate steps during both the thinking and response generation phases (reported as percentages of token generation within each phase; e.g., t50%t_50\% denotes probing halfway through the thinking phase, after 50% of think tokens have been generated). We investigate at which generation stages llm encode the most salient information for research idea novelty judgments when applying tpr. For this analysis, we allow the models to generate their full reasoning and response exactly as instructed by the prompt in Figure 4, without inserting the probing classifierâs prediction as an intermediate control signal. This setup enables a clean examination of when novelty-related information naturally emerges in the modelâs representations. We probe last-layer hidden states at multiple time steps during both the thinking and response generation phases and report the results in Table 3. Novelty Signals Peak at the End of the Thinking Phase Across nearly all models, probing at the end of the thinking phase (tnt_n) yields the strongest or near-strongest novelty judgment performance. This pattern holds consistently for both reasoning and non-reasoning models, indicating that novelty-related beliefs are most fully consolidated once the model has completed its internal reasoning process. In contrast, probing earlier thinking tokens (e.g., t25%t_25\% or t50%t_50\%) generally results in substantially lower performance, suggesting that novelty representations emerge progressively rather than being present at the onset of reasoning. Probing Is Robust Across Reasoning Lengths As shown in Table 4, models vary widely in the number of tokens generated during the thinking phase, from short sequences to long reasoning chains. Despite these differences, probing at tnt_n provides consistently strong novelty signals. Importantly, there is no clear correlation between the absolute length of reasoning chains and probing performance: models with shorter or longer chains achieve comparable F1F_1 scores at the final thinking token. This indicates that the position within the reasoning sequence (final token) matters more than absolute length for capturing novelty-related beliefs. # Think Tokens # Response Tokens Model Min Max Avg. Min Max Avg. Non-Reason. Gemma-3-4B 380 1572 741.98 21 128 80.69 Gemma-3-12B 233 1672 631.21 44 105 69.25 Gemma-3-27B 206 804 431.64 48 107 71.81 Llama-3.1-8B 17 4358 435.33 21 152 67.79 Llama-3.1-70B 18 1034 313.76 23 120 57.23 Reason. Qwen3-4B 312 2525 1055.51 45 193 125.74 Qwen3-14B 256 3076 674.28 53 142 92.48 Qwen3-32B 275 2127 633.75 65 155 106.49 GPT-OSS-20B 7 265 57.66 21 231 126.13 Table 4: Number of tokens generated by different llm during research idea novelty prediction using our tpr approach. Note, we use GPT-OSS-20B with reasoning level âlowâ, resulting in a small number of think tokens. Response Generation Dilutes Novelty Representations While some models achieve local maxima at intermediate response steps (e.g., r50%r_50\% or r75%r_75\%), probing during the response generation phase produces competitive but generally weaker results than probing at tnt_n. This suggests that once the model transitions to response generation, the representations become increasingly influenced by surface realization and linguistic planning, diluting the underlying novelty signal. This trend holds for both reasoning and non-reasoning models. While some models exhibit local peaks at intermediate response steps, the final thinking token remains the most reliable and stable probing point overall. Takeaway Novelty judgments are primarily formed during the reasoning phase and are most reliably captured toward their later stages. 7 Conclusion We showed that llm are miscalibrated judges of research idea novelty. Although their rationales often align with human reasoning, their final judgments are biased towards medium novelty. To mitigate this, we proposed tpr, a lightweight approach that probes latent novelty judgments from hidden states during reasoning and conditions the final response on the probed judgment. Experiments demonstrate that tpr improves novelty judgment performance over strong baselines by 22.30% and reduces the prevalent medium-novelty bias. 8 Limitations Our experiments are conducted on rino, which focuses on machine learning research ideas and may not fully reflect novelty judgments in other scientific domains. In addition, tpr requires access to model hidden states, limiting its direct applicability to closed-source llm. Finally, novelty judgments are inherently subjective, and even expert annotations may reflect individual preferences or incomplete knowledge of the literature. 9 Ethical Considerations We emphasize that this work is intended for research and educational purposes only. Users should not use our models or approaches to make formal or high-stakes judgments of research ideas, as novelty judgments are inherently subjective and context-dependent. Our work is intended to advance AI-assisted scientific discovery by enabling models to reason about and explain novel contributions in research. However, automated predictions of research idea novelty should not replace human expert judgment. Our approaches are intended as tools to support, rather than replace human judgments of research ideas. Acknowledgments This work is supported by a scholarship of the German Academic Exchange Service (DAAD) - 57557629 and by the BMFTR through a Software Campus project with identification number 16|S23070. The authors acknowledge the financial support by the Federal Ministry of Research, Technology and Space of Germany (BMFTR) and by SĂ€chsische Staatsministerium fĂŒr Wissenschaft, Kultur und Tourismus in the programme Center of Excellence for AI-research âCenter for Scalable Data Analytics and Artificial Intelligence Dresden/Leipzigâ, project identification number: ScaDS.AI. We used AI-based assistance tools to support language editing, minor formatting, and coding tasks. These tools did not contribute to the intellectual content or scientific conclusions. All content was reviewed by the authors, who assume full responsibility for the publication. References Afzal et al. (2025) A. Afzal, F. Matthes, G. Chechik, and Y. Ziser Knowing before saying: LLM representations encode information about chain-of-thought success before completion. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 12791â12806. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: Appendix A. Afzal et al. (2026) O. M. Afzal, P. Nakov, T. Hope, and I. Gurevych Beyond ânot novel enoughâ: enriching scholarly critique with LLM-assisted feedback. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 2648â2671. External Links: Link, Document, ISBN 979-8-89176-380-7 Cited by: §1, §4. Amplayo et al. (2019) R. K. Amplayo, S. Hwang, and M. Song Evaluating research novelty detection: counterfactual approaches. In Proceedings of the Thirteenth Workshop on Graph-Based Methods for Natural Language Processing (TextGraphs-13), D. Ustalov, S. Somasundaran, P. Jansen, G. GlavaĆĄ, M. Riedl, M. Surdeanu, and M. Vazirgiannis (Eds.), Hong Kong, p. 124â133. External Links: Link, Document Cited by: §2. Anthropic (2025a) Anthropic Introducing Claude Opus 4.5 â anthropic.com. Note: https://w.anthropic.com/news/claude-opus-4-5[Accessed 05-01-2026] Cited by: §4. Anthropic (2025b) Anthropic Introducing Claude Sonnet 4.5 â anthropic.com. Note: https://w.anthropic.com/news/claude-sonnet-4-5[Accessed 05-01-2026] Cited by: §4. Baek et al. (2025) J. Baek, S. K. Jauhar, S. Cucerzan, and S. J. Hwang ResearchAgent: iterative research idea generation over scientific literature with large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 6709â6738. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2, §5.1. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, L. Marris, S. Petulla, C. Gaffney, A. Aharoni, N. Lintz, T. C. Pais, H. Jacobsson, I. Szpektor, N. Jiang, K. Haridasan, A. Omran, N. Saunshi, D. Bahri, G. Mishra, E. Chu, T. Boyd, B. Hekman, A. Parisi, C. Zhang, K. Kawintiranon, T. Bedrax-Weiss, O. Wang, Y. Xu, O. Purkiss, U. Mendlovic, I. Deutel, N. Nguyen, A. Langley, F. Korn, L. Rossazza, A. RamĂ©, S. Waghmare, H. Miller, N. Byrd, A. Sheshan, R. Hadsell, S. Bhardwaj, P. Janus, T. Rissa, D. Horgan, A. Abdagic, L. Belenki, J. Allingham, A. Singh, T. Guidroz, S. Srinivasan, H. Schmit, K. Chiafullo, A. Elisseeff, N. Jha, P. Kolhar, L. Berrada, F. Ding, X. Si, S. B. Mallick, F. Och, S. Erell, E. Ni, T. Latkar, S. Yang, P. Sirkovic, Z. Feng, R. Leland, R. Hornung, G. Wu, C. Blundell, H. Alvari, P. Huang, C. Yip, S. Deur, L. Liu, G. Surita, P. Duque, D. Damen, J. Jia, A. Guez, M. Mircea, A. Sinha, A. Magni, P. Stradomski, T. Marian, V. GaliÄ, W. Chen, H. Husain, A. Singhal, D. Grewe, F. Aubet, S. Song, L. Blanco, L. Rechis, L. Ho, R. Munoz, K. Zheng, J. Hamrick, K. Mather, H. Taitelbaum, E. Rutherford, Y. Lei, K. Chen, A. Shukla, E. Moreira, E. Doi, B. Isik, N. Shabat, D. RogoziĆska, K. Kolipaka, J. Chang, E. VuĆĄak, S. Venkatachary, S. Noghabi, T. Bharti, Y. Jun, A. Zaks, S. Green, J. Challagundla, W. Wong, M. Mohammad, D. Hirsch, Y. Cheng, I. Naim, L. Proleev, D. Vincent, A. Singh, M. Krikun, D. Krishnan, Z. Ghahramani, A. Atias, R. Aggarwal, C. Kirov, D. Vytiniotis, C. Koh, A. Chronopoulou, P. Dogra, V. Ion, G. Tyen, J. Lee, F. Weissenberger, T. Strohman, A. Balakrishna, J. Rae, M. Velic, R. de Liedekerke, O. Elyada, W. Yuan, C. Liu, L. Shani, S. Kishchenko, B. Alessio, Y. Li, R. Song, S. Kwei, O. Jankowski, A. Pappu, Y. Namiki, Y. Ma, N. Tripuraneni, C. Cherry, M. Ikonomidis, Y. Ling, C. Ji, B. Westberg, A. Wright, D. Yu, D. Parkinson, S. Ramaswamy, J. Connor, S. H. Yeganeh, S. Grover, G. Kenwright, L. Litchev, C. Apps, A. Tomala, F. Halim, A. Castro-Ros, Z. Li, A. Boral, P. Sho, M. Yarom, E. Malmi, D. Klinghoffer, R. Lin, A. Ansell, P. K. S, S. Zhao, S. Zuo, A. Santoro, H. Cheng, S. Demmessie, Y. Liu, N. Brichtova, A. Culp, N. Braun, D. Graur, W. Ng, N. Mehta, A. Phillips, P. Sundberg, V. Godbole, F. Liu, Y. Katariya, D. Rim, M. Seyedhosseini, S. Ammirati, J. Valfridsson, M. Malihi, T. Knight, A. Toor, T. Lampe, A. Ittycheriah, L. Chiang, C. Yeung, A. FrĂ©chette, J. Rao, H. Wang, H. Srivastava, R. Zhang, R. Rhodes, A. Brand, D. Weesner, I. Figotin, F. Gimeno, R. Fellinger, P. Marcenac, J. Leal, E. Marcus, V. Cotruta, R. Cabrera, S. Luo, D. Garrette, V. Axelrod, S. Baltateanu, D. Barker, D. Chen, H. Toma, B. Ingram, J. Riesa, C. Kulkarni, Y. Zhang, H. Liu, C. Wang, M. Polacek, W. Wu, K. Hui, A. N. Reyes, Y. Su, M. Barnes, I. Malhi, A. Siddiqui, Q. Feng, M. Damaschin, D. Pighin, A. Steiner, S. Yang, R. S. Boppana, S. Ivanov, A. Kandoor, A. Shah, A. Mujika, D. Huang, C. A. Choquette-Choo, M. Patel, T. Yu, T. Creswell, Jerry, Liu, C. Barros, Y. Razeghi, A. Roy, P. Culliton, B. Xiong, J. Pan, T. Strohmann, T. Powell, B. Seal, D. DeCarlo, P. Shyam, K. Katircioglu, X. Wang, C. Hardin, I. Odisho, J. Broder, O. Chang, A. Nair, A. Shtefan, M. OâBrien, M. Agarwal, S. Potluri, S. Goyal, A. Jhindal, S. Thakur, Y. Stuken, J. Lyon, K. Toutanova, F. Feng, A. Wu, B. Horn, A. Wang, A. Cullum, G. Taubman, D. Shrivastava, C. Shi, H. Tomlinson, R. Patel, T. Tu, A. M. Oflazer, F. Pongetti, M. Yang, A. A. TaĂŻga, V. Perot, N. W. Pierse, F. Han, Y. Drori, I. Iturrate, A. Chakrabarti, L. Yeung, D. Dopson, Y. Chen, A. Kulshreshtha, T. Guo, P. Pham, T. Schuster, J. Chen, A. Polozov, J. Xing, H. Zhou, P. Kacham, D. Kukliansky, A. Miech, S. Yaroshenko, E. Chi, S. Douglas, H. Fei, M. Blondel, P. Myla, L. Madmoni, X. Wu, D. Keysers, K. Kjems, I. Albuquerque, L. Yu, J. Dâsa, M. Plantan, V. Ionescu, J. S. Elias, A. Gupta, M. R. Vuyyuru, F. Alcober, T. Zhou, K. Ji, F. Hartmann, S. Puttagunta, H. Song, E. Amid, A. Stefanoiu, A. Lee, P. Pucciarelli, E. Wang, A. Raul, S. Petrov, I. Tian, V. Anklin, N. Nti, V. Gomes, M. Schumacher, G. Vesom, A. Panagopoulos, K. Bousmalis, D. Andor, J. Jacob, Y. Zhang, B. Rosgen, M. Kecman, M. Tung, A. Belias, N. Goodman, P. Covington, B. Wieder, N. Saxena, E. Davoodi, M. Huang, S. Maddineni, V. Roulet, F. Campbell-Ajala, P. G. Sessa, Xintian, Wu, G. Lai, P. Collins, A. Haig, V. Sakenas, X. Xu, M. Giustina, L. E. Shafey, P. Charoenpanit, S. Garg, J. Ainslie, B. Severson, M. G. Arenas, S. Pathak, S. Rajayogam, J. Feng, M. Bakker, S. Li, N. Wichers, J. Rogers, X. Geng, Y. Li, R. Jagerman, C. Jia, N. Olmert, D. Sharon, M. Mauger, S. Mariserla, H. Ma, M. Mohabey, K. Kim, A. Andreev, S. Pollom, J. Love, V. Jain, P. Agrawal, Y. Schroecker, A. Fortin, M. Warmuth, J. Liu, A. Leach, I. Blok, G. P. Girirajan, R. Aharoni, B. Uria, A. Sozanschi, D. Goldberg, L. Ionita, M. T. Ribeiro, M. Zlocha, V. Birodkar, S. Lachgar, L. Yuan, H. Choudhury, M. Ginsberg, F. Zheng, G. Dibb, E. Graves, S. Lokhande, G. Rasskin, G. Muraru, C. Quick, S. Tata, P. Sermanet, A. Chawla, I. Karo, Y. Wang, S. Zhang, O. Keller, A. Dragan, G. Su, I. Chou, X. Liu, Y. Tao, S. Prabhakara, M. Wilson, R. Liu, S. Wang, G. Evans, D. Du, A. Castaño, G. Prasad, M. E. Mahdy, S. Gerlach, M. Reid, J. Kahn, A. Zait, T. S. Pillai, T. Ulrich, G. Wang, J. Wassenberg, E. Farkash, K. Yalasangi, C. Wang, M. Bauza, S. Bucher, T. Liu, J. Yan, G. Leung, V. Sindhwani, P. Barnes, A. Singh, I. Jurin, J. Chang, N. K. Bhumihar, S. Eiger, G. Citovsky, B. Withbroe, Z. Li, S. Xue, N. D. Santo, G. Stoyanov, Y. Raimond, S. Zheng, Y. Gao, V. ListĂk, S. Kwasiborski, R. Saputro, A. Ozturel, G. Mallya, K. Majmundar, R. West, P. Caron, J. Wei, L. Castrejon, S. Vikram, D. Ramachandran, N. Dhawan, J. Park, S. Smoot, G. van den Driessche, Y. Blau, C. Malik, W. Liang, R. Hirsch, C. N. dos Santos, E. Weinstein, A. van den Oord, S. Lall, N. FitzGerald, Z. Jiang, X. Yang, D. Webster, A. Elqursh, A. Pope, G. Rotival, D. Raposo, W. Zhu, J. Dean, S. Alabed, D. Tran, A. Gupta, Z. Gleicher, J. Austin, E. Rosseel, M. Umekar, D. Das, Y. Sun, K. Chen, K. Misiunas, X. Zhou, Y. Di, A. Loo, J. Newlan, B. Li, V. Ramasesh, Y. Xu, A. Chen, S. Gandhe, R. Soricut, N. Gupta, S. Hu, S. El-Sayed, X. Garcia, I. Brusilovsky, P. Chen, A. Bolt, L. Huang, A. Gurney, Z. Zhang, A. Pritzel, J. Wilkiewicz, B. Seybold, B. K. Shamanna, F. Fischer, J. Dean, K. Gill, R. Mcilroy, A. Bhowmick, J. Selier, A. Yang, D. Cheng, V. Magay, J. Tan, D. Varma, C. Walder, T. Kocisky, R. Nakashima, P. Natsev, M. Kwong, I. Gog, C. Zhang, S. Dieleman, T. Jimma, A. Ryabtsev, S. Brahma, D. Steiner, D. Du, A. ĆœuĆŸul, M. ĆœaniÄ, M. Raghavachari, W. Gierke, Z. Zheng, D. Petrova, Y. Dauphin, Y. Liu, I. Kessler, S. Hand, C. Duvarney, S. Kim, H. Lee, L. Hussenot, J. Hui, J. Smith, D. Jain, J. Xia, G. S. Tomar, K. Amiri, D. Phan, F. Fuchs, T. Weyand, N. Tomasev, A. Cordell, X. Liu, J. Mallinson, P. Joshi, A. Crawford, A. Suggala, S. Chien, N. Fernando, M. Sanchez-Vargas, D. Williams, P. Crone, X. Luo, I. Karpov, J. Shan, T. Thurk, R. Strudel, P. Voigtlaender, P. Patil, T. Dozat, A. Khodaei, S. Singla, P. Ambroszczyk, Q. Wu, Y. Chang, B. Roark, C. Hegde, T. Ding, A. Filos, Z. Wu, A. S. Pinto, S. Liu, S. Khanna, A. Pandey, S. Mcloughlin, Q. Li, S. Haves, A. Zhou, E. Buchatskaya, I. Leal, P. de Boursac, N. Akazawa, N. Anderson, T. Chen, K. Somandepalli, C. Liang, S. Goenka, S. Winkler, A. Grushetsky, Y. Ding, J. Smith, F. Ye, J. Pont-Tuset, E. Li, R. Li, T. Golany, D. Wegner, T. Jiang, O. Barak, Y. Shangguan, E. VĂ©rtes, R. Wong, J. Bornschein, A. Tudor, M. Bevilacqua, T. Schaul, A. S. Rawat, Y. Zhao, K. Axiotis, L. Meng, C. McLean, J. Lai, J. Beattie, N. Kushman, Y. Liu, B. Kutzman, F. Lang, J. Ye, P. Netrapalli, P. Mishra, M. Khan, M. Goel, R. Willoughby, D. Tian, H. Zhuang, J. Chen, Z. Tsai, T. Kementsietsidis, A. Khare, J. Keeling, K. Xu, N. Waters, F. AltchĂ©, A. Popat, B. Mittal, D. Saxton, D. E. Badawy, M. Mathieu, Z. Zheng, H. Zhou, N. Ranka, R. Shin, Q. Duan, T. Salimans, I. Mihailescu, U. Shaham, M. Chang, Y. Assael, N. Dikkala, M. Izzard, V. Cohen-Addad, C. Graves, V. Feinberg, G. Chung, D. Strouse, D. Karmon, S. Sharifzadeh, Z. Ashwood, K. Pham, J. Blanton, A. Vasiloff, J. Barber, M. Geller, A. Zhou, F. Zubach, T. Huang, L. Zhang, H. Gupta, M. Young, J. Proskurnia, R. Votel, V. Gabeur, G. Barcik, A. Tripathi, H. Yu, G. Yan, B. Changpinyo, F. PavetiÄ, A. Coyle, Y. Fujii, J. G. Mendez, T. Zhou, H. Rajamani, B. Hechtman, E. Cao, D. Juan, Y. Tan, V. Dalibard, Y. Du, N. Clay, K. Yao, W. Jia, D. Vijaykumar, Y. Zhou, X. Bai, W. Hung, S. Pecht, G. Todorov, N. Khadke, P. Gupta, P. Lahoti, A. Autef, K. Duddu, J. Lee-Thorp, A. Bykovsky, T. Misiunas, S. Flennerhag, S. Thangaraj, J. McGiffin, Z. Nado, M. Kunesch, A. Noever, A. Hertz, M. Liang, V. Stone, E. Palmer, S. Daruki, A. Pramanik, S. PĂ”der, A. Kyker, M. Khan, E. Sluzhaev, M. Ritter, A. Ruderman, W. Zhou, C. Nagpal, K. Vodrahalli, G. Necula, P. Barham, E. Pavlick, J. Hartford, I. Shafran, L. Zhao, M. MikuĆa, T. Eccles, H. Shimokawa, K. Garg, L. Vilnis, H. Chen, I. Shumailov, K. Lee, A. Abdelhamed, M. Xie, V. Cohen, E. Hlavnova, D. Malkin, C. Sitawarin, J. Lottes, P. Coquinot, T. Yu, S. Kumar, J. Zhang, A. Mahendru, Z. Ahmed, J. Martens, T. Chen, A. Boag, D. Peng, C. Devin, A. Klimovskiy, M. Phuong, D. Vainstein, J. Xie, B. Ramabhadran, N. Howard, X. Yu, G. Goswami, J. Cui, S. Shleifer, M. Pinto, C. Yeh, M. Yang, S. Javanmardi, D. Ethier, C. Lee, J. Orbay, S. Kotecha, C. Bromberg, P. Shaw, J. Thornton, A. G. Rosenthal, S. Gu, M. Thomas, I. Gemp, A. Ayyar, A. Ushio, A. Selvan, J. Wee, C. Liu, M. Majzoubi, W. Yu, J. Abernethy, T. Liechty, R. Pan, H. Nguyen, Qiong, Hu, S. Perrin, A. Arora, E. Pitler, W. Wang, K. Shivakumar, F. Prost, B. Limonchik, J. Wang, Y. Gao, T. Cour, S. Buch, H. Gui, M. Ivanova, P. Neubeck, K. Chan, L. Kim, H. Chen, N. Goyal, D. Chung, L. Liu, Y. Su, A. Petrushkina, J. Shen, A. Joulin, Y. Xu, S. X. Lin, Y. Kulizhskaya, C. Chelba, S. Vasudevan, E. Collins, V. Bashlovkina, T. Lu, D. Fritz, J. Park, Y. Zhou, C. Su, R. Tanburn, M. Sushkov, M. Rasquinha, J. Li, J. Prendki, Y. Li, P. LV, S. Sharma, H. Fitoussi, H. Huang, A. Dai, P. Dao, M. Burrows, H. Prior, D. Qin, G. Pundak, L. L. Sjoesund, A. Khurshudov, Z. Zhu, A. Webson, E. Kemp, T. Tan, S. Agrawal, S. Sargsyan, L. Cheng, J. Stephan, T. Kwiatkowski, D. Reid, A. Byravan, A. H. Michaely, N. Heess, L. Zhou, S. Goenka, V. Carpenter, A. Levskaya, B. Wang, R. Roberts, R. Leblond, S. Chikkerur, S. Ginzburg, M. Chang, R. Riachi, Chuqiao, Xu, Z. Borsos, M. Pliskin, J. Pawar, M. Lustman, H. Kirkwood, A. Anand, A. Chaudhary, N. Kalb, K. Milan, S. Augenstein, A. Goldie, L. Prince, K. Raman, Y. Sun, V. Xia, A. Cohen, Z. Huo, J. Camp, S. Ellis, L. Zilka, D. V. Torres, L. Patel, S. Arora, B. Chan, J. Adler, K. Ayoub, J. Liang, F. Jamil, J. Jiang, S. Baumgartner, H. Sun, Y. Karov, Y. Akulov, H. Zheng, I. Cai, C. Fantacci, J. Rubin, A. R. Acha, M. Wang, N. DâSouza, R. Sathyanarayana, S. Dai, S. Rowe, A. Simanovsky, O. Goldman, Y. Kuang, X. Pan, A. Rosenberg, T. Rojas-Esponda, P. Dutta, A. Zeng, I. Jurenka, G. Farquhar, Y. Bansal, S. Iqbal, B. Roelofs, G. Joung, P. Beak, C. Ryu, R. Poplin, Y. Wu, J. Alayrac, S. Buthpitiya, O. Ronneberger, C. Habtegebriel, W. Li, P. Cavallaro, A. Wei, G. Bensky, T. Denk, H. Ganapathy, J. Stanway, P. Joshi, F. Bertolini, J. Lo, O. Ma, Z. Charles, G. Sampemane, H. Sahni, X. Chen, H. Askham, D. Gaddy, P. Young, J. Tan, M. Eyal, A. BraĆŸinskas, L. Zhong, Z. Wu, M. Epstein, K. Bailey, A. Hard, K. Lee, S. Goldshtein, A. Ruiz, M. Badawi, M. Lochbrunner, J. Kearns, A. Brown, F. Pardo, T. Weber, H. Yang, P. Jiang, B. Akin, Z. Fu, M. Wainwright, C. Zou, M. Gaba, P. Manzagol, W. Kan, Y. Song, K. Zainullina, R. Lin, J. Ko, S. Deshmukh, A. Jindal, J. Svensson, D. Tyam, H. Zhao, C. Kaeser-Chen, S. Baird, P. Moradi, J. Hall, Q. Guo, V. Tsang, B. Liang, F. Pereira, S. Ganesh, I. Korotkov, J. Adamek, S. Thiagarajan, V. Tran, C. Chen, C. Tar, S. Jain, I. Dasgupta, T. Bilal, D. Reitter, K. Zhao, G. Vezzani, Y. Gehman, P. Mehta, L. Beltrone, X. Dotiwalla, S. Guadarrama, Z. Abbas, S. Karp, P. Georgiev, C. Ferng, M. Brockschmidt, L. Peng, C. Hirnschall, V. Verma, Y. Bi, Y. Xiao, A. Dabush, K. Xu, P. Wallis, R. Parker, Q. Wang, Y. Xu, I. Safarli, D. Tewari, Y. Zhang, S. Kim, A. Gesmundo, M. Thomas, S. Levi, A. Chowdhury, K. Rao, P. Garst, S. Conway-Rahman, H. Ran, K. McKinney, Z. Xiao, W. Yu, R. Agrawal, A. Stjerngren, C. Ionescu, J. Chen, V. Sharma, J. Chiu, F. Liu, K. Franko, C. Sanford, X. Cai, P. Michel, S. Ganapathy, J. Labanowski, Z. Garrett, B. Vargas, S. Sun, B. Gale, T. Buschmann, G. Desjardins, N. Ghelani, P. Jain, M. Verma, C. Asawaroengchai, J. Eisenschlos, J. Harlalka, H. Kazawa, D. Metzler, J. Howland, Y. Jian, J. Ades, V. Shah, T. Gangwani, S. Lee, R. Ring, S. M. Hernandez, D. Reich, A. Sinha, A. Sathe, J. Kovac, A. Gill, A. Kannan, A. Dâolimpio, M. Sevenich, J. Whang, B. Kim, K. C. Sim, J. Chen, J. Zhang, S. Lall, Y. Matias, B. Jia, A. Friesen, S. Nasso, A. Thapliyal, B. Perozzi, T. Yu, A. Shekhawat, S. Huda, P. Grabowski, E. Wang, A. Sreevatsa, H. Dib, M. Hassen, P. Schuh, V. Milutinovic, C. Welty, M. Quinn, A. Shah, B. Wang, G. Barth-Maron, J. Frye, N. Axelsson, T. Zhu, Y. Ma, I. Giannoumis, H. Sedghi, C. Ye, Y. Luan, K. Aydin, B. Chandra, V. Sampathkumar, R. Huang, V. Lavrenko, A. Eleryan, Z. Hong, S. Hansen, S. M. Carthy, B. Samanta, D. Äevid, X. Wang, F. Li, M. Voznesensky, M. Hoffman, A. Terzis, V. Sehwag, G. Fidel, L. He, M. Cai, Y. He, A. Feng, M. Nikoltchev, S. Phatale, J. Chase, R. Lawton, M. Zhang, T. Ouyang, M. Tragut, M. H. Manshadi, A. Narayanan, J. Shen, X. Gao, T. Bolukbasi, N. Roy, X. Li, D. Golovin, L. Panait, Z. Qin, G. Han, T. Anthony, S. Kudugunta, V. Patraucean, A. Ray, X. Chen, X. Yang, T. Bhatia, P. Talluri, A. Morris, A. RaĆŸnatoviÄ, B. Brownfield, J. An, S. Peng, P. Kane, C. Zheng, N. Duduta, J. Kessinger, J. Noraky, S. Liu, K. Rong, P. VeliÄkoviÄ, K. Rush, A. Goldin, F. Wei, S. M. R. Garlapati, C. Pantofaru, O. Kwon, J. Ni, E. Noland, J. D. Trapani, F. Beaufays, A. G. Roy, Y. Chow, A. Turker, G. Cideron, L. Mei, J. Clark, Q. Dou, M. BoĆĄnjak, R. Leith, Y. Du, A. Yazdanbakhsh, M. Nasr, C. Kwak, S. S. Sheth, A. Kaskasoli, A. Anand, B. Lakshminarayanan, S. Jerome, D. Bieber, C. Chu, A. Senges, T. Shen, M. Sridhar, N. Ndebele, B. Beyret, S. Mohamed, M. Chen, M. Freitag, J. Guo, L. Liu, P. Roit, H. Chen, S. Yan, T. Stone, J. Co-Reyes, J. Cole, S. Scellato, S. Azizi, H. Hashemi, A. Jin, A. Iyer, M. Valentine, A. György, A. Ahuja, D. H. Diaz, C. Lee, N. Clement, W. Kong, D. Garmon, I. Watts, K. Bhatia, K. Gupta, M. Miecnikowski, H. Vallet, A. Taly, E. Loper, S. Joshi, J. Atwood, J. Chick, M. Collier, F. Iliopoulos, R. Trostle, B. Gunel, R. Leal-Cavazos, A. M. Hrafnkelsson, M. Guzman, X. Ju, A. Forbes, J. Emond, K. Chauhan, B. Caine, L. Xiao, W. Zeng, A. Moufarek, D. Murphy, M. Meng, N. Gupta, F. Riedel, A. Das, E. Lawal, S. Narayan, T. Sosea, J. Swirhun, L. Friso, B. Neyshabur, J. Lu, S. Girgin, M. Wunder, E. Yvinec, A. Pyne, V. Carbune, S. Rijhwani, Y. Guo, T. Doshi, A. Briukhov, M. Bain, A. Hitron, X. Wang, A. Gupta, K. Chen, C. Du, W. Zhang, D. Shah, A. Akula, M. Dylla, A. Kachra, W. Kuo, T. Zou, L. Wang, L. Xu, J. Zhu, J. Snyder, S. Menon, O. Firat, I. Mordatch, Y. Yuan, N. Ponomareva, R. Blevins, L. Moore, W. Wang, P. Chen, M. Scholz, A. Dwornik, J. Lin, S. Li, D. Antognini, T. I, X. Song, M. Miller, U. Kalra, A. Raveret, O. Akerlund, F. Wu, A. Nystrom, N. Godbole, T. Liu, H. DeBalsi, J. Zhao, B. Liu, A. Caciularu, L. Lax, U. Khandelwal, V. Langston, E. Bailey, S. Lattanzi, Y. Wang, N. Kovelamudi, S. Mondal, G. Guruganesh, N. Hua, O. Roval, P. WesoĆowski, R. Ingale, J. Halcrow, T. Sohn, C. Angermueller, B. Raad, E. Stickgold, E. Lu, A. Kosik, J. Xie, T. Lillicrap, A. Huang, L. L. Zhang, D. Paulus, C. Farabet, A. Wertheim, B. Wang, R. Joshi, C. Ko, Y. Wu, S. Agrawal, L. Lin, X. Sheng, P. Sung, T. Breland-King, C. Butterfield, S. Gawde, S. Singh, Q. Zhang, R. Apte, S. Shetty, A. Hutter, T. Li, E. Salesky, F. Lebron, J. Kanerva, M. Paganini, A. Nguyen, R. Vallu, J. Peter, S. Velury, D. Kao, J. Hoover, A. Bortsova, C. Bishop, S. Jakobovits, A. Agostini, A. Agarwal, C. Liu, C. Kwong, S. Tavakkol, I. Bica, A. Greve, A. GP, J. Marcus, L. Hou, T. Duerig, R. Moroshko, D. Lacey, A. Davis, J. Amelot, G. Wang, F. Kim, T. Strinopoulos, H. Wan, C. L. Lan, S. Krishnan, H. Tang, P. Humphreys, J. Bai, I. H. Shtacher, D. Machado, C. Pang, K. Burke, D. Liu, R. Aravamudhan, Y. Song, E. Hirst, A. Singh, B. Jou, L. Bai, F. Piccinno, C. K. Fu, R. Alazard, B. Meiri, D. Winter, C. Chen, M. Zhang, J. Heitkaemper, J. Lambert, J. Lee, A. Frömmgen, S. Rogulenko, P. Nair, P. Niemczyk, A. Bulyenov, B. Xu, H. Shemtov, M. Zadimoghaddam, S. Toropov, M. Wirth, H. Dai, S. Gollapudi, D. Zheng, A. Kurakin, C. Lee, K. Bullard, N. Serrano, I. Balazevic, Y. Li, J. Schalkwyk, M. Murphy, M. Zhang, K. Sequeira, R. Datta, N. Agrawal, C. Sutton, N. Attaluri, M. Chiang, W. Farhan, G. Thornton, K. Lin, T. Choma, H. Nguyen, K. Dasgupta, D. Robinson, I. ComĆa, M. Riley, A. Pillai, B. Mustafa, B. Golan, A. Zandieh, J. Lespiau, B. Porter, D. Ross, S. Rajayogam, M. Agarwal, S. Venugopalan, B. Shahriari, Q. Yan, H. Xu, T. Tobin, P. Dubov, H. Shi, A. Recasens, A. Kovsharov, S. Borgeaud, L. Dery, S. Vasanth, E. Gribovskaya, L. Qiu, M. Mahdieh, W. Skut, E. Nielsen, C. Zheng, A. Yu, C. G. Bostock, S. Gupta, A. Archer, C. Rawles, E. Davies, A. Svyatkovskiy, T. Tsai, Y. Halpern, C. Reisswig, B. Wydrowski, B. Chang, J. Puigcerver, M. H. Taege, J. Li, E. Schnider, X. Li, D. Dena, Y. Xu, U. Telang, T. Shi, H. Zen, K. Kastner, Y. Ko, N. Subramaniam, A. Kumar, P. Blois, Z. Dai, J. Wieting, Y. Lu, Y. Zeldes, T. Xie, A. Hauth, A. Ćąifrea, Y. Li, S. El-Husseini, D. Abolafia, H. Zhou, W. Ding, S. Ghalebikesabi, C. GuĂa, A. Maksai, Ă. Weisz, S. Arik, N. Sukhanov, A. Ćwietlik, X. Jia, L. Yu, W. Wang, M. Brand, D. Bloxwich, S. Kirmani, Z. Chen, A. Go, P. Sprechmann, N. Kannen, A. Carin, P. Sandhu, I. Edkins, L. Nooteboom, J. Gupta, L. Maggiore, J. Azizi, Y. Pritch, P. Yin, M. Gupta, D. Tarlow, D. Smith, D. Ivanov, M. Babaeizadeh, A. Goel, S. Kambala, G. Chu, M. Kastelic, M. Liu, H. Soltau, A. Stone, S. Agrawal, M. Kim, K. Soparkar, S. Tadepalli, O. Bunyan, R. Soh, A. Kannan, D. Kim, B. J. Chen, A. Halumi, S. Roy, Y. Wang, O. Sercinoglu, G. Gibson, S. Bhatnagar, M. Sano, D. von Dincklage, Q. Ren, B. Mitrevski, M. OlĆĄĂĄk, J. She, C. Doersch, Jilei, Wang, B. Liu, Q. Tan, T. Yakar, T. Warkentin, A. Ramirez, C. Lebsack, J. Dillon, R. Mathews, T. Cobley, Z. Wu, Z. Chen, J. Simon, S. Nath, T. Sainath, A. Bendebury, R. Julian, B. Mankalale, D. Äurko, P. Zacchello, A. R. Brown, K. Sodhia, H. Howard, S. Caelles, A. Gupta, G. Evans, A. Bulanova, L. Katzen, R. Goldenberg, A. Tsitsulin, J. Stanton, B. Schillings, V. Kovalev, C. Fry, R. Shah, K. Lin, S. Upadhyay, C. Li, S. Radpour, M. Maggioni, J. Xiong, L. Haas, J. Brennan, A. Kamath, N. Savinov, A. Nagrani, T. Yacovone, R. Kappedal, K. Andriopoulos, L. Lao, Y. Li, G. Rozhdestvenskiy, K. Hashimoto, A. Audibert, S. Austin, D. Rodriguez, A. Ruoss, G. Honke, D. Karkhanis, X. Xiong, Q. Wei, J. Huang, Z. Leng, V. Premachandran, S. Bileschi, G. Evangelopoulos, T. Mensink, J. Pavagadhi, D. Teplyashin, P. Chang, L. Xue, G. Tanzer, S. Goldman, K. Patel, S. Li, J. Wiesner, I. Zheng, I. Stewart-Binks, J. Han, Z. Li, L. Luo, K. Lenc, M. LuÄiÄ, F. Xue, R. Mullins, A. Guseynov, C. Chang, I. Galatzer-Levy, A. Zhang, G. Bingham, G. Hu, A. Hartman, Y. Ma, J. Griffith, A. Irpan, C. Radebaugh, S. Yue, L. Fan, V. Ungureanu, C. Sorokin, H. Teufel, P. Li, R. Anil, D. Paparas, T. Wang, C. Lin, H. Peng, M. Shum, G. Petrovic, D. Brady, R. Nguyen, K. Macherey, Z. Li, H. Singh, M. Yenugula, M. Iinuma, X. Chen, K. Kopparapu, A. Stern, S. Dave, C. Thekkath, F. Perot, A. Kumar, F. Li, Y. Xiao, M. Bilotti, M. H. Bateni, I. Noble, L. Lee, A. VĂĄzquez-Reina, J. Salazar, X. Yang, B. Wang, E. Gruzewska, A. Rao, S. Raghuram, Z. Xu, E. Ben-David, J. Mei, S. Dalmia, Z. Zhang, Y. Liu, G. Bansal, H. Pankov, S. Schwarcz, A. Burns, C. Chan, S. Sanghai, R. Liang, E. Liang, A. He, A. Stuart, A. Narayanan, Y. Zhu, C. Frank, B. Fatemi, A. Sabne, O. Lang, I. Bhattacharya, S. Settle, M. Wang, B. McMahan, A. Tacchetti, L. B. Soares, M. Hadian, S. Cabi, T. Chung, N. Putikhin, G. Li, J. Chen, A. Tarango, H. Michalewski, M. Kazemi, H. Masoom, H. Sheftel, R. Shivanna, A. Vadali, R. Comanescu, D. Reid, J. Moore, A. Neelakantan, M. Sander, J. Herzig, A. Rosenberg, M. Dehghani, J. Choi, M. Fink, R. Hayes, E. Ge, S. Weng, C. Ho, J. Karro, K. Krishna, L. N. Thiet, A. Skerry-Ryan, D. Eppens, M. Andreetto, N. Sarma, S. Bonacina, B. K. Ayan, M. Nawhal, Z. Shan, M. Dusenberry, S. Thakoor, S. Gubbi, D. D. Nguyen, R. Tsarfaty, S. Albanie, J. MitroviÄ, M. Gandhi, B. Chen, A. Epasto, G. Stephanov, Y. Jin, S. Gehman, A. Amini, J. Weber, F. Behbahani, S. Xu, M. Allamanis, X. Chen, M. Ott, C. Sha, M. Jastrzebski, H. Qi, D. Greene, X. Wu, A. Toki, D. Vlasic, J. Shapiro, R. Kotikalapudi, Z. Shen, T. Saeki, S. Xie, A. Cassirer, S. Bharadwaj, T. Kiyono, S. Bhojanapalli, E. Rosenfeld, S. Ritter, J. Mao, J. G. Oliveira, Z. Egyed, B. Bandemer, E. Parisotto, K. Kinoshita, J. Pluto, P. Maniatis, S. Li, Y. Guo, G. Ghiasi, J. Tarbouriech, S. Chatterjee, J. Jin, Katrina, Xu, J. Palomaki, S. Arnold, M. Sewak, F. Piccinini, M. Sharma, B. Albrecht, S. Purser-haskell, A. Vaswani, C. Chen, M. Wisniewski, Q. Cao, J. Aslanides, N. M. Phu, M. Sieb, L. Agubuzu, A. Zheng, D. Sohn, M. Selvi, A. Andreassen, K. Subudhi, P. Eruvbetine, O. Woodman, T. Mery, S. Krause, X. Ren, X. Ma, J. Luo, D. Chen, W. Fan, H. Griffiths, C. Schuler, A. Li, S. Zhang, J. Sarr, S. Luo, R. Patana, M. Watson, D. Naboulsi, M. Collins, S. Sidhwani, E. Hoogeboom, S. Silver, E. Caveness, X. Zhao, M. Rodriguez, M. Deines, L. Bai, P. Griffin, M. Tagliasacchi, E. Xue, S. R. Babbula, B. Pang, N. Ding, G. Shen, E. Peake, R. Crocker, S. S. Raghvendra, D. Swisher, W. Han, R. Singh, L. Wu, V. Pchelin, T. Munkhdalai, D. Alon, G. Bacon, E. Robles, J. Bulian, M. Johnson, G. Powell, F. T. Ferreira, Y. Li, F. Benzing, M. VelimiroviÄ, H. Soyer, W. Kong, Tony, NguyĂȘn, Z. Yang, J. Liu, J. van Amersfoort, D. Gillick, B. Sun, N. Rauschmayr, K. Zhang, S. Zhan, T. Zhou, A. Frolov, C. Yang, D. Vnukov, L. Rouillard, H. Li, A. Mandhane, N. Fallen, R. Venkataraman, C. H. Hu, J. Brennan, J. Lee, J. Chang, M. Sundermeyer, Z. Pan, R. Ke, S. Tong, A. Fabrikant, W. Bono, J. Gu, R. Foley, Y. Mao, M. Delakis, D. Bhaswar, R. Frostig, N. Li, A. Zipori, C. Hope, O. Kozlova, S. Mishra, J. Djolonga, C. Schiff, M. A. Merey, E. Briakou, P. Morgan, A. Wan, A. Hassidim, R. Skerry-Ryan, K. Sengupta, M. Jasarevic, P. Kallakuri, P. Kunkle, H. Brennan, T. Lieber, H. Mansoor, J. Walker, B. Zhang, A. Xie, G. ĆœuĆŸiÄ, A. Chukwuka, A. Druinsky, D. Cho, R. Yao, F. Naeem, S. Butt, E. Kim, Z. Jia, M. Jordan, A. Lelkes, M. Kurzeja, S. Wang, J. Zhao, A. Over, A. Chakladar, M. Prasetya, N. Jha, S. Ganapathy, Y. Cong, P. Shroff, C. Saroufim, S. Miryoosefi, M. Hammad, T. Nasir, W. Xi, Y. Gao, Y. Maeng, B. Hora, C. Cheng, P. Haghani, Y. Lewenberg, C. Lu, M. Matysiak, N. Raisinghani, H. Wang, L. Baugher, R. Sukthankar, M. Giang, J. Schultz, N. Fiedel, M. Chen, C. Lee, T. Dey, H. Zheng, S. Paul, C. Smith, A. Ly, Y. Wang, R. Bansal, B. Perz, S. Ricco, S. Blank, V. Keshava, D. Sharma, M. Chow, K. Lad, K. Jalan, S. Osindero, C. Swanson, J. Scott, A. IliÄ, X. Li, S. R. Jonnalagadda, A. S. Soudagar, Y. Xiong, B. Batsaikhan, D. Jarrett, N. Kumar, M. Shah, M. Lawlor, A. Waters, M. Graham, R. May, S. Ramos, S. Lefdal, Z. Cankara, N. Cano, B. OâDonoghue, J. Borovik, F. Liu, J. Grimstad, M. Alnahlawi, K. Tsihlas, T. Hudson, N. Grigorev, Y. Jia, T. Huang, T. P. Igwe, S. Lebedev, X. Tang, I. Krivokon, F. Garcia, M. Tan, E. Jia, P. Stys, S. Vashishth, Y. Liang, B. Venkatraman, C. Gu, A. Kementsietsidis, C. Zhu, J. Jung, Y. Bai, M. J. Hosseini, F. Ahmed, A. Gupta, X. Yuan, S. Ashraf, S. Nigam, G. Vasudevan, P. Awasthi, A. M. Gilady, Z. Mariet, R. Eskander, H. Li, H. Hu, G. Garrido, P. Schlattner, G. Zhang, R. Saxena, P. DeviÄ, K. Muralidharan, A. Murthy, Y. Zhou, M. Choi, A. Wongpanich, Z. Wang, P. Shah, Y. Xu, Y. Huang, S. Spencer, A. Chen, J. Cohan, J. Wang, J. Tompson, J. Wu, R. Haroun, H. Li, B. Huergo, F. Yang, T. Yin, J. Wendt, M. Bendersky, R. Chaabouni, J. Snaider, J. Ferret, A. Jindal, T. Thompson, A. Xue, W. Bishop, S. M. Phal, A. Sharma, Y. Sung, P. Radhakrishnan, M. Shomrat, R. Ingle, R. Vij, J. Gilmer, M. D. Istin, S. Sobell, Y. Lu, E. Nottage, D. Sadigh, J. Willcock, T. Zhang, S. Xu, S. Brown, K. Lee, G. Wang, Y. Zhu, Y. Tay, C. Kim, A. Gutierrez, A. Sharma, Y. Xian, S. Seo, C. Cui, E. Pochernina, C. Baetu, K. JastrzÄbski, M. Ly, M. Elhawaty, D. Suh, E. Sezener, P. Wang, N. Yuen, G. Tucker, J. Cai, Z. Yang, C. Wang, A. Muzio, H. Qian, J. Yoo, D. Lockhart, K. R. McKee, M. Guo, M. Mehrotra, A. Mendonça, S. V. Mehta, S. Ben, C. Tekur, J. Mu, M. Zhu, V. Krakovna, H. Lee, A. Maschinot, S. Cevey, H. Choe, A. Bai, H. Srinivasan, D. Gasaway, N. Young, P. Siegler, D. Holtmann-Rice, V. Piratla, K. Baumli, R. Yogev, A. Hofer, H. van Hasselt, S. Grant, Y. Chervonyi, D. Silver, A. Hogue, A. Agarwal, K. Wang, P. Singh, F. Flynn, J. Lipschultz, R. David, L. Bellot, Y. Yang, L. Le, F. Graziano, K. Olszewska, K. Hui, A. Maurya, N. Parotsidis, W. Chen, T. Oguntebi, J. Kelley, A. Baddepudi, J. Mauerer, G. Shaw, A. Siegman, L. Yang, S. Shetty, S. Roy, Y. Song, W. Stokowiec, R. Burnell, O. Savant, R. Busa-Fekete, J. Miao, S. Ghosh, L. MacDermed, P. Lippe, M. Dektiarev, Z. Behrman, F. Mentzer, K. Nguyen, M. Wei, S. Verma, C. Knutsen, S. Dasari, Z. Yan, P. Mitrichev, X. Wang, V. Shejwalkar, J. Austin, S. Sunkara, N. Potti, Y. Virin, C. Wright, G. Liu, O. Riva, E. Pot, G. Kochanski, Q. Le, G. Balasubramaniam, A. Dhar, Y. Liao, A. Bloniarz, D. Shukla, E. Cole, J. Lee, S. Zhang, S. Kafle, S. Vashishtha, P. Mahmoudieh, G. Chen, R. Hoffmann, P. Srinivasan, A. D. Lago, Y. B. Shalom, Z. Wang, M. Elabd, A. Sharma, J. Oh, S. Kothawade, M. Le, M. Monteiro, S. Yang, K. Alarakyia, R. Geirhos, D. Mincu, H. Garnes, H. Kobayashi, S. Mariooryad, K. Krasowiak, Zhixin, Lai, S. Mourad, M. Wang, F. Bu, O. Aharoni, G. Chen, A. Goyal, V. Zubov, A. Bapna, E. Dabir, N. Kothari, K. Lamerigts, N. D. Cao, J. Shar, C. Yew, N. Kulkarni, D. Mahaarachchi, M. Joshi, Z. Zhu, J. Lichtarge, Y. Zhou, H. Muckenhirn, V. Selo, O. Vinyals, P. Chen, A. Brohan, V. Mehta, S. Cogan, R. Wang, T. Geri, W. Ko, W. Chen, F. Viola, K. Shivam, L. Wang, M. C. Elish, R. A. Popa, S. Pereira, J. Liu, R. Koster, D. Kim, G. Zhang, S. Ebrahimi, P. Talukdar, Y. Zheng, P. Poklukar, A. Mikhalap, D. Johnson, A. Vijayakumar, M. Omernick, M. Dibb, A. Dubey, Q. Hu, A. Suman, V. Aggarwal, I. Kornakov, F. Xia, W. Lowe, A. Kolganov, T. Xiao, V. Nikolaev, S. Hemingray, B. Li, J. Iljazi, M. RybiĆski, B. Sandhu, P. Lu, T. Luong, R. Jenatton, V. Govindaraj, Hui, Li, G. Dulac-Arnold, W. Park, H. Wang, A. Modi, J. Pouget-Abadie, K. Greller, R. Gupta, R. Berry, P. Ramachandran, J. Xie, L. McCafferty, J. Wang, K. Gupta, H. Lim, B. BrataniÄ, A. Brock, I. Akolzin, J. Sproch, D. Karliner, D. Kim, A. Goedeckemeyer, N. Shazeer, C. Schmid, D. Calandriello, P. Bhatia, K. Choromanski, C. Montgomery, D. Dua, A. Ramalho, H. King, Y. Gao, L. Nguyen, D. Lindner, D. Pitta, O. Johnson, K. Salama, D. Ardila, M. Han, E. Farnese, S. Odoom, Z. Wang, X. Ding, N. Rink, R. Smith, H. T. Lehri, E. Cohen, N. Vats, T. He, P. Gopavarapu, A. Paszke, M. Patel, W. V. Gansbeke, L. Loher, L. Castro, M. Voitovich, T. von Glehn, N. George, S. Niklaus, Z. Eaton-Rosen, N. RakiÄeviÄ, E. Jue, S. Perel, C. Zhang, Y. Bahat, A. Pouget, Z. Xing, F. Huot, A. Shenoy, T. Bos, V. Coriou, B. Richter, N. Noy, Y. Wang, S. Ontanon, S. Qin, G. Makarchuk, D. Hassabis, Z. Li, M. Sharma, K. Venkatesan, I. Kemaev, R. Daniel, S. Huang, S. Shah, O. Ponce, Warren, Chen, M. Faruqui, J. Wu, S. AndaÄiÄ, S. Payrits, D. McDuff, T. Hume, Y. Cao, M. Tessler, Q. Wang, Y. Wang, I. Rendulic, E. Agustsson, M. Johnson, T. Lando, A. Howard, S. G. S. Padmanabhan, M. Daswani, A. Banino, M. Kilgore, J. Heek, Z. Ji, A. Caceres, C. Li, N. Kassner, A. Vlaskin, Z. Liu, A. Grills, Y. Hou, R. Sukkerd, G. Cheon, N. Shetty, L. Markeeva, P. Stanczyk, T. Iyer, Y. Gong, S. Gao, K. Gopalakrishnan, T. Blyth, M. Reynolds, A. Bhoopchand, M. Bilenko, D. Gharibian, V. Zayats, A. Faust, A. Singh, M. Ma, H. Jiao, S. Vijayanarasimhan, L. Aroyo, V. Yadav, S. Chakera, A. Kakarla, V. Meshram, K. Gregor, G. Botea, E. Senter, D. Jia, G. Kovacs, N. Sharma, S. Baur, K. Kang, Y. He, L. Zhuo, M. Kostelac, I. Laish, S. Peng, L. OâBryan, D. Kasenberg, G. R. Rao, E. Leurent, B. Zhang, S. Stevens, A. Salazar, Y. Zhang, I. Lobov, J. Walker, A. Porter, M. Redshaw, H. Ke, A. Rao, A. Lee, H. Lam, M. Moffitt, J. Kim, S. Qiao, T. Koo, R. Dadashi, X. Song, M. Sundararajan, P. Xu, C. Kawamoto, Y. Zhong, C. Barbu, A. Reddy, M. Verzetti, L. Li, G. Papamakarios, H. Klimczak-PluciĆska, M. Cassin, K. Kavukcuoglu, R. Swavely, A. Vaucher, J. Zhao, R. Hemsley, M. Tschannen, H. Ge, G. Menghani, Y. Yu, N. Ha, W. He, X. Wu, M. Song, R. Sterneck, S. Zinke, D. A. Calian, A. Marsden, A. C. Ruiz, M. Hessel, A. Gueta, B. Lee, B. Farris, M. Gupta, Y. Li, M. Saleh, V. Misra, K. Xiao, P. Mendolicchio, G. Buttimore, V. Krayvanova, N. Nayakanti, M. Wiethoff, Y. Pande, A. Mirhoseini, N. Lao, J. Liu, Y. Hua, A. Chen, Y. Malkov, D. Kalashnikov, S. Gupta, K. Audhkhasi, Y. Zhai, S. Kopalle, P. Jain, E. Ofek, C. Meyer, K. Baatarsukh, H. StrejÄek, J. Qian, J. Freedman, R. Figueira, M. Sokolik, O. Bachem, R. Lin, D. Kharrat, C. Hidey, P. Xu, D. Duan, Y. Li, M. Ersoy, R. Everett, K. Cen, R. Santamaria-Fernandez, A. Taubenfeld, I. Mackinnon, L. Deng, P. Zablotskaia, S. Viswanadha, S. Goel, D. Yates, Y. Deng, P. Choy, M. Chen, A. Sinha, A. Mossin, Y. Wang, A. Szlam, S. Hao, P. K. Rubenstein, M. Toksoz-Exley, M. Aperghis, Y. Zhong, J. Ahn, M. Isard, O. Lacombe, F. Luisier, C. Anastasiou, Y. Kalley, U. Prabhu, E. Dunleavy, S. Bijwadia, J. Mao-Jones, K. Chen, R. Pasumarthi, E. Wood, A. Dostmohamed, N. Hurley, J. Simsa, A. Parrish, M. Pajarskas, M. Harvey, O. Skopek, Y. Kochinski, J. Rey, V. Rieser, D. Zhou, S. J. Lee, T. Acharya, G. Li, J. Jiang, X. Zhang, B. Gipson, E. Mahintorabi, M. Gelmi, N. Khajehnouri, A. Yeh, K. Lee, L. Matthey, L. Baker, T. Pham, H. Fu, A. Pak, P. Gupta, C. Vasconcelos, A. Sadovsky, B. Walker, S. Hsiao, P. Zochbauer, A. Marzoca, N. Velan, J. Zeng, G. Baechler, D. Driess, D. Jain, Y. Huang, L. Tao, J. Maggs, N. Levine, J. Schneider, E. Gemzer, S. Petit, S. Han, Z. Fisher, D. Zelle, C. Biles, E. Ie, A. Fadeeva, C. Liu, J. V. Franco, A. Collister, H. Zhang, R. Wang, R. Zhao, L. Kieliger, K. Shuster, R. Zhu, B. Gong, L. Chan, R. Sun, S. Basu, R. Zimmermann, J. Hayes, A. Bapna, J. Snoek, W. Yang, P. Datta, J. A. Abdallah, K. Kilgour, L. Li, S. Mah, Y. Jun, M. RiviĂšre, A. Karmarkar, T. Spalink, T. Huang, L. Gonzalez, D. Tran, A. Nowak, J. Palowitch, M. Chadwick, E. Talius, H. Mehta, T. Sellam, P. FrĂ€nken, M. Nicosia, K. He, A. Kini, D. Amos, S. Basu, H. Jobe, E. Shaw, Q. Xu, C. Evans, D. Ikeda, C. Yan, L. Jin, L. Wang, S. Yadav, I. Labzovsky, R. Sampath, A. Ma, C. Schumann, A. Siddhant, R. Shah, J. Youssef, R. Agarwal, N. Dabney, A. Tonioni, M. Ambar, J. Li, I. Guyon, B. Li, D. Soergel, B. Fang, G. Karadzhov, C. Udrescu, T. Trinh, V. Raunak, S. Noury, D. Guo, S. Gupta, M. Finkelstein, D. Petek, L. Liang, G. Billock, P. Sun, D. Wood, Y. Song, X. Yu, T. Matejovicova, R. Cohen, K. Andra, D. DâAmbrosio, Z. Deng, V. Nallatamby, E. Songhori, R. Dangovski, A. Lampinen, P. Botadra, A. Hillier, J. Cao, N. Baddi, A. Kuncoro, T. Yoshino, A. Bhagatwala, M. Ranzato, R. Schaeffer, T. Liu, S. Ye, O. Sarvana, J. Nham, C. Kuang, I. Gao, J. Baek, S. Mittal, A. Wahid, A. Gergely, B. Ni, J. Feldman, C. Muir, P. Lamblin, W. Macherey, E. Dyer, L. Kilpatrick, V. Campos, M. Bhutani, S. Fort, Y. Ahmad, A. Severyn, K. Chatziprimou, O. Ferludin, M. Dimarco, A. Kusupati, J. Heyward, D. Bahir, K. Villela, K. Millican, D. Marcus, S. Bahargam, C. Unlu, N. Roth, Z. Wei, S. Gopal, D. Ghoshal, E. Lee, S. Lin, J. Lees, D. Lee, A. Hosseini, C. Fan, S. Neel, M. Wu, Y. Altun, H. Cai, E. Piqueras, J. Woodward, A. Bissacco, S. Haykal, M. Bordbar, P. Sundaram, S. Hodkinson, D. Toyama, G. Polovets, A. Myers, A. Sinha, T. Levinboim, K. Krishnakumar, R. Chhaparia, T. Sholokhova, N. B. Gundavarapu, G. Jawahar, H. Qureshi, J. Hu, N. Momchev, M. Rahtz, R. Wu, A. P. S, K. Dhamdhere, M. Guo, U. Gupta, A. Eslami, M. Schain, M. Blokzijl, D. Welling, D. Orr, L. Bolelli, N. Perez-Nieves, M. Sirotenko, A. Prasad, A. Kar, B. D. B. Pigem, T. Terzi, G. Weisz, D. Ghosh, A. Mavalankar, D. Madeka, K. Daugaard, H. Adam, V. Shah, D. Berman, M. Tran, S. Baker, E. Andrejczuk, G. Chole, G. Raboshchuk, M. Mirzazadeh, T. Kagohara, S. Wu, C. Schallhart, B. Orlando, C. Wang, A. Rrustemi, H. Xiong, H. Liu, A. Vezer, N. Ramsden, S. Chang, S. Mudgal, Y. Li, N. Vieillard, Y. Hoshen, F. Ahmad, A. Slone, A. Hua, N. Potikha, M. Rossini, J. Stritar, S. Prakash, Z. Wang, X. Dong, A. Nazari, E. Nehoran, K. Tekelioglu, Y. Li, K. Badola, T. Funkhouser, Y. Li, V. Yerram, R. Ganeshan, D. Formoso, K. Langner, T. Shi, H. Li, Y. Yamamori, A. Panda, A. Saade, A. S. Scarpati, C. Breaux, C. Carey, Z. Zhou, C. Hsieh, S. Bridgers, A. Butryna, N. Gupta, V. Tulsyan, S. Woo, E. Eltyshev, W. Grathwohl, C. Parks, S. Benjamin, R. Panigrahy, S. Dodhia, D. D. Freitas, C. Sauer, W. Song, F. Alet, J. Tolins, C. Paduraru, X. Zhou, B. Albert, Z. Zhang, L. Shu, M. Bansal, S. Nguyen, A. Globerson, O. Xiao, J. Manyika, T. Hennigan, R. Rong, J. Matak, A. Bakalov, A. Sharma, D. Sinopalnikov, A. Pierson, S. Roller, G. Brown, M. Gao, T. Fukuzawa, A. Ghafouri, K. Vassigh, I. Barr, Z. Wang, A. Korsun, R. Jayaram, L. Ren, T. Zaman, S. Khan, Y. Lunts, D. Deutsch, D. Uthus, N. Katz, M. Samsikova, A. Khalifa, N. Sethi, J. Sun, L. Tang, U. Alon, X. Luo, D. Yu, A. Nayyar, B. Petrini, W. Truong, V. Hellendoorn, N. Chinaev, C. Alberti, W. Wang, J. Hu, V. Mirrokni, A. Balashankar, A. Aharon, A. Mehta, A. Iscen, J. Kready, L. Manning, A. Mohananey, Y. Chen, A. Tripathi, A. Wu, I. Petrovski, D. Hwang, M. Baeuml, S. Chandrakaladharan, Y. Liu, R. Coaguila, M. Chen, S. Ma, P. Tafti, S. Tatineni, T. Spitz, J. Ye, P. Vicol, M. Rosca, A. PuigdomĂšnech, Z. Yahav, S. Ghemawat, H. Lin, P. Kirk, Z. Nabulsi, S. Brin, B. Bohnet, K. Caluwaerts, A. S. Veerubhotla, D. Zheng, Z. Dai, P. Petrov, Y. Xu, R. Mehran, Z. Xu, L. Zintgraf, J. Choi, S. A. Hombaiah, R. Thoppilan, S. Reddi, L. Lew, L. Li, K. Webster, K. Sawhney, L. Lamprou, S. Shakeri, M. Lunayach, J. Chen, S. Bagri, A. Salcianu, Y. Chen, Y. Donchev, C. Magister, S. NĂžrly, V. Rodrigues, T. Izo, H. Noga, J. Zou, T. Köppe, W. Zhou, K. Lee, X. Long, D. Eisenbud, A. Chen, C. Schenck, C. M. To, P. Zhong, E. Taropa, M. Truong, O. Levy, D. Martins, Z. Zhang, C. Semturs, K. Zhang, A. Yakubovich, P. Moreno, L. McConnaughey, D. Lu, S. Redmond, L. Weerts, Y. Bitton, T. Refice, N. Lacasse, A. Conmy, C. Tallec, J. Odell, H. Forbes-Pollard, A. Socala, J. Hoech, P. Kohli, A. Walton, R. Wang, M. Sazanovich, K. Zhu, A. Kapishnikov, R. Galt, M. Denton, B. Murdoch, C. Sikora, K. Mohamed, W. Wei, U. First, T. McConnell, L. C. Cobo, J. Qin, T. Avrahami, D. Balle, Y. Watanabe, A. Louis, A. Kraft, S. Ariafar, Y. Gu, E. Rives, C. Yoon, A. Rusu, J. Cobon-Kerr, C. Hahn, J. Luo, Yuvein, Zhu, N. Ahuja, R. Benenson, R. L. Kaufman, H. Yu, L. Hightower, J. Zhang, D. Ni, L. A. Hendricks, G. Wang, G. Yona, L. Jain, P. Barrio, S. Bhupatiraju, S. Velusamy, A. Dafoe, S. Riedel, T. Thomas, Z. Yuan, M. Bellaiche, S. Panthaplackel, K. Kloboves, S. Jauhari, C. Akbulut, T. Davchev, E. Gladchenko, D. Madras, A. Chuklin, T. Hill, Q. Yuan, M. Madhavan, L. Leonhard, D. Scandinaro, Q. Chen, N. Niu, A. Douillard, B. Damoc, Y. Onoe, F. Pedregosa, F. Bertsch, C. Leichner, J. Pagadora, J. Malmaud, S. Ponda, A. Twigg, O. Duzhyi, J. Shen, M. Wang, R. Garg, J. Chen, U. Evci, J. Lee, L. Liu, K. Kojima, M. Yamaguchi, A. Rajendran, A. Piergiovanni, V. K. Rajendran, M. Fornoni, G. Ibagon, H. Ragan, S. M. Khan, J. Blitzer, A. Bunner, G. Sun, T. Kosakai, S. Lundberg, N. Elue, K. Guu, S. Park, J. Park, A. Narayanaswamy, C. Wu, J. Mudigonda, T. Cohn, H. Mu, R. Kumar, L. Graesser, Y. Zhang, R. Killam, V. Zhuang, M. GimĂ©nez, W. A. Jishi, R. Ley-Wild, A. Zhai, K. Osawa, D. Cedillo, J. Liu, M. Upadhyay, M. Sieniek, R. Sharma, T. Paine, A. Angelova, S. Addepalli, C. Parada, K. Majumder, A. Lamp, S. Kumar, X. Deng, A. Myaskovsky, T. SaboliÄ, J. Dudek, S. York, F. de Chaumont Quitry, J. Nie, D. Cattle, A. Gunjan, B. Piot, W. Khawaja, S. Bang, S. Wang, S. Khodadadeh, R. R, P. Rawlani, R. Powell, K. Lee, J. Griesser, G. Oh, C. Magalhaes, Y. Li, S. Tokumine, H. N. Vogel, D. Hsu, A. BC, D. Jindal, M. Cohen, Z. Yang, J. Yuan, D. de Cesare, T. Bruguier, J. Xu, M. Roy, A. Jacovi, D. Belov, R. Arya, P. Meadowlark, S. Cohen-Ganor, W. Ye, P. Morris-Suzuki, P. Banzal, G. Song, P. Ponnuramu, F. Zhang, G. Scrivener, S. Zaiem, A. R. Rochman, K. Han, B. Ghazi, K. Lee, S. Drath, D. Suo, A. Girgis, P. Shenoy, D. Nguyen, D. Eck, S. Gupta, L. Yan, J. Carreira, A. Gulati, R. Sang, D. Mirylenka, E. Cooney, E. Chou, M. Ling, C. Fan, B. Coleman, G. Tubone, R. Kumar, J. Baldridge, F. Hernandez-Campos, A. Lazaridou, J. Besley, I. Yona, N. Bulut, Q. Wellens, A. Pierigiovanni, J. George, R. Green, P. Han, C. Tao, G. Clark, C. You, A. Abdolmaleki, J. Fu, T. Chen, A. Chaugule, A. Chandorkar, A. Rahman, W. Thompson, P. Koanantakool, M. Bernico, J. Ren, A. Vlasov, S. Vassilvitskii, M. Kula, Y. Liang, D. Kim, Y. Huang, C. Ye, D. Lepikhin, and W. Helmholz Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. External Links: 2507.06261, Link Cited by: §4. Feng et al. (2025) T. Feng, Y. Sun, and J. You GraphEval: a lightweight graph-based LLM framework for idea evaluation. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2. Fortunato et al. (2018) S. Fortunato, C. T. Bergstrom, K. Börner, J. A. Evans, D. Helbing, S. MilojeviÄ, A. M. Petersen, F. Radicchi, R. Sinatra, B. Uzzi, A. Vespignani, L. Waltman, D. Wang, and A. BarabĂĄsi Science of science. Science 359 (6379), p. eaao0185. External Links: Document, Link, https://w.science.org/doi/pdf/10.1126/science.aao0185 Cited by: §1. GĂłmez-PĂ©rez et al. (2022) J. M. GĂłmez-PĂ©rez, A. GarcĂa-Silva, R. Leone, M. Albani, M. Fontaine, C. Poncet, L. Summerer, A. Donati, I. Roma, and S. Scaglioni Artificial intelligence and natural language processing and understanding in space: a methodological framework and four esa case studies. External Links: 2210.03640, Link Cited by: §2. Google et al. (2025) Google, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. RamĂ©, M. RiviĂšre, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. György, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. Klimczak-PluciĆska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. PĂ”der, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot Gemma 3 technical report. External Links: 2503.19786, Link Cited by: §5.1. Google (2025) Google A new era of intelligence with Gemini 3 â blog.google. Note: https://blog.google/products/gemini/gemini-3/#note-from-ceo[Accessed 05-01-2026] Cited by: §4. Gottesman and Geva (2024) D. Gottesman and M. Geva Estimating knowledge in large language models without generating a single token. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 3994â4019. External Links: Link, Document Cited by: Appendix A. Gottweis et al. (2026) J. Gottweis, W. Weng, A. Daryin, T. Tu, P. Sirkovic, A. Myaskovsky, G. Glowaty, F. Weissenberger, A. Orlandi, D. Popovici, A. Palepu, K. Rong, R. Tanno, K. Saab, F. Zhang, J. Blum, A. Carroll, K. Kulkarni, N. TomaĆĄev, D. Zverinski, I. Rendulic, E. Vedadi, F. Hasler, L. Rimanic, M. Boia, I. Budiselic, B. Feinstein, M. Bellaiche, T. Sheffer, J. Freyberg, J. Ratcliff, O. Bertolli, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. PenadĂ©s, G. Peltz, Y. Matias, J. Manyika, D. Hassabis, Y. Xu, P. Kohli, A. Pawlosky, A. Karthikesalingam, and V. Natarajan Accelerating scientific discovery with co-scientist. Nature. External Links: ISSN 1476-4687, Document, Link Cited by: §1. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. GuzmĂĄn, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Ăelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §5.1. Gurnee and Tegmark (2024) W. Gurnee and M. Tegmark Language models represent space and time. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: Appendix A. He et al. (2024) L. He, P. Chen, E. Nie, Y. Li, and J. R. Brennan Decoding probing: revealing internal linguistic structures in neural language models using minimal pairs. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Eds.), Torino, Italia, p. 4488â4497. External Links: Link Cited by: Appendix A. Hou et al. (2026) J. Hou, H. Deng, W. Jiao, X. Liu, X. Ke, and M. Zhang NoveltyAgent: autonomous novelty reporting agent with point-wise novelty analysis and self-validation. External Links: 2603.20884, Link Cited by: §2. Hu et al. (2022) E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: Link Cited by: Appendix C, §5.1. Jin et al. (2025) M. Jin, Q. Yu, J. Huang, Q. Zeng, Z. Wang, W. Hua, H. Zhao, K. Mei, Y. Meng, K. Ding, F. Yang, M. Du, and Y. Zhang Exploring concept depth: how large language models acquire knowledge and concept at different layers?. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, p. 558â573. External Links: Link Cited by: Appendix A. Ju et al. (2024) T. Ju, W. Sun, W. Du, X. Yuan, Z. Ren, and G. Liu How large language models encode context knowledge? a layer-wise probing study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Eds.), Torino, Italia, p. 8235â8246. External Links: Link Cited by: Appendix A. Klerings et al. (2025) A. Klerings, J. Brinkmann, D. Ruffinelli, and S. P. Ponzetto Steering language models in multi-token generation: a case study on tense and aspect. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 8621â8639. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: Appendix A. Li et al. (2023) K. Li, A. K. Hopkins, D. Bau, F. ViĂ©gas, H. Pfister, and M. Wattenberg Emergent world representations: exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: Appendix A. Li et al. (2025) L. Li, W. Xu, J. Guo, R. Zhao, X. Li, Y. Yuan, B. Zhang, Y. Jiang, Y. Xin, R. Dang, Y. Rong, D. Zhao, T. Feng, and L. Bing Chain of ideas: revolutionizing research via novel idea development with LLM agents. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 8971â9004. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §2. Li et al. (2024a) L. Li, W. Xu, J. Guo, R. Zhao, X. Li, Y. Yuan, B. Zhang, Y. Jiang, Y. Xin, R. Dang, D. Zhao, Y. Rong, T. Feng, and L. Bing Chain of ideas: revolutionizing research via novel idea development with llm agents. External Links: 2410.13185, Link Cited by: §1. Li et al. (2024b) Z. Li, Y. Cao, and J. C.K. Cheung Do llms build world representations? probing through the lens of state abstraction. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37, p. 98009â98032. External Links: Document, Link Cited by: Appendix A. Lin et al. (2025) E. Lin, Z. Peng, and Y. Fang Evaluating and enhancing large language models for novelty assessment in scholarly publications. In Proceedings of the 1st Workshop on AI and Scientific Discovery: Directions and Opportunities, P. Jansen, B. Dalvi Mishra, H. Trivedi, B. Prasad Majumder, T. Hope, T. Khot, D. Downey, and E. Horvitz (Eds.), Albuquerque, New Mexico, USA, p. 46â57. External Links: Link, Document, ISBN 979-8-89176-224-4 Cited by: §2. Liu et al. (2025) Y. Liu, Z. Yang, S. Poria, T. Nguyen, and E. Cambria Harnessing large language models for scientific novelty detection. External Links: 2505.24615, Link Cited by: §2. Lu et al. (2024) C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha The ai scientist: towards fully automated open-ended scientific discovery. External Links: 2408.06292, Link Cited by: §1, §5.1. Lu et al. (2026) C. Lu, C. Lu, R. T. Lange, Y. Yamada, S. Hu, J. Foerster, D. Ha, and J. Clune Towards end-to-end automation of ai research. Nature 651 (8107), p. 914â919. External Links: ISSN 1476-4687, Document, Link Cited by: §1. Maas et al. (2011) A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, D. Lin, Y. Matsumoto, and R. Mihalcea (Eds.), Portland, Oregon, USA, p. 142â150. External Links: Link Cited by: Appendix A. Maiya et al. (2025) S. Maiya, Y. Liu, R. Debnath, and A. Korhonen Improving preference extraction in LLMs by identifying latent knowledge through classifying probes. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 9061â9081. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: footnote 1. Marks and Tegmark (2024) S. Marks and M. Tegmark The geometry of truth: emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, External Links: Link Cited by: Appendix A. Mostafa et al. (2026) A. Mostafa, T. H. Nguyen, and Z. Ahmadi What is novel? a knowledge-driven framework for bias-aware literature originality evaluation. External Links: 2602.06054, Link Cited by: §1. Mysore et al. (2022) S. Mysore, A. Cohan, and T. Hope Multi-vector models with textual guidance for fine-grained scientific document similarity. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M. de Marneffe, and I. V. Meza Ruiz (Eds.), Seattle, United States, p. 4453â4470. External Links: Link, Document Cited by: §2. OpenAI et al. (2025) OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V. Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, L. Gross, K. G. Guzman, J. Hallman, J. Hehir, J. Heidecke, A. Helyar, H. Hu, R. Huet, J. Huh, S. Jain, Z. Johnson, C. Koch, I. Kofman, D. Kundel, J. Kwon, V. Kyrylov, E. Y. Le, G. Leclerc, J. P. Lennon, S. Lessans, M. Lezcano-Casado, Y. Li, Z. Li, J. Lin, J. Liss, Lily, Liu, J. Liu, K. Lu, C. Lu, Z. Martinovic, L. McCallum, J. McGrath, S. McKinney, A. McLaughlin, S. Mei, S. Mostovoy, T. Mu, G. Myles, A. Neitz, A. Nichol, J. Pachocki, A. Paino, D. Palmie, A. Pantuliano, G. Parascandolo, J. Park, L. Pathak, C. Paz, L. Peran, D. Pimenov, M. Pokrass, E. Proehl, H. Qiu, G. Raila, F. Raso, H. Ren, K. Richardson, D. Robinson, B. Rotsted, H. Salman, S. Sanjeev, M. Schwarzer, D. Sculley, H. Sikchi, K. Simon, K. Singhal, Y. Song, D. Stuckey, Z. Sun, P. Tillet, S. Toizer, F. Tsimpourlas, N. Vyas, E. Wallace, X. Wang, M. Wang, O. Watkins, K. Weil, A. Wendling, K. Whinnery, C. Whitney, H. Wong, L. Yang, Y. Yang, M. Yasunaga, K. Ying, W. Zaremba, W. Zhan, C. Zhang, B. Zhang, E. Zhang, and S. Zhao Gpt-oss-120b & gpt-oss-20b model card. External Links: 2508.10925, Link Cited by: §B.2, §5.1. OpenAI (2026) OpenAI GPT-5.4 thinking system card. OpenAI. Note: Accessed: 2026-05-20 External Links: Link Cited by: §4. Pedregosa et al. (2011) F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay Scikit-learn: machine learning in Python. Journal of Machine Learning Research 12, p. 2825â2830. Cited by: Appendix D. Picard et al. (2025) C. Picard, K. M. Edwards, A. C. Doris, B. Man, G. Giannone, M. F. Alam, and F. Ahmed From concept to manufacturing: evaluating vision-language models for engineering design. Artificial Intelligence Review 58 (9), p. 288. External Links: Document, ISBN 1573-7462, Link Cited by: §1. Sarica et al. (2020) S. Sarica, J. Luo, and K. L. Wood TechNet: technology semantic network based on patent data. Expert Systems with Applications 142, p. 112995. External Links: ISSN 0957-4174, Document, Link Cited by: §2. Schopf and FĂ€rber (2026) T. Schopf and M. FĂ€rber Is this idea novel? an automated benchmark for judgment of research ideas. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), S. Piperidis, N. Bel, H. van den Heuvel, N. Ide, S. Krek, and A. Toral (Eds.), Palma, Mallorca, Spain, p. 4716â4727. External Links: Document Cited by: Appendix B, §1, §3, §4. Shahid et al. (2025) S. Shahid, M. Radensky, R. Fok, P. Siangliulue, D. S. Weld, and T. Hope Literature-grounded novelty assessment of scientific ideas. In Proceedings of the Fifth Workshop on Scholarly Document Processing (SDP 2025), T. Ghosal, P. Mayr, A. Singh, A. Naik, G. Rehm, D. Freitag, D. Li, S. Schimmler, and A. De Waard (Eds.), Vienna, Austria, p. 96â113. External Links: Link, Document, ISBN 979-8-89176-265-7 Cited by: §5.1. Si et al. (2025) C. Si, D. Yang, and T. Hashimoto Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers. In International Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025, p. 94003â94092. External Links: Link Cited by: §1, §2, §4, §5.1. Singh et al. (2026) A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram, A. Nathan, A. Luo, A. Helyar, A. Madry, A. Efremov, A. Spyra, A. Baker-Whitcomb, A. Beutel, A. Karpenko, A. Makelov, A. Neitz, A. Wei, A. Barr, A. Kirchmeyer, A. Ivanov, A. Christakis, A. Gillespie, A. Tam, A. Bennett, A. Wan, A. Huang, A. M. Sandjideh, A. Yang, A. Kumar, A. Saraiva, A. Vallone, A. Gheorghe, A. G. Garcia, A. Braunstein, A. Liu, A. Schmidt, A. Mereskin, A. Mishchenko, A. Applebaum, A. Rogerson, A. Rajan, A. Wei, A. Kotha, A. Srivastava, A. Agrawal, A. Vijayvergiya, A. Tyra, A. Nair, A. Nayak, B. Eggers, B. Ji, B. Hoover, B. Chen, B. Chen, B. Barak, B. Minaiev, B. Hao, B. Baker, B. Lightcap, B. McKinzie, B. Wang, B. Quinn, B. Fioca, B. Hsu, B. Yang, B. Yu, B. Zhang, B. Brenner, C. R. Zetino, C. Raymond, C. Lugaresi, C. Paz, C. Hudson, C. Whitney, C. Li, C. Chen, C. Cole, C. Voss, C. Ding, C. Shen, C. Huang, C. Colby, C. Hallacy, C. Koch, C. Lu, C. Kaplan, C. Kim, C. Minott-Henriques, C. Frey, C. Yu, C. Czarnecki, C. Reid, C. Wei, C. Decareaux, C. Scheau, C. Zhang, C. Forbes, D. Tang, D. Goldberg, D. Roberts, D. Palmie, D. Kappler, D. Levine, D. Wright, D. Leo, D. Lin, D. Robinson, D. Grabb, D. Chen, D. Lim, D. Salama, D. Bhattacharjee, D. Tsipras, D. Li, D. Yu, D. Strouse, D. Williams, D. Hunn, E. Bayes, E. Arbus, E. Akyurek, E. Y. Le, E. Widmann, E. Yani, E. Proehl, E. Sert, E. Cheung, E. Schwartz, E. Han, E. Jiang, E. Mitchell, E. Sigler, E. Wallace, E. Ritter, E. Kavanaugh, E. Mays, E. Nikishin, F. Li, F. P. Such, F. de Avila Belbute Peres, F. Raso, F. Bekerman, F. Tsimpourlas, F. Chantzis, F. Song, F. Zhang, G. Raila, G. McGrath, G. Briggs, G. Yang, G. Parascandolo, G. Chabot, G. Kim, G. Zhao, G. Valiant, G. Leclerc, H. Salman, H. Wang, H. Sheng, H. Jiang, H. Wang, H. Jin, H. Sikchi, H. Schmidt, H. Aspegren, H. Chen, H. Qiu, H. Lightman, I. Covert, I. Kivlichan, I. Silber, I. Sohl, I. Hammoud, I. Clavera, I. Lan, I. Akkaya, I. Kostrikov, I. Kofman, I. Etinger, I. Singal, J. Hehir, J. Huh, J. Pan, J. Wilczynski, J. Pachocki, J. Lee, J. Quinn, J. Kiros, J. Kalra, J. Samaroo, J. Wang, J. Wolfe, J. Chen, J. Wang, J. Harb, J. Han, J. Wang, J. Zhao, J. Chen, J. Yang, J. Tworek, J. Chand, J. Landon, J. Liang, J. Lin, J. Liu, J. Wang, J. Tang, J. Yin, J. Jang, J. Morris, J. Flynn, J. Ferstad, J. Heidecke, J. Fishbein, J. Hallman, J. Grant, J. Chien, J. Gordon, J. Park, J. Liss, J. Kraaijeveld, J. Guay, J. Mo, J. Lawson, J. McGrath, J. Vendrow, J. Jiao, J. Lee, J. Steele, J. Wang, J. Mao, K. Chen, K. Hayashi, K. Xiao, K. Salahi, K. Wu, K. Sekhri, K. Sharma, K. Singhal, K. Li, K. Nguyen, K. Gu-Lemberg, K. King, K. Liu, K. Stone, K. Yu, K. Ying, K. Georgiev, K. Lim, K. Tirumala, K. Miller, L. Ahmad, L. Lv, L. Clare, L. Fauconnet, L. Itow, L. Yang, L. Romaniuk, L. Anise, L. Byron, L. Pathak, L. Maksin, L. Lo, L. Ho, L. Jing, L. Wu, L. Xiong, L. Mamitsuka, L. Yang, L. McCallum, L. Held, L. Bourgeois, L. Engstrom, L. Kuhn, L. Feuvrier, L. Zhang, L. Switzer, L. Kondraciuk, L. Kaiser, M. Joglekar, M. Singh, M. Shah, M. Stratta, M. Williams, M. Chen, M. Sun, M. Cayton, M. Li, M. Zhang, M. Aljubeh, M. Nichols, M. Haines, M. Schwarzer, M. Gupta, M. Shah, M. Y. Guan, M. Huang, M. Dong, M. Wang, M. Glaese, M. Carroll, M. Lampe, M. Malek, M. Sharman, M. Zhang, M. Wang, M. Pokrass, M. Florian, M. Pavlov, M. Wang, M. Chen, M. Wang, M. Feng, M. Bavarian, M. Lin, M. Abdool, M. Rohaninejad, N. Soto, N. Staudacher, N. LaFontaine, N. Marwell, N. Liu, N. Preston, N. Turley, N. Ansman, N. Blades, N. Pancha, N. Mikhaylin, N. Felix, N. Handa, N. Rai, N. Keskar, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, O. Gleeson, P. Mishkin, P. Lesiewicz, P. Baltescu, P. Belov, P. Zhokhov, P. Pronin, P. Guo, P. Thacker, Q. Liu, Q. Yuan, Q. Liu, R. Dias, R. Puckett, R. Arora, R. T. Mullapudi, R. Gaon, R. Miyara, R. Song, R. Aggarwal, R. Marsan, R. Yemiru, R. Xiong, R. Kshirsagar, R. Nuttall, R. Tsiupa, R. Eldan, R. Wang, R. James, R. Ziv, R. Shu, R. Nigmatullin, S. Jain, S. Talaie, S. Altman, S. Arnesen, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Yoo, S. Heon, S. Ethersmith, S. Grove, S. Taylor, S. Bubeck, S. Banesiu, S. Amdo, S. Zhao, S. Wu, S. Santurkar, S. Zhao, S. R. Chaudhuri, S. Krishnaswamy, Shuaiqi, Xia, S. Cheng, S. Anadkat, S. P. Fishman, S. Tobin, S. Fu, S. Jain, S. Mei, S. Egoian, S. Kim, S. Golden, S. Mah, S. Lin, S. Imm, S. Sharpe, S. Yadlowsky, S. Choudhry, S. Eum, S. Sanjeev, T. Khan, T. Stramer, T. Wang, T. Xin, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Degry, T. Shadwell, T. Fu, T. Gao, T. Garipov, T. Sriskandarajah, T. Sherbakov, T. Korbak, T. Kaftan, T. Hiratsuka, T. Wang, T. Song, T. Zhao, T. Peterson, V. Kharitonov, V. Chernova, V. Kosaraju, V. Kuo, V. Pong, V. Verma, V. Petrov, W. Jiang, W. Zhang, W. Zhou, W. Xie, W. Zhan, W. McCabe, W. DePue, W. Ellsworth, W. Bain, W. Thompson, X. Chen, X. Qi, X. Xiang, X. Shi, Y. Dubois, Y. Yu, Y. Khakbaz, Y. Wu, Y. Qian, Y. T. Lee, Y. Chen, Y. Zhang, Y. Xiong, Y. Tian, Y. Cha, Y. Bai, Y. Yang, Y. Yuan, Y. Li, Y. Zhang, Y. Yang, Y. Jin, Y. Jiang, Y. Wang, Y. Wang, Y. Liu, Z. Stubenvoll, Z. Dou, Z. Wu, and Z. Wang OpenAI gpt-5 system card. External Links: 2601.03267, Link Cited by: §4. Su et al. (2025) H. Su, R. Chen, S. Tang, Z. Yin, X. Zheng, J. Li, B. Qi, Q. Wu, H. Li, W. Ouyang, P. Torr, B. Zhou, and N. Dong Many heads are better than one: improved scientific idea generation by a LLM-based multi-agent system. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 28201â28240. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1. Tang et al. (2025) J. Tang, L. Xia, Z. Li, and C. Huang AI-researcher: autonomous scientific innovation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2. Uzzi et al. (2013) B. Uzzi, S. Mukherjee, M. Stringer, and B. Jones Atypical combinations and scientific impact. Science 342 (6157), p. 468â472. External Links: Document, Link, https://w.science.org/doi/pdf/10.1126/science.1240474 Cited by: §2. Valois et al. (2025) P. H. V. Valois, L. S. Souza, E. K. Shimomoto, and K. Fukui Frame representation hypothesis: multi-token llm interpretability and concept-guided text generation. Transactions of the Association for Computational Linguistics 13, p. 1436â1458. External Links: ISSN 2307-387X, Document, Link, https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.48/2561673/tacl.a.48.pdf Cited by: Appendix A. Wang et al. (2017) J. Wang, R. Veugelers, and P. Stephan Bias against novelty in science: a cautionary tale for users of bibliometric indicators. Research Policy 46 (8), p. 1416â1436. External Links: ISSN 0048-7333, Document, Link Cited by: §2. Wang et al. (2019) K. Wang, B. Dong, and J. Ma Towards computational assessment of idea novelty. In Proceedings of the 52nd Hawaii International Conference on System Sciences, External Links: ISBN 978-0-9981331-2-6, Link Cited by: §2. Wang et al. (2025) W. Wang, L. Gu, L. Zhang, Y. Luo, Y. Dai, C. Shen, L. Xie, B. Lin, X. He, and J. Ye SciPIP: an llm-based scientific paper idea proposer. External Links: 2410.23166, Link Cited by: §2. Wolf et al. (2020) T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush Transformers: state-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Q. Liu and D. Schlangen (Eds.), Online, p. 38â45. External Links: Link, Document Cited by: Appendix D. Wu et al. (2025) W. Wu, C. Zhang, and Y. Zhao Automated novelty evaluation of academic paper: a collaborative approach integrating human and large language model knowledge. Journal of the Association for Information Science and Technology 76 (11), p. 1452â1469. External Links: Document, Link, https://asistdl.onlinelibrary.wiley.com/doi/pdf/10.1002/asi.70005 Cited by: §2. Wu et al. (2026) W. Wu, Y. Zhao, Y. Wang, S. Li, J. Shao, Y. Long, and C. Zhang NovBench: evaluating large language models on academic paper novelty assessment. External Links: 2604.11543, Link Cited by: §1. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: §5.1. Yang et al. (2024) Z. Yang, X. Du, J. Li, J. Zheng, S. Poria, and E. Cambria Large language models for automated open-domain scientific hypotheses discovery. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 13545â13565. External Links: Link, Document Cited by: §5.1. Zhang et al. (2025) Y. Zhang, H. Diddee, S. Holm, H. Liu, X. Liu, V. Samuel, B. Wang, and D. Ippolito NoveltyBench: evaluating creativity and diversity in language models. In Second Conference on Language Modeling, External Links: Link Cited by: §2. Zur et al. (2025) A. Zur, A. Geiger, E. S. Lubana, and E. Bigelow Are language models aware of the road not taken? token-level uncertainty and hidden state dynamics. External Links: 2511.04527, Link Cited by: Appendix A. Appendix A About Probing llm During Generation Probing approaches quantify the extent to which llm representations encode specific knowledge. While extensive research investigates internal knowledge across diverse domains such as sentiment Maas et al. (2011) and factual knowledge Marks and Tegmark (2024), spatial and temporal understanding Gurnee and Tegmark (2024), and world models Li et al. (2023), such existing studies predominantly focus on layer-wise localization of internal knowledge (He et al., 2024; Li et al., 2024b; Ju et al., 2024; Jin et al., 2025, inter alia). Work on probing llm during different generation steps is scarce and primarily addresses steering text generation (Valois et al., 2025; Zur et al., 2025; Klerings et al., 2025, inter alia). Although some works investigate the encoded information in llm before generating the first token Gottesman and Geva (2024); Afzal et al. (2025), they do not distinguish between functional phases during generation. We address this gap by providing the first comparison of llm representations during the reasoning (âthinkingâ) phase versus the response generation phase via probing, demonstrating that llm encode more information about research idea novelty judgments while thinking than when producing the actual response. Score Degree of Novelty 1 The idea is not novel. All aspects already exist in prior work. 2 The idea is marginally novel. It represents only a minor variation of existing work. 3 The idea is somewhat novel. Aspects already exist in prior work. However, it might combine known approaches in new ways, apply them to new contexts, or propose incremental updates. 4 The idea is novel. It introduces new aspects not present in existing work. 5 The idea is highly innovative and novel. It is not present in existing work and potentially encourages new thinking or opens up new research directions. Table 5: Novelty Judgment Rubric Appendix B Evaluation Metrics We briefly summarize the rino evaluation metrics, which we adopt in this work. For full details, see Schopf and FĂ€rber (2026). The metrics evaluate both numerical novelty scores and textual justifications. B.1 Novelty Score Metrics We evaluate predicted novelty scores using macro-F1F_1, class-wise F1F_1, and mae. Macro-F1F_1 measures overall classification performance across the five novelty categories, class-wise F1F_1 shows performance for each individual score, and mae measures the average distance between predicted and human gold scores on the ordinal 1â5 scale. B.2 Justification Metrics For textual justifications, rino distinguishes between known aspects, which describe overlaps with prior work, and novelty aspects, which describe new contributions of the research idea. Following rino, these metrics are computed using an llm-as-a-judge approach that compares model-generated justifications against human gold-standard justifications. In this work, we use the GPT-OSS-120B OpenAI et al. (2025) model for evaluation. Alignment Alignment measures whether the model-generated justification follows reasoning consistent with the human gold justification and supports a similar novelty judgment. Scores range from 0 to 1, where higher is better: 1 indicates strong agreement with the human rationale, while 0 indicates no alignment. Recall Recall measures how many known-aspect and novelty-aspect arguments from the human gold justification are captured by the model-generated justification. Scores range from 0 to 100, where higher is better: 100 means all relevant gold arguments are covered, while 0 means none are covered. Additional Ratio Additional ratio measures how many extra known-aspect and novelty-aspect arguments the model adds beyond the gold justification, while still being grounded in the related works or research idea. Scores are non-negative percentages, where 0% means no additional grounded arguments are added and higher values indicate more extra grounded content. This metric is not inherently good or bad: moderate or high values can indicate useful additional evidence, but very high values may also reflect overly verbose justifications. Hallucination Rate Hallucination rate measures the proportion of generated known-aspect and novelty-aspect arguments that are not supported by the related works or research idea. Scores range from 0% to 100%, where lower is better: 0% indicates justifications that are fully grounded in the research idea and related works, while higher values indicate more unsupported or hallucinated content. Appendix C LoRA Fine-tuning Details For the FineTune approach in Section 5, we fine-tune the base llm using Low-Rank Adaptation (LoRA; Hu et al., 2022). We apply LoRA to all major projection layers in the transformer, including the query, key, value, and output projections of the attention mechanism, as well as the gate, up, and down projections in the feed-forward network. We use a rank of r=16r=16, a scaling factor α=32α=32, and a LoRA dropout of 0.10.1. Training is performed for two epochs using a per-device batch size of 1 and gradient accumulation over 8 steps, resulting in an effective batch size of 8. We employ a learning rate of 2Ă10â42Ă 10^-4 with a short warmup of 20 steps. To reduce memory consumption, gradient checkpointing is enabled, and training is conducted in bfloat16 precision. Appendix D Experimental Details All experiments were conducted on two NVIDIA A100 (80GB) GPUs. The probing classifier was implemented using scikit-learn Pedregosa et al. (2011). Hidden states were extracted using the Hugging Face Transformers library Wolf et al. (2020). Appendix E On the Choice of Novelty Score Metrics We use macro-F1F_1 as the primary metric for evaluating novelty score predictions, rather than mae. Although mae is useful for measuring the average ordinal distance between predicted and gold scores, it is less informative in our setting because the novelty scale is small (1â5) and model predictions are strongly concentrated around the middle categories. As shown in Table 6, different models and approaches obtain very similar mae values, typically around one. This makes mae unsuitable for evaluation in our setting. Model Zero-shot Few-shot CoT Moose Research Agent AI Scientist AI Researcher FineTune tpr Non-Reas. Gemma-3-4B 1.0 0.9 1.0 0.9 0.9 0.9 0.9 1.1 1.0 Gemma-3-12B 1.0 0.9 0.9 1.0 0.9 0.9 0.9 1.1 1.1 Gemma-3-27B 1.0 1.0 0.9 1.0 0.9 1.0 0.9 1.1 1.1 Llama-3.1-8B 1.0 1.0 1.0 0.9 0.9 0.9 0.9 1.1 1.1 Llama-3.1-70B 1.0 1.0 1.1 1.0 1.0 1.0 0.9 1.1 1.1 Reas. Qwen3-4B 1.0 1.0 1.0 1.0 0.9 1.0 1.0 1.0 1.1 Qwen3-14B 1.0 1.0 1.1 1.0 1.0 1.0 1.0 1.1 1.0 Qwen3-32B 1.0 1.0 1.0 1.0 0.9 1.0 1.0 1.0 1.1 GPT-OSS-20B 0.9 0.9 1.0 0.9 0.9 1.0 0.9 1.2 1.1 Table 6: mae scores for different approaches and llm on the rino test set. The limitation arises because a model that repeatedly predicts a middle score, such as 3, can achieve a deceptively low mae: many gold labels are only one or two points away on a five-point Likert scale. Conversely, a less biased model that also predicts more extreme novelty categories may occasionally incur larger absolute errors, even if it produces more accurate novelty judgments overall. Optimizing for mae can therefore favor conservative middle-ground predictions, precisely the behavior we aim to mitigate. Macro-F1F_1 better reflects our evaluation objective. It treats all novelty classes equally, regardless of their frequency, and explicitly rewards models for correctly predicting low, medium, and high novelty judgments. This is crucial for evaluating whether a model can correctly judge extreme novelty categories, such as ânot novelâ and âhighly novelâ, rather than merely staying close to the center of the scale. We therefore report mae for completeness in Table 6, but use macro-F1F_1 and class-wise F1F_1 as the main indicators of novelty judgment performance. Appendix F Justification Evaluation Beyond novelty judgment performance, we also evaluate the quality of textual justifications generated by open-source llm using tpr, with the results shown in Table 7. ALI Recall Add. Ratio Hall. Rate Model KA NA KA NA KA NA Non-Reas. Gemma-3-4B 0.25 39.4 35.7 40.8 51.9 38.8 44.2 Gemma-3-12B 0.39 58.2 50.4 56.7 72.4 8.6 6.8 Gemma-3-27B 0.42 62.4 57.2 60.9 80.5 6.5 2.0 Llama-3.1-8B 0.36 52.0 45.3 40.0 74.1 12.6 10.5 Llama-3.1-70B 0.30 51.0 50.0 35.1 55.6 12.0 11.1 Reas. Qwen3-4B 0.43 61.9 61.5 78.9 108.3 7.8 2.6 Qwen3-14B 0.43 63.6 68.8 105.4 137.2 9.2 5.4 Qwen3-32B 0.54 67.9 59.0 94.5 98.7 8.0 1.1 GPT-OSS-20B 0.52 66.6 63.6 88.4 94.8 10.6 2.8 Table 7: Evaluation of textual justifications generated by different llm using tpr. Alignment The alignment scores of open-source models using tpr are comparable to those of substantially larger proprietary llm in Table 1. In particular, the strongest reasoning-capable models achieve ALI values around 0.5, close to the range observed for Claude and GPT models under zero-shot prompting. This indicates that tpr does not merely improve numerical novelty prediction, but also enables open-source models to generate justifications whose reasoning remains broadly aligned with human gold-standard rationales. Recall Similarly to propriety llm, smaller open-source models exhibit relatively high recall, indicating substantial overlap between model-generated and human-annotated justification arguments. Comparing the results to the ones in Table 1, recall in open-source models is slightly lower than in the top-performing OpenAI models but remains competitive with Gemini Pro models. Additional Ratio The Additional Ratio is generally higher for reasoning-capable models than for non-reasoning models, indicating that reasoning models produce more elaborate justifications. While proprietary OpenAI llm exhibit even higher additional ratios, the best-performing Qwen3 models remain competitive with Gemini Pro models, highlighting that tpr enables smaller open-source models to generate rich novelty judgment justifications. Hallucination Rate Reasoning models show low hallucination rates similar to proprietary models, whereas non-reasoning models, particularly Gemma-3-4B, display higher hallucination rates when attempting to justify novelty judgments. This suggests that tpr is most effective in producing grounded justifications when paired with reasoning llm. Takeaway Overall, tpr allows open-source LLMs to generate high-quality, human-aligned novelty justifications, similar to what is achievable with large proprietary models. However, performance depends strongly on the model: reasoning-capable llm consistently produce more accurate, elaborate, and reliable justifications, whereas smaller non-reasoning llm may struggle to generate appropriate novelty judgment justifications. Gold Novelty Predicted Novelty Gold Justification LLM-generated Justification 11 33 â The idea of using neural networks to learn Greenâs functions is already known and the proposed contribution is incremental, offering no new aspects beyond existing approaches. The proposed idea combines fundamental solutions [âŠ] represents a somewhat novel synthesis, but the core components [âŠ] are well-established in the literature. 22 44 â The approach adds [âŠ] extensions of existing communication mechanisms rather than fundamentally new concepts, resulting in only marginal novelty. The research idea proposes [âŠ] meaningful novelty [âŠ], though it builds on existing concepts in communication [âŠ] 33 33 â The approach primarily assembles existing components [âŠ], resulting in a somewhat novel contribution. The research idea combines several existing concepts [âŠ] adds incremental novelty, [âŠ] 44 44 â The approach is novel [âŠ] extending diffusion models beyond the usual focus on learning only the reverse process. [âŠ] has not been presented in prior work [âŠ] [âŠ] introduces a novel approach by jointly parameterizing both forward and reverse diffusion processes [âŠ] related works [âŠ] primarily focus on the reverse process [âŠ] 55 44 â The idea is highly novel because it uncovers a previously unreported generalization phenomenon and establishes a new theoretical link [âŠ] [âŠ] calibration literature doesnât explicitly connect ensemble disagreement to generalization [âŠ] builds incrementally on existing calibration and ensemble concepts. Table 8: Selected comparison of gold novelty judgments and LLM(Claude Opus 4.5)-generated novelty judgments. Green denotes alignment of model and gold justifications. Red highlights miscalibration, where the modelâs novelty judgment diverges from the gold judgment despite exhibiting a rationale aligned with the gold justification. Symbols indicate novelty overestimation (â ), correct prediction (â ), and underestimation (â ). ⏠system prompt = "You are an expert researcher experienced in judging the novelty of a research idea." user prompt = f""" You are an expert in machine learning research evaluation. You will be given two inputs: 1. A research idea with objective, problem statement, and solution approach. 2. A list of related works, each with a title and abstract. Your task is to **assess the novelty of the research idea** compared to the related works. ### Instructions: - Analyze the research idea and summarize its key contributions. - Compare it with the related works to identify overlaps and differences. - Specifically, assess whether the idea introduces **significant new aspects** not present in existing work, or if it is largely a variation on known approaches. - Provide your output as a **JSON object only**, with: - "reasoning": a short paragraph (2-4 sentences) explaining the reasoning behind the novelty score. - "novelty_score": an integer between 1-5 where: novelty_rubric ### Inputs: **Research Idea:** research_idea **Related Works:** related_works ### Output Format: ââjson "reasoning": <short explanation>, "novelty_score": <1|2|3|4|5> """ Figure 3: Prompts for the zero-shot approach to judging the novelty of research ideas. Here, an llm receives a research idea, its related works, and the rino novelty rubric, and is asked to generate both a numerical novelty score and a textual justification. ⏠system prompt = "You are ReviewerGPT, an intelligent assistant that helps researchers evaluate the novelty of their ideas." user prompt = f""" You are given some papers similar to the proposed idea (<IDEA> and </IDEA>). Your task is to evaluate the ideaâs novelty using the related papers (<PAPER> and </PAPER>) only. ### Novelty types: novelty_class_descriptions ### Instructions: - Use the example review below to write a review for the provided idea by comparing it to the related papers. - Donât assume any prior knowledge about the idea. - Make sure the generated review follows the format in example review provided below. - The review should be concise - around 60 to 100 words. ### Research Idea: research_idea ### Related Papers: related_papers ### Example Review: example_review ### Output Format: <REVIEW> concise review </REVIEW> Think step by step before generating the review! """ Figure 4: Instruction used for our tpr approach as introduced in Section 5. The instruction to reason step by step is included only for models that do not generate think tokens by default and is omitted otherwise.