Paper deep dive
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model
Ergan Shang, Weijing Tang, Yinqiu He
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.
Tags
Links
- Source: https://arxiv.org/abs/2608.22295v1
- Canonical: https://arxiv.org/abs/2608.22295v1
Trouble viewing inline? Open PDF directly →
Full Text
33,324 characters extracted from source content.
Expand or collapse full text
LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model Ergan Shang Affiliation: Department of Statistics and Data Science, Carnegie Mellon University Weijing Tang Affiliation: Department of Statistics and Data Science, Carnegie Mellon University Correspondence to: weijingt@andrew.cmu.edu Yinqiu He Affiliation: Department of Statistics, University of Madison Correspondence to: yinqiu.he@wisc.edu Abstract Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge. Keywords: LLM evaluation, out-of-sample performance prediction, multidimensional item response theory, contextual embeddings 1 Introduction Evaluation of large language models (LLMs) is increasingly shaped by rapid model iteration, evolving benchmarks, and deployment settings that change faster than full-scale annotation pipelines can keep up (14; 29; 27; 22). In this setting, a central question is no longer only how a model performed on a fixed benchmark, but how well observed evaluation results can predict behavior on new questions. Reliable prediction on unseen questions is crucial for deployment decisions, because it helps distinguish genuine generalization from benchmark-specific tuning and supports more trustworthy model selection and risk assessment (24; 12). This is especially important in high-stakes or rapidly changing applications, where collecting new gold-standard labels may require costly human annotators or expensive judge models (19). In such cases, leveraging existing or small-scale evaluations to predict performance on new questions can substantially improve evaluation efficiency and guide resource allocation (16; 25). Reliable prediction is challenging, however, because benchmark items are heterogeneous. Questions can vary substantially in difficulty, context, and the underlying capabilities they demand (33). As a result, simply averaging correctness across questions can confound model capabilities with question characteristics, obscuring structures that matter for extrapolation to unseen questions (10). Item response theory (IRT) serves as a natural approach to address this problem (3; 6). Prior studies have used IRT to build evaluation scales, estimate latent item parameters from model response patterns, analyze benchmark quality, and enable more reliable or amortized evaluation (10; 11; 9; 16). More recently, several works have combined IRT modeling with question embeddings to estimate item characteristics and support prediction on previously unseen questions (18; 17; 24; 23). However, most existing approaches remain largely based on unidimensional ability, which may not capture the multiple latent capabilities that drive LLM performance. They also leave open a harder question: whether predictive evaluation can generalize across heterogeneous scenarios rather than only to held-out questions within the same benchmark setting. This broader scope of generalization is especially important when evaluating LLMs on new tasks or benchmarks (20; 13; 32). In this work, we study LLM evaluation on unseen questions through a contextual multidimensional IRT (MIRT) model. Our model jointly captures question characteristics and model capabilities, allowing semantic information from question text to inform item parameters while providing a richer description of latent capability structure. Our main findings are threefold. First, incorporating question embeddings into MIRT models improves predictive performance relative to model-free alternatives. Second, allowing multiple latent dimensions yields a richer description of model capability than a purely unidimensional approach. Third, although the proposed framework performs encouragingly when predicting held-out questions within the same scenario, its cross-scenario performance has larger variation, suggesting that transfer learning across domains remains a major open challenge. In summary, these results show context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while cautioning that within-scenario predictive success may not necessarily translate to robust cross-scenario generalization. 1.1 Related work Benchmark-based evaluation remains dominant for comparing LLMs, but its limitations in both generalizability and efficiency are increasingly recognized. Prior studies have argued that conclusions drawn from static benchmarks may not directly transfer to changing evaluation protocols (2) or scenarios (8; 15). Meanwhile, a growing line of work seeks to reduce evaluation cost by recovering full-benchmark conclusions from only a small subset of instances (16; 19; 12; 25; 30). Our work shares the goal of moving beyond retrospective scoring on a fixed benchmark, but studies a structured model-based approach for predicting performance on unseen questions. Achieving this goal requires separating model-side abilities from heterogeneous question-side characteristics. IRT has long offered a principled way to disentangle item properties from participant abilities in psychometrics (3; 4). In natural language processing (NLP), 10; 11 introduce IRT-based evaluation scales. The recent tutorial by 9 highlights growing interest in IRT as a general framework for model assessment in language technology. In the LLM setting, an emerging number of studies have used IRT-based modeling to study benchmark measurement quality or interpretations (31; 28; 5). A parallel line of work leverages question texts or embedding-based features to model item characteristics and support generalization to unseen items. In educational assessment, such approaches have been used for predicting difficulty (1; 17) or unseen items (18; 7). For LLM evaluation, the work most closely related to ours is 24, who combine Rasch-style psychometric modeling and question embeddings to support reliable and efficient amortized evaluation. Our method builds on this line of work but differs in two key respects. First, we explicitly model multivariate latent capability dimensions of LLMs rather than relying on a unidimensional Rasch-model view or hard-to-interpret blackbox predictors. Second, we focus on the generalizability of evaluation prediction on unseen questions, including the more difficult setting of cross-scenario transfer rather than only efficient estimation on a fixed benchmark. 2 Method To predict evaluations of LLMs on unseen questions, we use the following contextual multidimensional item response theory (C-MIRT) model that separates model-side latent abilities from question-side characteristics across scenarios. C-MIRT model. Consider a benchmark dataset where models are evaluated across S different scenarios, i.e., different task types or contextual domains. For each scenario s∈[S]s∈[S], suppose n LLMs are evaluated on psp_s questions. Let yij(s)∈0,1y_ij^(s)∈\0,1\ denote whether model i answers question j correctly in scenario s. Each question j is associated with a contextual embedding ej∈ℝde_j ^d. To model performance on these questions, we map this embedding into an r-dimensional latent question representation through a feature map ϕ(s):ℝd→ℝrφ^(s):R^d ^r. Then, the response probability in the C-MIRT is modeled through a logistic link as yij(s)∣θij(s)∼Bernoulli(σ(θij(s))),σ(x)=11+e−x,y_ij^(s) _ij^(s) \! (σ( _ij^(s)) ), σ(x)= 11+e^-x, with the parameter θij(s)=αi(s)+ui(s)⊤ϕ(s)(ej). _ij^(s)= _i^(s)+u_i^(s) φ^(s)(e_j). (1) Here αi(s)∈ℝ _i^(s) is a scenario-specific intercept for model i, capturing its baseline success rate in scenario s, and ui(s)∈ℝru_i^(s) ^r is a scenario-specific latent capability vector. The transformed embedding ϕ(s)(ej)φ^(s)(e_j) represents the latent characteristics of question j in scenario s. The bilinear form ui(s)⊤ϕ(s)(ej)u_i^(s) φ^(s) (e_j ) captures how well the capabilities of model i align with the requirements of question j. Different specifications of ϕ(s)φ^(s) can be used to incorporate contextual information about the questions. For example, 23 models ϕ(s)φ^(s) as a smooth function in a reproducing kernel Hilbert space, so that questions with similar contextual embeddings are encouraged to have similar latent representations. The C-MIRT formulation goes beyond a unidimensional difficulty-based view by allowing model performance to vary along multiple latent traits. In particular, the C-MIRT model is closely related to the Rasch-style formulation in 24, which models θij(s)=αi(s)−βj(s), _ij^(s)= _i^(s)- _j^(s), with a scalar difficulty parameter βj(s) _j^(s) for question j in scenario s. In that formulation, all item heterogeneity is absorbed into a single difficulty dimension. In contrast, C-MIRT in (1) takes a multivariate bilinear form, which allows question characteristics and model capabilities to interact through multiple latent dimensions. An additional advantage of C-MIRT is that it naturally supports prediction on unseen questions. Once the feature map ϕ(s)φ^(s) is learned from training questions in scenario s, a new question can be embedded into the same latent space and used to predict response probabilities. This allows us to investigate the generalization within and across scenarios. Model estimation. Given a source scenario s, we estimate the model parameters (αi(s),ui(s),ϕ(s))(α_i^(s),u_i^(s),φ^(s)) by minimizing the logistic negative log-likelihood under the C-MIRT model, where the feature map ϕ(s)φ^(s) is parameterized by a multilayer perceptron. Because the parameters in MIRT are only identifiable up to a linear transformation, we adopt a two-step estimation procedure; details are deferred to the Appendix. Predicting responses to unseen questions. Suppose α^i(s),u^i(s),ϕ^(s) α_i^(s), u_i^(s), φ^(s) are estimated from the training split of the source scenario s. For a question j from target scenario t, with embedding eje_j, we compute θ^ij(s→t)=α^i(s)+u^i(s)⊤ϕ^(s)(ej),p^ij(s→t)=σ(θ^ij(s→t)). θ_ij^(s→ t)= α_i^(s)+ u_i^(s) φ^(s)(e_j), p_ij^(s→ t)=σ\! ( θ_ij^(s→ t) ). When t=st=s, this gives within-scenario predictions for previously unseen questions, enabled by the learned feature map ϕ^(s) φ^(s). When t≠st≠ s, this gives cross-scenario predictions. This directly evaluates out-of-scenario generalization, and the performance depends on how well the latent structure learned under scenario s transfers to scenario t, as we demonstrate empirically in Section 3.1. 3 Experiments We adopt the benchmark setup in 24, which leveraged 22 datasets from 5 HELM repositories: Classic, Lite, AIR-Bench, Thai Exam, and MMLU. We filtered out scenarios where the question descriptions are incomplete, e.g., questions with missing multiple-choice options. After filtering, we retain 11 scenarios for evaluation. Examples of scenarios include “wikifact” and “math”. The number of questions per scenario ranges from 436 to 29407. We construct a contextual embedding eje_j for each benchmarking question using BERT-based language models (26; 21). All methods are evaluated using 5-fold cross-validation. Within each scenario, questions are split into five folds. In each run, one fold, containing 20% of questions, is used for testing, and the remaining four folds, containing 80% of questions, are used for training. For each training-test scenario pair and each fold, the fitted model produces a predicted score matrix, where rows correspond to n LLMs and columns correspond to ps,testp_s,test testing questions in scenario s. We evaluate this matrix in two ways. The first is question-wise AUC, or column-wise AUC, where for each fixed test question, we compute the AUC using the predicted scores across n LLMs. This reflects how well the predicted scores rank LLMs on the same question. Repeating this over ps,testp_s,test questions yields a distribution of question-wise AUC values. The second is LLM-wise AUC, or row-wise AUC, where for a fixed LLM, we compute the AUC using the predicted scores across ps,testp_s,test questions. This measures how well the predicted scores rank questions for that LLM. Repeating this over n LLMs yields a distribution of LLM-wise AUC values. To summarize either distribution, we report its 9090th percentile. 3.1 Within- and cross-scenario prediction by C-MIRT This section evaluates the prediction performance of the C-MIRT. Figures 1 and 2 report the mean 90th percentile AUC over the five cross-validation folds. In both figures, rows indicate the training scenario and columns indicate the test scenario. Diagonal entries therefore represent within-scenario prediction, whereas off-diagonal entries represent cross-scenario prediction. Figure 1: Heatmap of the 90th percentile of question-wise AUC distribution predicted by the C-MIRT method across scenarios. Figure 1 shows the results for question-wise AUC. The diagonal entries are consistently high, indicating strong within-scenario performance in ranking LLMs for fixed questions. Several off-diagonal entries are also relatively large, suggesting that the learned contextual structure can transfer across scenarios to some extent. For example, training scenarios such as babi_qa and wikifact, achieve comparatively strong performance on multiple test scenarios. Meanwhile, the substantial variation among off-diagonal cells indicates that cross-scenario transfer depends on the similarity between the source and target scenarios. Figure 2 shows the results for LLM-wise AUC. Compared with Figure 1, the diagonal entries remain strong, although the off-diagonal entries are much closer to 0.5. This suggests that ranking question difficulty for a fixed LLM is more scenario-specific and less transferable across scenarios than ranking LLMs for a fixed question. Overall, C-MIRT model performs well for within-scenario prediction in both settings, while cross-scenario generalization is notably stronger for question-wise AUC than for LLM-wise AUC. Figure 2: Heatmap of the 90th percentile of LLM-wise AUC distribution predicted by the C-MIRT method across scenarios. 3.2 Comparison with competing methods We compare C-MIRT with three baselines. The first is the contextual Rasch-style model in 24. The second is a lasso-penalized logistic regression model with parameters θij(s)=γi0(s)+γi(s)⊤ej. _ij^(s)= _i0^(s)+ _i^(s) e_j. This baseline also uses contextual embeddings eje_j, but unlike C-MIRT, does not impose a low-rank latent structure in the coefficient matrix. The third is a simple mean baseline: within each scenario, the predicted score for every test question is set to the average training-set accuracy. Because this baseline assigns the same score to all questions for a given LLM, we use it only for question-wise AUC. For each method, we compute the 90th percentile AUC in the same way as in Section 3.1. For each pair of methods, 5-fold cross-validation produces 5 paired differences between C-MIRT and the comparator. We assess the significance of these differences using paired t-tests. In Figures 3 and 4, each heatmap cell corresponds to a scenario–method pair. The number in each cell is the average, across the five folds, of the difference in the 90th percentile of the AUC distribution between C-MIRT and the competing method. A cell is shaded gray when its paired t-test yields a p-value >0.05>0.05, indicating the difference is not statistically significant at the 5% level. Among the remaining cells, red indicates that C-MIRT performs better, whereas blue indicates that the competing method performs better. Figure 3: Heatmap of mean differences in the 90th percentile of question-wise AUC between C-MIRT and competing methods across scenarios. Figure 4: Heatmap of mean differences in the 90th percentile of LLM-wise AUC between C-MIRT and competing methods across scenarios. Figure 3 shows that for question-wise AUC, C-MIRT consistently outperforms the Mean and Rasch baselines across different scenarios. In contrast, its comparison with the lasso method is more scenario-dependent. Figure 4 shows that, however, for LLM-wise AUC, C-MIRT remains competitive across the majority of scenarios. These results highlight the robustness of the proposed method. 4 Conclusions and future work This paper proposes to use C-MIRT for predicting evaluations of LLMs on unseen questions. We show that C-MIRT enhances within-scenario prediction, whereas cross-scenario generalizability depends on the source-target scenarios and criterion. It opens an interesting future direction to further investigate new-scenario prediction for LLMs. Impact Statement This paper studies a more efficient and interpretable approach to evaluating LLMs on unseen questions. A potential positive impact of this work is that it can reduce the cost of AI evaluation, both in annotation effort and computation, while still providing structured information about model strengths and weaknesses. However, the method also presents risks. Predicted performance may be mistaken for direct evidence of reliability, especially in high-stakes applications, and embedding-based item models may inherit biases from benchmark data or pretrained text representations. Our results also show that cross-scenario prediction is substantially harder than within-scenario prediction, so these methods should not be treated as a substitute for direct testing on representative data. We therefore view this work as a tool for supplementing evaluation pipelines, not replacing careful benchmark design, independent validation, or human oversight. References AlKhuzaey et al. (2024) S. AlKhuzaey, F. Grasso, T. R. Payne, and V. Tamma Text-based question difficulty prediction: a systematic review of automatic approaches. International Journal of Artificial Intelligence in Education 34 (3), p. 862–914. Cited by: §1.1. Alzahrani et al. (2024) N. Alzahrani, H. A. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushaykeh, F. Mirza, N. Alotaibi, N. Altwairesh, A. Alowisheq, M. S. Bari, and H. Khan When benchmarks are targets: revealing the sensitivity of large language model leaderboards. arXiv preprint arXiv:2402.01781. Cited by: §1.1. Baker (2001) F. B. Baker The basics of item response theory. ERIC. Cited by: §1.1, §1. Cai et al. (2016) L. Cai, K. Choi, M. Hansen, and L. Harrell Item response theory. Annual Review of Statistics and Its Application 3 (1), p. 297–321. Cited by: §1.1. Cai et al. (2025) P. Cai, C. Cui, F. M. Polo, S. Somerstep, L. Choshen, M. Yurochkin, M. Banerjee, Y. Sun, K. M. Tan, and G. Xu A latent variable framework for scaling laws in large language models. arXiv preprint arXiv:2512.06553. Cited by: §1.1. Chen et al. (2025) Y. Chen, X. Li, J. Liu, and Z. Ying Item response theory—a statistical framework for educational and psychological measurement. Statistical Science 40 (2), p. 167–194. Cited by: §1. Khan et al. (2025) A. Khan, N. Li, T. Shen, and A. N. Rafferty Just read the question: enabling generalization to new assessment items with text awareness. arXiv preprint arXiv:2507.08154. Cited by: §1.1. Kiela et al. (2021) D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, et al. Dynabench: rethinking benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 4110–4124. Cited by: §1.1. Lalor et al. (2024) J. P. Lalor, P. Rodriguez, J. Sedoc, and J. Hernández-Orallo Item response theory for natural language processing. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts, p. 9–13. Cited by: §1.1, §1. Lalor et al. (2016) J. P. Lalor, H. Wu, and H. Yu Building an evaluation scale using item response theory. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, p. 648–657. Cited by: §1.1, §1, §1. Lalor et al. (2019) J. P. Lalor, H. Wu, and H. Yu Learning latent parameters without human response patterns: item response theory with artificial crowds. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 4249–4259. Cited by: §1.1, §1. Li et al. (2025) Y. Li, J. Ma, M. Ballesteros, Y. Benajiba, and G. Horwood Active evaluation acquisition for efficient LLM benchmarking. In Proceedings of the 42nd International Conference on Machine Learning, p. 35581–35602. Cited by: §1.1, §1. Li et al. (2026) Z. Li, Z. Li, Y. Shi, R. Wang, J. Yang, Z. Liu, X. Wu, A. Li, Y. Yu, N. Liu, et al. Long-horizon-terminal-bench: testing the limits of agents on long-horizon terminal tasks with dense reward-based grading. arXiv preprint arXiv:2607.08964. Cited by: §1. Liang et al. (2023) P. Liang, R. Bommasani, T. Lee, et al. Holistic evaluation of language models. Transactions on Machine Learning Research. Cited by: §1. Lin et al. (2025) B. Y. Lin, Y. Deng, K. Chandu, A. Ravichander, V. Pyatkin, N. Dziri, R. L. Bras, and Y. Choi WildBench: benchmarking LLMs with challenging tasks from real users in the wild. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1.1. Maia Polo et al. (2024) F. Maia Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin TinyBenchmarks: evaluating LLMs with fewer examples. In Proceedings of the 41st International Conference on Machine Learning, p. 34303–34326. Cited by: §1.1, §1, §1. Marinho et al. (2023) W. Marinho, E. W. G. Clua, L. Martí, and K. Marinho Predicting item response theory parameters using question statements texts. In Proceedings of the 13th International Learning Analytics and Knowledge Conference, p. 1–10. Cited by: §1.1, §1. McCarthy et al. (2021) A. D. McCarthy, K. P. Yancey, G. T. LaFlair, J. Egbert, M. Liao, and B. Settles Jump-starting item parameters for adaptive language tests. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 883–899. Cited by: §1.1, §1. Pacchiardi et al. (2024) L. Pacchiardi, L. G. Cheke, and J. Hernández-Orallo 100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances. arXiv preprint arXiv:2409.03563. Cited by: §1.1, §1. Phan et al. (2025) L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al. Humanity’s last exam. arXiv preprint arXiv:2501.14249. Cited by: §1. Reimers and Gurevych (2019) N. Reimers and I. Gurevych Sentence-BERT: sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 3982–3992. Cited by: §3. Shang and Truzzi (2026) E. Shang and F. S. Truzzi ERASE: early backpropagation schedule for faster training of modern recommendation systems. External Links: 2608.18469, Link Cited by: §1. Tang et al. (2026) W. Tang, M. Yuan, Z. Xia, and T. Cai Knowledge-embedded latent projection for robust representation learning. arXiv preprint arXiv:2602.16709. Cited by: Appendix A, §1, §2. Truong et al. (2025) S. T. Truong, Y. Tu, P. Liang, B. Li, and S. Koyejo Reliable and efficient amortized model-based evaluation. In Proceedings of the 42nd International Conference on Machine Learning, p. 60238–60265. Cited by: §1.1, §1, §1, §2, §3.2, §3. Wang et al. (2025) S. Wang, C. Wang, W. Fu, Y. Min, M. Feng, I. Guan, X. Hu, C. He, C. Wang, K. Yang, X. Ren, F. Huang, D. Liu, and L. Zhang Rethinking LLM evaluation: can we evaluate LLMs with 200x less data?. arXiv preprint arXiv:2510.10457. Cited by: §1.1, §1. Wang et al. (2020) W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou MiniLM: deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems 33, p. 5776–5788. Cited by: §3. White et al. (2025) C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Dey, Shubh-Agrawal, S. S. Sandha, S. V. Naidu, C. Hegde, Y. LeCun, T. Goldstein, W. Neiswanger, and M. Goldblum LiveBench: a challenging, contamination-limited LLM benchmark. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1. Yao et al. (2025) L. H. Yao, N. Jarvis, T. Zhan, S. Ghosh, L. Liu, and T. Jiang JE-IRT: a geometric lens on LLM abilities through joint embedding item response theory. arXiv preprint arXiv:2509.22888. Cited by: §1.1. You et al. (2024) J. You, M. Liu, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro LLM-Evolve: evaluation for LLM’s evolving capability on benchmarks. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 16937–16942. Cited by: §1. Zhong et al. (2025) X. Zhong, C. Yi, and H. Ye Efficient evaluation of large language models via collaborative filtering. arXiv preprint arXiv:2504.08781. Cited by: §1.1. Zhou et al. (2025) H. Zhou, H. Huang, Z. Zhao, L. Han, H. Wang, K. Chen, M. Yang, W. Bao, J. Dong, B. Xu, C. Zhu, H. Cao, and T. Zhao Lost in benchmarks? rethinking large language model benchmarking with item response theory. arXiv preprint arXiv:2505.15055. Cited by: §1.1. Zhou et al. (2026) J. Zhou, Z. Sun, B. Li, J. Zhou, Y. Pan, H. Wang, H. Ren, X. Jia, X. Zhou, X. Cao, Y. Chen, Y. Feng, J. Wu, C. Zhang, S. Chen, H. Xue, C. You, H. Wang, K. Wu, P. Gao, J. Wu, W. Li, E. Shang, Q. Zheng, J. Zhou, R. Jia, Y. Xu, H. Zhang, X. Ma, Z. Cheng, Y. Hao, L. Mai, X. Ji, W. Zhang, Z. Chen, Y. Huang, C. Wang, W. Hua, Y. Hao, Y. Zhai, Z. Zhao, and J. Xie ASI-bench: at the dawn of artificial superintelligence. External Links: 2608.17271, Link Cited by: §1. Zhuang et al. (2025) Y. Zhuang, Q. Liu, Z. Pardos, P. C. Kyllonen, J. Zu, Z. Huang, S. Wang, and E. Chen Position: AI evaluation should learn from how we test humans. In Proceedings of the 42nd International Conference on Machine Learning, p. 82483–82508. Cited by: §1. Appendix A Appendix Matrix representation and model identifiability. For a fixed scenario s, let Y(s)∈0,1n×psY^(s)∈\0,1\^n× p_s be the response matrix, let (s)=(α1(s),…,αn(s))⊤∈ℝn, α^(s)=( _1^(s),…, _n^(s)) ^n, let (s)∈ℝn×rU^(s) ^n× r collect the row vectors ui(s)⊤u_i^(s) , and let (s)∈ℝps×rV^(s) ^p_s× r collect the transformed embeddings ϕ(s)(ej)⊤φ^(s)(e_j) . When the latent question factors are estimated freely rather than parameterized through embeddings, the model reduces to (s)=(s)ps⊤+(s)(s)⊤. ^(s)= α^(s)1_p_s +U^(s)V^(s) . In that case, to remove the usual rotational ambiguity, one may impose (s)⊤ps=r,(s)⊤(s)=(s)⊤(s),V^(s) 1_p_s=0_r, ^(s) U^(s)=V^(s) V^(s), under which the factorization is identifiable up to an orthogonal transformation (23). Specifically, if there exists another set of parameters ¯(s),¯(s),¯(s)\ α^(s), U^(s), V^(s)\ satisfying the same constraints such that (s)ps⊤+(s)(s)⊤=¯ps⊤+¯(s)¯(s)⊤ α^(s)1_p_s +U^(s)V^(s) = α1_p_s + U^(s) V^(s) , then (s)=¯(s),(s)=¯(s),and(s)=¯(s),for some ∈(r), α^(s)= α^(s), ^(s)= U^(s)O, ^(s)= V^(s)O, some O (r), where (r)O(r) collects all orthonormal matrices in ℝr×rR^r× r. Estimation procedure. Under the Bernoulli model with logistic link, we estimate the C-MIRT parameters using a two-stage procedure. In the first stage, for each scenario s, we fit a low-rank logistic factorization model to obtain latent row and column representations ((s),(s),(s))( α^(s),U^(s),V^(s)). In the second stage, we learn the mapping ϕ(s)φ^(s) from contextual embeddings to the estimated question factors by supervised regression. For a fixed scenario s, the negative log-likelihood under the Bernoulli model with logistic link is given by ℒ(s)((s),(s),(s))=∑i=1n∑j∈s[log(1+exp(θij(s)))−yij(s)θij(s)], ^(s) ( α^(s),U^(s),V^(s) )= _i=1^n _j _s [ \! (1+ \! ( _ij^(s) ) )-y_ij^(s) _ij^(s) ], (2) where sT_s, the index set, indicates those questions in the training data of scenario s. To address the identifiability issue, we impose the centering constraint (s)⊤ps=r,V^(s) 1_p_s=0_r, and we encourage balanced factorizations through the regularizer g((s),(s))=‖(s)⊤(s)−(s)⊤(s)‖F2.g (U^(s),V^(s) )= \|U^(s) U^(s)-V^(s) V^(s) \|_F^2. Accordingly, in stage 1 we solve the regularized optimization problem min(s),(s),(s)ℒR(s)((s),(s),(s)):=ℒ(s)((s),(s),(s))+14g((s),(s)), _ α^(s),\,U^(s),\,V^(s) _R^(s) ( α^(s),U^(s),V^(s) ):=L^(s) ( α^(s),U^(s),V^(s) )+ 14g (U^(s),V^(s) ), (3) subject to (s)⊤ps=r.V^(s) 1_p_s=0_r. We optimize (3) using projected gradient descent. At each iteration, we update (s) α^(s), (s)U^(s), and (s)V^(s) by gradient descent, followed by a projection step that enforces (s)⊤ps=rV^(s) 1_p_s=0_r. After obtaining the stage-1 estimator ^(s)=(v^1⊤,⋯,v^j⊤,⋯)j∈s∈ℝps×r, V^(s)=( v_1 ,·s, v_j ,·s)_j _s ^p_s× r, we proceed to stage 2 and estimate the feature map ϕ(s)φ^(s) by regressing v^j(s) v_j^(s) on the corresponding contextual embedding eje_j. Specifically, we parameterize ϕ(s)φ^(s) by a multilayer perceptron and minimize the regression loss ℒreg(s)=1|sb|∑j∈sb‖ϕ(s)(ej)−v^j(s)‖22L_reg^(s)= 1|T_s^b| _j _s^b \|φ^(s)(e_j)- v_j^(s) \|_2^2 over mini-batches sb⊆sT_s^b _s. This yields an estimator ϕ^(s) φ^(s), which can then be used to predict latent representations for unseen questions and hence out-of-sample response probabilities through θ^ij(s)=α^i(s)+u^i(s)⊤ϕ^(s)(ej). θ_ij^(s)= α_i^(s)+ u_i^(s) φ^(s)(e_j).