Paper deep dive
UniRank: A Multi-Agent Calibration Pipeline for Estimating University Rankings from Anonymized Bibliometric Signals
Pedram Riyazimehr, Seyyed Ehsan Mahmoudi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 9:16:30 PM
Summary
The paper introduces UniRank, a multi-agent LLM pipeline that estimates university rankings using anonymized bibliometric data from OpenAlex and Semantic Scholar. The system employs a three-stage architecture: zero-shot estimation, tool-augmented calibration, and final synthesis. By anonymizing institutional identities and hiding ground-truth ranks during evaluation, the authors demonstrate that the model performs genuine analytical reasoning rather than memorizing training data, achieving a Memorization Index of zero and significant correlation with actual rankings (Spearman ρ = 0.769) on the THE World University Rankings.
Entities (9)
Relation Signals (7)
UniRank → evaluatedon → Times Higher Education (THE)
confidence 97% · On the Times Higher Education (THE) World University Rankings (n=352), the system achieves...
UniRank → usesdatasource → OpenAlex
confidence 95% · publicly available bibliometric data from OpenAlex and Semantic Scholar
UniRank → usesdatasource → Semantic Scholar
confidence 95% · publicly available bibliometric data from OpenAlex and Semantic Scholar
UniRank → achievesmetric → Memorization Index
confidence 92% · Memorization Index of exactly zero
UniRank → prevents → LLM Memorization
confidence 90% · preventing LLM memorization from confounding results
UniRank → usesmodel → GPT-5.2
confidence 90% · The pipeline uses GPT-5.2
UniRank → inspiredby → MAgICoRe
confidence 85% · A novel three-stage multi-agent architecture (MAgICoRe-inspired)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present UniRank, a multi-agent LLM pipeline that estimates university positions across global ranking systems using only publicly available bibliometric data from OpenAlex and Semantic Scholar. The system employs a three-stage architecture: (a) zero-shot estimation from anonymized institutional metrics, (b) per-system tool-augmented calibration against real ranked universities, and (c) final synthesis. Critically, institutions are anonymized -- names, countries, DOIs, paper titles, and collaboration countries are all redacted -- and their actual ranks are hidden from the calibration tools during evaluation, preventing LLM memorization from confounding results. On the Times Higher Education (THE) World University Rankings ($n=352$), the system achieves MAE = 251.5 rank positions, Median AE = 131.5, PNMAE = 12.03%, Spearman $\rho = 0.769$, Kendall $\tau = 0.591$, hit rate @50 = 20.7%, hit rate @100 = 39.8%, and a Memorization Index of exactly zero (no exact-match zero-width predictions among all 352 universities). The systematic positive-signed error (+190.1 positions, indicating the system consistently predicts worse ranks than actual) and monotonic performance degradation from elite tier (MAE = 60.5, hit@100 = 90.5%) to tail tier (MAE = 328.2, hit@100 = 20.8%) provide strong evidence that the pipeline performs genuine analytical reasoning rather than recalling memorized rankings. A live demo is available at this https URL .
Tags
Links
- Source: https://arxiv.org/abs/2602.18824v1
- Canonical: https://arxiv.org/abs/2602.18824v1
Trouble viewing inline? Open PDF directly →
Full Text
55,595 characters extracted from source content.
Expand or collapse full text
UniRank: A Multi-Agent Calibration Pipeline for Estimating University Rankings from Anonymized Bibliometric Signals Pedram Riyazimehr* Seyyed Ehsan Mahmoudi NotionWave pedram.riazi@notionwave.com (February 2026) Abstract We present UniRank, a multi-agent LLM pipeline that estimates university positions across global ranking systems using only publicly available bibliometric data from OpenAlex and Semantic Scholar. The system employs a three-stage architecture: (a) zero-shot estimation from anonymized institutional metrics, (b) per-system tool-augmented calibration against real ranked universities, and (c) final synthesis. Critically, institutions are anonymized—names, countries, DOIs, paper titles, and collaboration countries are all redacted—and their actual ranks are hidden from the calibration tools during evaluation, preventing LLM memorization from confounding results. On the Times Higher Education (THE) World University Rankings (n=352n=352), the system achieves MAE=251.5=251.5 rank positions, Median AE=131.5=131.5, PNMAE=12.03%=12.03\%, Spearman ρ=0.769ρ=0.769, Kendall τ=0.591τ=0.591, hit rate @50=20.7%=20.7\%, hit rate @100=39.8%=39.8\%, and a Memorization Index of exactly zero (no exact-match zero-width predictions among all 352 universities). The systematic positive signed error (+190.1+190.1 positions, indicating the system consistently predicts worse ranks than actual) and monotonic performance degradation from elite tier (MAE=60.5=60.5, hit@100=90.5%=90.5\%) to tail tier (MAE=328.2=328.2, hit@100=20.8%=20.8\%) provide strong evidence that the pipeline performs genuine analytical reasoning rather than recalling memorized rankings. A live demo is available at https://unirank.scinito.ai. Keywords: university rankings, multi-agent systems, LLM evaluation, bibliometrics, anonymization, tool-augmented reasoning, OpenAlex, ranking estimation, decontamination, MAgICoRe 1 Introduction 1.1 Problem Statement Global university rankings—the QS World University Rankings, Times Higher Education (THE) World University Rankings, and the Academic Ranking of World Universities (ARWU)—are among the most influential instruments in higher education (Hazelkorn, 2015). They shape student enrollment decisions, institutional funding allocations, government policy, and cross-border academic partnerships (Marginson, 2014). Yet these rankings rely heavily on data that is expensive to collect, partially proprietary, and methodologically opaque: • Survey data: QS allocates 45% of its weight to academic and employer reputation surveys (Quacquarelli Symonds, 2025); THE allocates ∼18% 18\% to research reputation (Times Higher Education, 2026). • Institution-reported data: Student–staff ratios, international student counts, and financial data are self-reported with limited independent verification. • Proprietary scoring: Each system uses different weights and normalization procedures, producing substantially different rankings for the same institution. These rankings cover only ∼1,500 1,500 of the world’s ∼30,000 30,000+ higher education institutions, leaving the vast majority unranked (Marginson, 2014). 1.2 Research Question Can a multi-agent LLM pipeline estimate university ranking positions from publicly available bibliometric data alone, without access to survey data, reputation signals, or proprietary institutional data? 1.3 Why This Is Not Simple LLM Memorization A naïve approach—“Ask GPT where MIT ranks”—fails for three reasons: (1) LLMs have seen ranking lists in their training data, making any direct query a memorization exercise (Carlini et al., 2021); (2) training data is temporally stale; (3) a memorized rank provides no methodological insight. Our approach addresses all three: • Anonymization: The LLM receives only an opaque identifier (e.g., INST-A3F2B1C4) with numeric metrics. All identifying information—institution name, country, paper titles, DOIs, collaboration countries—is redacted. • Data hiding: During evaluation, the target university is physically removed from the ranking data store, so calibration tools cannot return the ground-truth answer. • Tool-augmented reasoning: The LLM must use tools to compare the anonymous institution’s metrics against real ranked universities—it cannot retrieve a memorized ranking. 1.4 Contributions 1. A novel three-stage multi-agent architecture (MAgICoRe-inspired) for ranking estimation. 2. A comprehensive anonymization and data-hiding protocol that isolates LLM reasoning from memorized knowledge. 3. A bibliometric feature set (16 indicators) derived entirely from open data sources (OpenAlex, Semantic Scholar). 4. An evaluation framework with stratified test sets, multiple accuracy metrics, and Wilson confidence interval-based accuracy claims. 5. The Memorization Index (MI)—a novel metric for quantifying evidence of LLM memorization in estimation tasks. 6. A live, publicly accessible demo at https://unirank.scinito.ai. 1.5 Paper Organization Section 2 reviews related work. Section 3 describes the system architecture. Section 4 presents the evaluation framework. Section 5 reports results on the THE ranking system. Section 6 discusses memorization evidence and error analysis. Section 7 outlines future work, and Section 8 concludes. 2 Related Work 2.1 University Ranking Systems The three major global ranking systems differ substantially in methodology. QS allocates 50% to research metrics, 45% to reputation surveys, and 5% to sustainability (Quacquarelli Symonds, 2025). THE emphasizes research quality (30%), research environment (29.5%), teaching (15%), international outlook (7.5%), and industry impact (4%) (Times Higher Education, 2026). ARWU uses exclusively hard research metrics: highly cited researchers (20%), Nature/Science publications (20%), total publications (20%), Nobel/Fields winners (30%), and per-capita performance (10%) (Shanghai Ranking Consultancy, 2025). The divergence between systems means the same university can rank very differently depending on how heavily reputation, teaching, and research are weighted (Hazelkorn, 2015; Marginson, 2014). 2.2 Bibliometric Analysis and Scientometrics OpenAlex provides an open catalog of 240M+ scholarly works, 100K+ institutions, and their bibliometric relationships (Priem et al., 2022). Semantic Scholar contributes the influential citation metric—identifying citations that meaningfully impact the citing paper’s methodology or results (Valenzuela et al., 2015). The h-index (Hirsch, 2005) remains widely used despite known limitations (field dependence, career-stage bias). Field-Weighted Citation Impact (FWCI) provides field-normalized impact measurement, and we use the closely related 2-year mean citedness from OpenAlex as a proxy. 2.3 LLMs for Knowledge-Intensive Tasks Tool-augmented LLMs such as ReAct (Yao et al., 2023) and Toolformer (Schick et al., 2023) demonstrate that LLMs can perform complex reasoning when equipped with external tools. Multi-agent architectures—AutoGen (Wu et al., 2024), CAMEL (Li et al., 2023), MetaGPT (Hong et al., 2024)—show that distributing tasks across specialized agents improves performance on complex problems. MAgICoRe (Multi-Agent, Iterative, Coarse-to-Fine Refinement) (Chen et al., 2025) provides the direct architectural inspiration for our pipeline, demonstrating that iterative refinement across agents produces better results than single-pass estimation. 2.4 LLM Evaluation and Decontamination The risk of training data contamination in LLM evaluation is well-documented (Carlini et al., 2021; Magar and Schwartz, 2022). Our anonymization-based approach complements existing decontamination techniques (n-gram filtering, canary insertion) by operating at the input level: the LLM never sees identifiable information, making memorization-based shortcuts impossible regardless of what the model has seen during training. 3 System Architecture 3.1 Architecture Overview Figure 1 presents the end-to-end UniRank pipeline. Data from OpenAlex and Semantic Scholar is aggregated, normalized, and anonymized before entering the three-stage LLM pipeline. During evaluation, the target university is hidden from the ranking store to prevent data leakage. Figure 1: UniRank system architecture. Data from OpenAlex and Semantic Scholar is aggregated, normalized, and anonymized before entering the three-stage LLM pipeline. During evaluation, the target university is hidden from the ranking store (dashed line) to prevent data leakage. 3.1.1 Formal Problem Definition Let U be the set of all universities. For each ranking system s∈QS,THE,ARWUs∈\QS,THE,ARWU\, let πs:→ℕ _s:U be the ground-truth ranking function. For a target university u∗∈u^* , we observe a feature vector u∗∈ℝdx_u^* ^d computed from publicly available bibliometric data, where d=16d=16 features (Table 1). For each system s, we produce a predicted rank range [r^smin,r^smax][ r_s , r_s ] and a point estimate r^s=(r^smin+r^smax)/2 r_s=( r_s + r_s )/2. Let α:→′α:U be the anonymization function that maps u∗u^* to an opaque identifier while preserving only u∗x_u^*. During evaluation of u∗u^*, the ranking store ℛsR_s is modified to ℛs∖u∗R_s \u^*\. Objective: Minimize MAE=1||∑u∈|r^s(u)−πs(u)|MAE= 1|T| _u | r_s(u)- _s(u)| over a stratified test set ⊂T , subject to the anonymization and data-hiding constraints. 3.2 Data Sources OpenAlex (primary): We query institution profiles, disciplinary distributions, international collaboration counts, open access breakdowns, top cited works with FWCI, and research excellence counts (top 10% by FWCI). All queries are filtered to publication years 2020–2025 (Priem et al., 2022). Semantic Scholar (enrichment): For the top 10 cited works, we retrieve influential citation counts and ratios via the batch endpoint (Valenzuela et al., 2015). The influential citation ratio (influential/total) provides a quality signal unavailable from OpenAlex. Ground truth rankings: QS 2026, THE 2024, and ARWU 2025 data from publicly available CSV datasets serve as both calibration anchors and evaluation ground truth. 3.3 Metric Computation and Normalization Table 1 lists the 16 bibliometric indicators computed for each institution. Table 1: Bibliometric feature set (16 indicators). Metric Source Description worksCount OA Total publications citedByCount OA Total citations hIndex OA Hirsch index i10Index OA Papers with 10+ citations 2yr mean citedness OA FWCI proxy citationsPerWork OA Average citations/paper 5yr works growth OA Publication growth (%) 5yr citation growth OA Citation growth (%) researchExcellence% OA % in global top 10% FWCI intlCollaboration% OA % with intl. co-authors openAccess% OA Open access rate disciplinaryBreadth OA Shannon entropy normalizedResearch OA 0–100 research score normalizedImpact OA 0–100 impact score normalizedExcellence OA 0–100 excellence score influentialRatio S2 Influential/total citations Each metric is normalized against global tier benchmarks using: score=clamp(v−floorceiling−floor×100, 0, 100)score=clamp\! ( v-floorceiling-floor× 100,\;0,\;100 ) (1) where floor=p25floor=p_25 of Tier 4 (Top 500) and ceiling=p75ceiling=p_75 of Tier 1 (Top 10). For each metric, we additionally compute a Z-score relative to the closest tier: z=v−μclosestσclosestz= v- _closest _closest (2) producing signals like “h-index: +1.2σ+1.2σ from Tier 2 mean.” 3.4 Three-Stage Multi-Agent Pipeline Figure 2 illustrates the pipeline flow, inspired by MAgICoRe’s coarse-to-fine refinement pattern (Chen et al., 2025). Figure 2: Three-stage pipeline: Stage 1 produces coarse zero-shot estimates from anonymized metrics. Stage 2 refines per-system with tool-augmented calibration (parallel). Stage 3 synthesizes the final report. Stage 1: Initial Zero-Shot Estimation. The LLM receives anonymized metrics (∼2,000 2,000 tokens) and produces rank ranges for each system. No tools are available; estimation relies solely on the model’s understanding of each ranking methodology and the provided Z-scores. Stage 2: Per-System Calibration (Parallel). Three independent calibration engines run in parallel (one per system). Each has access to two tools: • get_ranking_samples(system, rankMin, rankMax, count): Returns real universities within the specified rank range, including names, actual ranks, and official sub-scores. During evaluation, the target is hidden from this data. • compute_metrics(universityName): Computes the full bibliometric feature set for a named university using the same pipeline, enabling direct metric comparison. Each engine executes up to 12 agentic steps: fetching samples, computing metrics, comparing, adjusting, and producing a calibrated range (target width: 30–50 positions). Stage 3: Final Report Synthesis. This stage receives the full (non-anonymized) university data along with Stage 1 and Stage 2 results, producing a structured analysis report. 3.5 Anonymization Protocol This is the core contribution addressing memorization concerns. The anonymization function transforms all identifying information before any LLM interaction (Figure 3): Figure 3: Anonymization before/after comparison. All identifying information (name, country, DOIs, paper titles, collaboration countries) is redacted; only numeric metrics are preserved. • Institution name → random hex identifier (INST-A3F2B1C4) • Country → REDACTED • Paper titles → [Work 1], [Work 2], … • DOIs → null • Collaboration countries → [Country 1], [Country 2], … • Published rankings (QS/THE/ARWU) → null • Semantic Scholar TLDRs → null Preserved fields (numeric only): all computed metrics, Z-scores, tier-relative positioning, yearly trends, field distribution percentages, citation counts, FWCI, citation percentiles, and influential citation data. 3.6 Model Configuration The pipeline uses GPT-5.2 (via @ai-sdk/openai) with reasoning effort set to medium. Structured outputs are enforced via Zod schemas through the Vercel AI SDK. The Helicone AI Gateway provides logging and monitoring. 4 Evaluation Framework 4.1 Test Set Construction The test set comprises 500 universities with OpenAlex ID resolution, constructed via stratified sampling across two dimensions: Tier stratification: Elite (ranks 1–25, 20%), Strong (26–150, 20%), Mid (151–400, 20%), Lower (401–700, 20%), Tail (700+, 20%). Tier assignment is based on consensus rank (median of available QS/THE/ARWU ranks). Regional stratification: North America (∼20% 20\%), Europe (∼25% 25\%), Asia-Pacific (∼25% 25\%), Rest of World (∼15% 15\%), Mixed/Other (∼15% 15\%). 4.2 Double-Blind Evaluation Protocol For each target university, evaluation proceeds in four steps: (1) Data hiding: the target’s record is physically removed from the in-memory ranking store, verified before proceeding; (2) Anonymization: the standard protocol (Section 3.5) strips all identifying information; (3) Pipeline execution: the three-stage pipeline runs against the anonymized, hidden data; (4) Restoration: the hidden university is restored to the ranking store. 4.3 Evaluation Metrics 4.3.1 Primary Metrics MAE =1n∑i=1n|mi−ai| = 1n _i=1^n|m_i-a_i| (3) Median AE =median(|mi−ai|) =median(|m_i-a_i|) (4) PNMAE =1n∑i=1n|mi−1N−1−ai−1N−1|×100 = 1n _i=1^n | m_i-1N-1- a_i-1N-1 |× 100 (5) where mim_i is the predicted midpoint, aia_i is the actual rank, and N is the total number of ranked universities in the system. 4.3.2 Hit Rate Metrics Hit@k=1n∑i=1n[|mi−ai|≤k]×100Hit@k= 1n _i=1^n1[|m_i-a_i|≤ k]× 100 (6) for k∈25,50,100k∈\25,50,100\. 4.3.3 Range Metrics Range coverage measures the fraction of actual ranks falling within the predicted range: 1n∑[rmin≤ai≤rmax] 1n 1[r_ ≤ a_i≤ r_ ]. Mean range width is 1n∑(rmax−rmin) 1nΣ(r_ -r_ ). 4.3.4 Memorization Index (MI) To explicitly quantify evidence of memorization: MI=|u∈:AE(u)=0∧W(u)=0|||MI= |\u :AE(u)=0 W(u)=0\||T| (7) where W(u)=r^max(u)−r^min(u)W(u)= r (u)- r (u) is the range width. A non-zero MI indicates suspiciously perfect predictions with zero-width ranges—a hallmark of memorized retrieval. We expect MI ≈0≈ 0 for a reasoning-based system. 4.3.5 Correlation and Error Decomposition Spearman’s ρ, Pearson’s r, and Kendall’s τ measure ordinal and linear correlation. We additionally report RMSE, signed error (directional bias), calibration slope β from OLS regression (r^=α+β⋅r r=α+β· r; β=1β=1 indicates perfect calibration), and Cohen’s κ for tier classification agreement. 4.4 Accuracy Claims with Confidence Intervals To avoid overclaiming, we compute Wilson score interval lower bounds (Wilson, 1927): L=p^+z22n−zp^(1−p^)+z24n1+z2nL= p+ z^22n-z p(1- p)+ z^24nn1+ z^2n (8) with z=1.96z=1.96 (95% CI). The claimed accuracy is ⌊L×100⌋% L× 100 \%. 4.5 Statistical Testing Plan • Wilcoxon signed-rank: Compare initial vs. calibrated absolute errors (paired). • McNemar’s test: Compare hit rates before/after calibration. • Kruskal–Wallis: Compare errors across tiers. All reported metrics include 95% bootstrapped confidence intervals (B=10,000B=10,000 resamples). 5 Results We present results on the THE World University Rankings—our primary evaluation target. The evaluation ran on 357 universities from the THE stratified test set; 352 produced successful predictions and 5 failed (pipeline errors, 1.4% failure rate). 5.1 Aggregate Metrics Table 2 reports aggregate performance over n=352n=352 successful predictions. THE contains approximately 2,092 ranked universities. Table 2: Aggregate evaluation metrics (THE, n=352n=352). Metric Value MAE 251.5 Median AE 131.5 RMSE 411.4 PNMAE 12.03% Signed Error +190.1+190.1 Spearman’s ρ 0.769 Pearson’s r 0.677 Kendall’s τ 0.591 Hit Rate @25 10.2% (36/352) Hit Rate @50 20.7% (73/352) Hit Rate @100 39.8% (140/352) Wilson-claimed @50 ≥ 16% Range Coverage 8.2% (29/352) Mean Range Width 42.9 positions Calibration Slope β 0.951 Calibration Intercept α 209.9 Cohen’s κ (tier) 0.349 (fair) Tier Agreement 49.7% (175/352) Memorization Index 0.000 (0/352) AE=0=0 (exact matches) 1 (Northwestern U.) The positive signed error (+190.1+190.1) indicates the system systematically predicts worse ranks (higher numbers) than actual. This is expected: bibliometrics alone cannot capture reputation and teaching signals, which tend to improve a university’s actual rank beyond what research metrics predict. 5.2 Per-Tier Performance Table 3 breaks down performance by evaluation tier. MI=0.000=0.000 across all tiers. Table 3: Per-tier performance breakdown (THE). Tier N MAE Med. @50 @100 SE Elite 21 60.5 33.0 57.1 90.5 ++60.2 Strong 80 211.6 93.0 30.0 53.8 ++205.8 Mid 99 247.8 122.5 18.2 38.4 ++214.0 Lower 99 287.0 192.0 11.1 29.3 ++198.5 Tail 53 328.2 309.0 15.1 20.8 ++157.2 Performance degrades monotonically from elite to tail, consistent with decreasing bibliometric signal density for lower-ranked universities. Elite tier performance (57.1% hit@50, 90.5% hit@100) contradicts what memorization would predict: if the LLM were memorizing, elite (most famous) universities would be recalled perfectly (AE≈ 0), not with MAE=60.5=60.5. Figure 4 shows the error distribution by tier, and Figure 5 compares hit rates across tiers and thresholds. Figure 4: Error distribution by evaluation tier (THE). Box plots show median (line), IQR (box), and outliers. Performance degrades monotonically from elite to tail. Figure 5: Hit rate comparison across tiers at three thresholds (@25, @50, @100). Elite tier achieves 90.5% hit@100. Tier confusion matrix. Table 4 shows actual vs. predicted tier classification. Table 4: Tier confusion matrix (actual × predicted). Predicted Actual Elite Strong Mid Lower Tail Elite 4 15 2 0 0 Strong 0 36 32 6 6 Mid 0 6 49 26 18 Lower 0 1 13 40 45 Tail 0 0 2 5 46 5.3 Predicted vs. Actual Rank Correlation Figure 6 plots predicted midpoint against actual rank for all 352 universities, colored by tier. The Spearman correlation (ρ=0.769ρ=0.769) and calibration slope (β=0.951β=0.951) indicate meaningful ordinal agreement with remarkably low systematic compression bias. Figure 6: Predicted vs. actual rank (THE, n=352n=352). Diagonal = perfect prediction; shaded band = ±50± 50 positions. Spearman ρ=0.769ρ=0.769, calibration slope β=0.951β=0.951. 5.4 Failure Taxonomy All failures share a common root cause: the system estimates rankings using research quality proxies alone, while THE incorporates industry impact (4%), reputation surveys (∼18% 18\%), and teaching quality (∼29.5% 29.5\%) that are entirely unobservable to our pipeline. Table 5 defines six failure modes; Table 5 summarizes their prevalence among the 20 worst predictions. Table 5: Failure mode definitions (left) and prevalence in top-20 worst predictions (right). ID Mode Description F1 Reputation Rank boosted by reputation/teaching signals unobservable from bibliometrics F2 Teaching Rank driven by teaching metrics (student–staff ratio, doctoral ratio) F3 Scale Small-but-excellent or large-but-average confusion F4 Regional Systematic regional over/under-estimation F5 Anchor Drift Calibration anchors on unrepresentative comparators F6 Data Gap OpenAlex provides fragmented or incorrect institutional data Mode Primary % Key pattern F1: Reputation 13 65% Strong brand, weak metrics F6: Data Gap 7 35% OpenAlex inconsistency F5: Anchor Drift 5 (2nd) 25% Amplifies F1/F6 F4: Regional 5 (2nd) 25% Russian, French institutions F3: Scale 3 (2nd) 15% Small specialized schools 5.5 Case Studies We present three representative case studies spanning distinct failure modes. University of Toulouse (F6, AE=1,879=1,879). OpenAlex fragments this multi-campus system into separate entities; the pipeline retrieved data for only Université Toulouse I – Paul Sabatier, yielding 2yr citedness=0.09=0.09 and h-index=18=18—catastrophically unrepresentative. The LLM’s reasoning was sound given its inputs; the inputs were wrong. Yale University (F1, AE=300=300). Yale’s 2yr citedness of 1.87 places it well below institutions like MIT (5.74) and Caltech (4.91) on pure research impact. The LLM correctly estimated ∼290−330 290-330 from bibliometrics alone, but Yale’s THE rank of #10 is driven substantially by reputation—one of the most recognized university brands globally. This error is entirely attributable to unobservable reputation signals. University of Michigan (F6+F1+F5, AE=982=982). Research excellence of 5.0% and international collaboration of 8.1% are implausibly low for THE #23. The initial estimate (∼200−500 200-500) was already compromised by bad input data (F6). Calibration then compared against #360–#1250 universities, found metric matches around #1000, and drifted the estimate to 980–1030 (F5), amplifying the error from ∼350 350 to 982—a textbook case of cascading failure. 5.6 Statistical Test Results Table 6 reports statistical comparisons between initial (Stage 1) and calibrated (Stage 1+2) predictions. Table 6: Statistical test results (initial vs. calibrated). Test Stat. p Interpretation Wilcoxon W=27,921W=27,921 0.343 Not significant McNemar χ2=1.786χ^2=1.786 0.181 Not significant Kruskal–Wallis H=42.09H=42.09 10−810^-8 Tiers differ∗ Calibration impact. Table 7 summarizes the effect of calibration. Overall, calibration produces a marginal improvement (MAE: 256.8→251.5256.8→ 251.5, −2.1%-2.1\%) that is not statistically significant (p=0.343p=0.343). However, the effect is bimodal: calibration reliably improves elite predictions (85.7% improved) and tail predictions (56.6% improved), but has mixed effects on mid-tier universities. Table 7: Calibration impact summary. Metric Initial Calibrated Δ MAE 256.8 251.5 −2.1%-2.1\% Median AE 149.5 130.8 −12.5%-12.5\% Spearman ρ 0.770 0.769 −0.001-0.001 Hit @50 17.9% 20.7% +2.8+2.8 p Hit @100 36.1% 39.8% +3.7+3.7 p Improved 191/352 (54.3%) Worsened 153/352 (43.5%) Unchanged 8/352 (2.3%) Figure 7 visualizes the per-university calibration effect. Figure 7: Initial vs. calibrated absolute error. Points below the diagonal indicate calibration helped. Calibration improved 54.3% of predictions overall. 6 Discussion 6.1 Beyond Memorization: Evidence and Arguments We structure the anti-memorization argument across six lines of evidence: 1. Anonymization prevents name-based recall. The LLM receives INST-A3F2B1C4 with numeric metrics only. Even if a metric profile seems familiar, the model must reason about rank implications—not retrieve a memorized answer. 2. Data hiding prevents tool-based leakage. The get_ranking_samples tool physically cannot return the target university’s rank. 3. Calibration requires reasoning, not retrieval. The multi-step process of fetching samples, computing metrics, comparing, and determining relative positioning is an analytical workflow, not a lookup. 4. Large errors on famous universities. The system produces its largest errors on well-known institutions—the opposite of what memorization would predict. Yale (THE #10, predicted 310, AE=300=300), Oxford (THE #1, predicted 95, AE=94=94), Michigan (THE #23, predicted 1005, AE=982=982), and UT Austin (THE #50, predicted 1813, AE=1,763=1,763) are among the most-discussed universities in training corpora. A memorizing system would recall these perfectly. 5. Memorization Index. MI=0.000=0.000 (0/352). Zero universities received an exact point prediction with a zero-width range. The single AE=0=0 case (Northwestern University) had a range width of 55 positions ([328, 383], actual=355=355)—a fortunate analytical prediction, not memorized retrieval. 6. Systematic positive signed error. The system consistently predicts worse ranks than actual (+190.1+190.1) across all tiers, consistent with missing reputation/teaching signals rather than memorized recall (which would produce signed error ≈0≈ 0). 6.2 Observable vs. Unobservable Data Figure 8 illustrates the fundamental challenge: UniRank can observe research metrics well but has zero signal on reputation, teaching, and industry dimensions. Figure 8: Data completeness radar. Green = what UniRank can measure; gray = what THE weights. The gap between series—particularly on Teaching (15%), Reputation (18%), and Industry (4%)—explains the systematic positive signed error. Table 8 details the missing data dimensions and their impact. Table 8: Missing data dimensions and ranking impact. Missing Data Impact Systems Reputation surveys 30% QS, 18% THE QS, THE Employer surveys 15% QS QS Student–staff ratios 5% QS, 15% THE QS, THE Intl. student/faculty 10% QS QS Nobel/Fields winners 30% ARWU ARWU Industry income 4% THE THE 6.3 Limitations 1. Sample size: 352 successful predictions from THE (∼17% 17\% of 2,092 ranked universities). 2. Single system: Results are for THE only; QS and ARWU evaluations remain future work. 3. Model dependency: Results are specific to GPT-5.2; different models may yield different results. 4. Temporal validity: Rankings and bibliometric data evolve annually. 5. Cost: Each evaluation requires 4–6 LLM calls with up to 12 tool-call steps each. 6. Single-run results: Due to LLM non-determinism, results may vary across runs. We report single-run results without repeated trials. 7 Future Work 1. Multi-system evaluation: Run stratified test sets across QS and ARWU. 2. Additional data sources: Student/staff data from government databases, patent data from Lens.org, Nobel/Fields data for ARWU, webometrics. 3. LLM-driven reputation gathering: Use agentic web search and deep-research workflows to let the LLM retrieve real-time reputation signals—news coverage, employer surveys, faculty awards. 4. Adaptive calibration: Select comparison universities by metric similarity rather than rank proximity. 5. TOPSIS-based aggregation: Integrate the Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) as a complementary multi-criteria decision-making layer, leveraging its ability to rank alternatives by geometric distance from ideal and anti-ideal profiles across heterogeneous indicator dimensions—providing a mathematically grounded aggregation baseline against which the LLM-based calibration can be compared and combined. 8 Conclusion We presented UniRank, a multi-agent calibration pipeline that estimates university rankings from publicly available bibliometric data, using anonymization and data hiding to isolate LLM reasoning from training data memorization. On the THE World University Rankings (n=352n=352), UniRank achieves MAE=251.5=251.5 rank positions, Spearman ρ=0.769ρ=0.769, hit rate @100=39.8%=39.8\%, and a Memorization Index of exactly zero. Performance is strongest in the elite tier (MAE=60.5=60.5, hit@100=90.5%=90.5\%) and degrades predictably toward the tail (MAE=328.2=328.2, hit@100=20.8%=20.8\%). The systematic positive signed error (+190.1+190.1) across all tiers confirms the system’s core limitation: it cannot observe reputation and teaching signals that account for approximately 50% of the THE ranking methodology. Despite this, the correlation metrics (ρ=0.769ρ=0.769, τ=0.591τ=0.591, r=0.677r=0.677) demonstrate meaningful ordinal agreement, and the calibration slope (β=0.951β=0.951) shows remarkably low systematic compression bias. Our failure analysis reveals two dominant error sources: F1 (Reputation Blind Spot)—accounting for 65% of the top-20 worst predictions—and F6 (Data Source Incompleteness)—35% of the top-20 worst. The primary contribution of this work is not a production ranking system, but an evaluation framework. UniRank establishes a rigorous, reproducible methodology—including the leave-one-out protocol, anonymization-based decontamination, memorization index measurement, and multi-dimensional failure taxonomy—that enables systematic assessment of future improvements. As we incorporate more complete data sources for reputation, teaching quality, and graduate outcomes, this evaluation framework will measure whether each addition genuinely improves ranking estimation or merely adds noise. The current results (Spearman ρ=0.769ρ=0.769 from research proxies alone) provide a clear baseline against which all future enhancements can be benchmarked. A live demo is available at https://unirank.scinito.ai. Acknowledgments We thank OpenAlex and Semantic Scholar for providing open access to scholarly metadata, and the QS, THE, and ARWU organizations for publishing their ranking methodologies. References N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel (2021) Extracting Training Data from Large Language Models. In 30th USENIX Security Symposium, p. 2633–2650. Cited by: §1.3, §2.4. J. C. Chen, A. Prasad, S. Saha, E. Stengel-Eskin, and M. Bansal (2025) MAgICoRe: multi-agent, iterative, coarse-to-fine refinement for reasoning. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), Miami, Florida, p. 32663–32686. External Links: Link, 2409.12147 Cited by: §2.3, §3.4. E. Hazelkorn (2015) Rankings and the Reshaping of Higher Education: The Battle for World-Class Excellence. Cited by: §1.1, §2.1. J. E. Hirsch (2005) An index to quantify an individual’s scientific research output. Proceedings of the National Academy of Sciences 102 (46), p. 16569–16572. Cited by: §2.2. S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber (2024) MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In The Twelfth International Conference on Learning Representations (ICLR), External Links: 2308.00352 Cited by: §2.3. G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: Communicative Agents for “Mind” Exploration of Large Language Model Society. In Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS), External Links: 2303.17760 Cited by: §2.3. I. Magar and R. Schwartz (2022) Data Contamination: From Memorization to Exploitation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), p. 157–168. Cited by: §2.4. S. Marginson (2014) University rankings and social science. European Journal of Education 49 (1), p. 45–59. Cited by: §1.1, §1.1, §2.1. J. Priem, H. Piwowar, and R. Orr (2022) OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv preprint arXiv:2205.01833. External Links: 2205.01833 Cited by: §2.2, §3.2. Quacquarelli Symonds (2025) QS World University Rankings 2026: Methodology. Note: Accessed February 2026 External Links: Link Cited by: 1st item, §2.1. T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom (2023) Toolformer: Language Models Can Teach Themselves to Use Tools. External Links: 2302.04761 Cited by: §2.3. Shanghai Ranking Consultancy (2025) Academic Ranking of World Universities 2025: Methodology. Note: Accessed February 2026 External Links: Link Cited by: §2.1. Times Higher Education (2026) THE World University Rankings 2026: Methodology. Note: Accessed February 2026 External Links: Link Cited by: 1st item, §2.1. M. Valenzuela, V. Ha, and O. Etzioni (2015) Identifying Meaningful Citations. In AAAI Workshop on Scholarly Big Data, Cited by: §2.2, §3.2. E. B. Wilson (1927) Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association 22 (158), p. 209–212. Cited by: §4.4. Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang (2024) AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. In Conference on Language Modeling (COLM), External Links: 2308.08155 Cited by: §2.3. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao (2023) ReAct: Synergizing Reasoning and Acting in Language Models. External Links: 2210.03629 Cited by: §2.3. Appendix A Reproducibility Checklist Item Detail Model GPT-5.2 (via @ai-sdk/openai, Helicone gateway) Reasoning effort medium Temperature Default (model-determined) Max calibration steps 12 per system Structured output Zod schemas via Vercel AI SDK Eval concurrency 3 parallel evaluations Eval delay 2000ms between batches Non-determinism Single-run results; may vary across runs Data snapshot OpenAlex Feb 2026; QS 2026, THE 2024, ARWU 2025 Code availability Live demo at https://unirank.scinito.ai Appendix B Full Tier Benchmark Table Table 9 shows the empirical benchmarks derived from OpenAlex data for four tiers. Table 9: Tier benchmarks (mean ± SD) for key metrics. Metric Tier 1 (Top 10) Tier 2 (Top 100) Tier 3 (Top 300) Tier 4 (Top 500) h-index 2135 ± 291 1196 ± 205 787 ± 145 560 ± 102 2yr citedness 4.60 ± 1.10 3.40 ± 0.70 2.95 ± 0.60 2.85 ± 0.77 Research excellence (%) 19.5 ± 3.5 20.5 ± 2.1 20.0 ± 2.8 19.5 ± 3.5 Intl. collab (%) 40.5 ± 9.2 40.5 ± 12.0 39.0 ± 12.7 38.0 ± 12.7 Figure 9: Tier benchmarks for key bibliometric indicators. Error bars show ±1± 1 SD. Appendix C Prompt Templates C.1 Stage 1: Initial Estimator Listing 1: Initial estimation prompt (abbreviated). ⬇ You are a university ranking estimation engine. You receive anonymized bibliometric data about an institution. Your job is to estimate where this institution would rank in QS, THE, and ARWU. ## What You Know About Each System ### QS 2025 (50% Research, 45% Reputation, 5% Sustainability) - Academic Reputation (30%): Proxy from h-index, citation volume - Citations per Faculty (20%): Citation metrics, NORMALIZED - Employer Reputation (15%): Infer from institution type - International Research Network (5%): International collaboration rate ### THE 2025 (30% Research Quality, 29.5% Env., 15% Teaching) - Research Excellence (5%): % of works in top 10% by FWCI - Citation Impact (15%): FWCI-based - Research Reputation (18%): Infer from h-index + citation volume - International Outlook (7.5%): Collaboration % ### ARWU (100% Hard Research Metrics, no reputation) - Highly Cited Researchers (20%): h-index as proxy - Nature/Science papers (20%): Top venue publications - Total publications (20%): Works count - Nobel/Fields (30%): Not in data ## Rules - Estimating based on METRICS ONLY - All data is anonymized -- do NOT try to identify the institution - Produce a RANGE (min-max), never a point estimate - Include confidence: "high" (+/-20), "medium" (+/-50), "low" (+/-100+) C.2 Stage 2: THE Calibration Engine Listing 2: THE calibration prompt (abbreviated). ⬇ You are a THE World University Rankings calibration engine. You receive anonymized bibliometric metrics and an initial rank estimate. Refine this estimate by comparing against real universities in THE. ## THE Ranking Weights - Research Quality (30%): Research Excellence %, FWCI, influential cites - Research Environment (29.5%): h-index, productivity, reputation - Teaching (15%): Not directly measurable - International Outlook (7.5%): International collaboration % - Industry Impact (4%): Limited data ## Calibration Protocol 1. Start with the initial estimate range 2. Call get_ranking_samples to fetch 3 universities near estimate center 3. Call compute_metrics for 2-3 sample universities 4. Compare target’s metrics against each sample 5. Decide if target ranks ABOVE, BETWEEN, or BELOW samples 6. Adjust range; MAY move outside initial estimate if warranted 7. Produce final calibrated range (target width: 30-50 positions) ## Rules - ALWAYS use tools; budget of 10 tool-call steps - Aim for range width of 30-50 positions - DO NOT try to identify the target institution C.3 Stage 3: Final Report Synthesis Listing 3: Final report prompt (abbreviated). ⬇ You are UniRank, an expert analyst in global higher education rankings. Writing the FINAL REPORT for a university analysis. The rank estimation has been done; you receive the results. Synthesize into a structured analysis. ## What You Receive 1. Full university profile with bibliometric data (unblinded) 2. Initial zero-shot estimate (Stage 1) 3. Calibration results from per-system engines (Stage 2) 4. Published rankings for systems where university is already ranked ## Output Format - Estimated Global Rank Range - Methodology Breakdown (per system) - Strengths (3-5, each referencing a metric) - Weaknesses (3-5, each referencing a metric or missing data) - Citation Quality Assessment (when S2 data available) - Comparable Peers (3-5 universities) - Strategic Recommendations (2-3 actionable suggestions) ## Rules - NEVER invent specific ranking numbers - ALWAYS cite specific data values - When data is MISSING, explicitly state this Appendix D Test Set Sampling Algorithm Input: Universe U of universities with OpenAlex IDs; target size N Output: Stratified test set T 1 1ex// Compute consensus rank for each university 2 foreach u∈u do 3 ru←median(rs(u):s∈QS,THE,ARWU,rs(u)≠null)r_u (\r_s(u):s∈\QS,THE,ARWU\,r_s(u) \); 4 5 1ex// Assign tiers 6 tiers←tiers←\Elite: [1,25][1,25], Strong: [26,150][26,150], Mid: [151,400][151,400], Lower: [401,700][401,700], Tail: [700+][700+]\; 7 foreach u∈u do 8 tier(u)←tier(u)← tier containing rur_u; 9 10 1ex// Assign regions from country code 11 regions←regions←\NA, Europe, Asia-Pacific, Rest of World, Mixed\; 12 foreach u∈u do 13 region(u)←mapCountryToRegion(country(u))region(u) (country(u)); 14 15 1ex// Stratified sampling: equal tier representation 16 ←∅T← ; 17 nper_tier←N/|tiers|n_per\_tier← N/|tiers|; 18 foreach t∈tierst do 19 t←u:tier(u)=tU_t←\u:tier(u)=t\; 20 ←∪sampleProportional(t,nper_tier,by=region)T (U_t,n_per\_tier,by=region); 21 return T Algorithm 1 Stratified Test Set Construction Appendix E Full Results Table Table LABEL:tab:full_results presents all 352 successful predictions sorted by absolute error (ascending). Best prediction: Northwestern University (AE=0=0). Worst prediction: University of Cape Town (AE=2,034=2,034). Table 10: Full evaluation results (THE, n=352n=352), sorted by AE ascending. # University Actual Pred. AE Tier 1 Northwestern University 30 30 0 strong 2 Tianjin University 208 208 1 mid 3 Univ. of Kwazulu-Natal 566 565 1 lower 4 Hong Kong Baptist Univ. 214 215 1 mid 5 Vrije Univ. Brussel 234 235 1 mid 6 Univ. of Alberta 119 120 1 strong 7 Univ. of Glasgow 84 85 1 strong 8 Lithuanian U. Health Sci. 1006 1005 1 tail 9 Brunel Univ. London 414 413 2 lower 10 Univ. of Maryland 117 115 2 strong 11 Nanyang Technological U. 31 28 4 elite 12 Hong Kong Polytechnic U. 82 88 6 strong 13 Harvard University 6 13 7 elite 14 Univ. of Minho 663 655 8 lower 15 New Jersey Inst. Tech. 591 600 9 tail 16 Durham University 177 168 10 mid 17 Univ. of Porto 415 405 10 mid 18 U. Politècnica Catalunya 670 660 10 lower 19 Univ. of Hull 525 515 10 lower 20 Gachon University 596 585 11 lower 21 East China Normal Univ. 286 275 11 mid 22 Univ. of Groningen 83 70 13 strong 23 UCL 22 35 13 elite 24 MIT 2 16 14 elite 25 Stanford University 5 20 15 elite 26 Univ. of Sci. & Tech. China 51 68 17 strong 27 Univ. of Calgary 202 185 17 mid 28 Old Dominion University 950 970 20 tail 29 Princeton University 3 24 21 elite 30 Cornell University 19 40 21 elite 31 Imperial College London 8 30 22 elite 32 Univ. of Padua 242 220 22 mid 33 Bharathidasan Univ. 1238 1215 23 tail 34 Univ. of Birmingham 100 78 23 strong 35 Heinrich Heine U. Düsseldorf 299 275 24 mid 36 Linköping University 241 265 24 mid 37 Heidelberg University 49 75 26 strong 38 NUS 17 43 26 strong 39 Wuhan University 124 98 27 strong 40 Univ. of Copenhagen 90 118 28 strong 41 Tsinghua University 12 40 28 elite 42 Southern U. Sci. & Tech. 162 190 28 mid 43 Queensland U. of Tech. 210 180 30 mid 44 Boğaziçi University 486 455 31 lower 45 King Saud University 266 235 31 mid 46 Univ. of Pennsylvania 14 45 31 elite 47 Univ. of Queensland 81 50 32 strong 48 Boston University 77 45 32 strong 49 Caltech 7 40 33 elite 50 Univ. of Amsterdam 62 95 33 strong 51 Hanyang University 295 263 33 mid 52 Zewail City 1416 1450 34 tail 53 Univ. of Trento 360 395 35 lower 54 Univ. of Melbourne 37 73 36 strong 55 Halmstad University 940 978 38 tail 56 Stony Brook University 313 275 38 mid 57 U. Teknologi Malaysia 437 475 38 mid 58 Univ. of Cambridge 4 43 39 elite 59 Sorbonne University 76 115 39 strong 60 Trinity College Dublin 175 135 40 mid 61 Carleton University 580 620 40 lower 62 Univ. of Freiburg 139 180 41 strong 63 Univ. of Minnesota 88 130 42 strong 64 Univ. of Delhi 802 760 42 lower 65 Natl. Taiwan Ocean U. 1449 1495 46 tail 66 U. Libre de Bruxelles 224 270 46 mid 67 Univ. of Liverpool 144 190 46 strong 68 Aston University 427 380 47 lower 69 The New School 1203 1250 47 tail 70 Univ. of Sydney 53 100 47 strong 71 Western Australia 154 105 49 strong 72 EPFL 35 85 50 strong 73 Shanghai Jiao Tong U. 40 90 50 strong 74 Kyung Hee University 254 305 51 mid 75 NW Polytechnical U. 274 325 51 mid 76 Univ. of Dundee 325 380 55 mid 77 Adelaide University 134 79 56 strong 78 Nanjing U. Info. Sci. 929 985 56 lower 79 Missouri U. Sci. & Tech. 549 605 56 lower 80 Xiamen University 276 220 56 mid 81 Univ. of Plymouth 553 498 56 lower 82 Univ. of Waikato 461 518 57 lower 83 Ulsan NIST 243 300 57 mid 84 Johns Hopkins Univ. 16 75 59 elite 85 Univ. of Florence 355 415 60 mid 86 Univ. of Basel 120 180 60 strong 87 KU Leuven 46 108 62 strong 88 Univ. of Bern 109 173 64 strong 89 Prince Sattam Bin Abdulaziz U. 440 375 65 lower 90 UIC 250 315 65 mid 91 UC Berkeley 9 75 66 elite 92 Univ. of Nottingham 146 215 69 strong 93 Univ. of Sussex 237 168 70 mid 94 King’s College London 38 108 70 strong 95 Ming Chi U. of Tech. 1053 1123 70 tail 96 Florida Intl. University 489 560 71 lower 97 Univ. of Toronto 21 93 72 elite 98 UC Santa Barbara 72 145 73 strong 99 Univ. of Nebraska-Lincoln 548 623 75 lower 100 Univ. of Tehran 490 565 75 lower 101 South China U. Tech. 280 205 75 mid 102 Univ. of Lausanne 125 200 75 mid 103 ETH Zurich 11 88 77 elite 104 Southeast University 292 215 77 mid 105 Zhejiang University 39 118 79 strong 106 Nanjing University 63 143 80 strong 107 Univ. of S. California 75 155 80 strong 108 Brown University 65 145 80 strong 109 Univ. of Jordan 639 720 81 lower 110 Carnegie Mellon Univ. 24 105 81 strong 111 Lincoln University 603 685 82 lower 112 Columbia University 20 103 83 elite 113 Univ. of Derby 793 710 83 tail 114 Jamia Millia Islamia 487 570 83 lower 115 Washington U. St. Louis 67 150 83 strong 116 Dalian U. of Tech. 404 320 84 mid 117 Duy Tan University 695 780 85 lower 118 Sichuan University 229 315 86 mid 119 Medical U. of Graz 248 335 87 mid 120 NYU 32 120 88 strong 121 Technion 343 255 88 mid 122 McMaster University 116 205 89 strong 123 Curtin University 285 195 90 mid 124 Tokyo U. Agri. & Tech. 1204 1295 91 tail 125 Beijing U. Chem. Tech. 451 360 91 lower 126 Univ. of Cyprus 493 585 92 lower 127 Radboud University 155 63 92 mid 128 Monash University 60 153 93 strong 129 Ewha Womans University 583 490 93 lower 130 Emory University 102 195 93 strong 131 Univ. of Massachusetts 112 205 93 strong 132 Univ. of Oxford 1 95 94 elite 133 Polytechnic U. Valencia 755 660 95 lower 134 Univ. of Otago 377 473 96 mid 135 Univ. Paris Cité 192 95 97 mid 136 Univ. of Chicago 15 113 98 elite 137 UC Merced 407 505 98 lower 138 Univ. College Cork 367 465 98 mid 139 Georgia Tech 42 140 98 strong 140 Oxford Brookes Univ. 819 920 101 lower 141 Univ. of Turin 468 368 101 mid 142 Tabriz U. Med. Sci. 626 728 102 lower 143 Shanghai University 508 405 103 lower 144 Harokopio U. Athens 696 800 104 lower 145 Univ. of British Columbia 45 150 105 strong 146 Konkuk University 559 665 106 lower 147 UC Riverside 329 223 107 mid 148 Univ. of Auckland 158 265 107 mid 149 Sun Yat-sen University 225 118 108 mid 150 Autonomous U. Madrid 384 493 109 mid 151 Dalhousie University 354 463 109 mid 152 Univ. of Virginia 171 280 109 mid 153 UCLA 18 128 110 strong 154 Canterbury Christ Church U. 1353 1240 113 tail 155 Rutgers U. New Brunswick 322 208 115 mid 156 Stellenbosch Univ. 326 443 117 mid 157 Case Western Reserve U. 148 265 117 mid 158 Glasgow Caledonian U. 913 795 118 tail 159 Portland State Univ. 985 868 118 tail 160 Sungkyunkwan Univ. 87 205 118 strong 161 Univ. of Edinburgh 29 148 119 strong 162 Stockholm University 204 85 119 mid 163 Univ. of Sheffield 111 230 119 strong 164 Coventry University 685 805 120 lower 165 Leiden University 70 190 120 strong 166 European U. Cyprus 854 975 121 tail 167 Macau U. Sci. & Tech. 260 383 123 mid 168 James Cook University 357 480 123 mid 169 Univ. of Nizwa 500 623 123 lower 170 Lund University 96 220 124 strong 171 Univ. of Granada 642 518 125 lower 172 Yazd University 1340 1465 125 tail 173 Univ. of Bath 263 390 127 mid 174 Univ. of Pavia 368 240 128 lower 175 Univ. of Florida 135 265 130 mid 176 Tech. U. Denmark 121 253 132 strong 177 Cardiff University 232 100 132 mid 178 Univ. of Leeds 118 250 132 strong 179 Aarhus University 101 238 137 strong 180 U. Teknologi Brunei 758 895 137 lower 181 Univ. of Sharjah 338 475 137 lower 182 Sunway University 307 445 138 mid 183 UC San Diego 47 185 138 strong 184 Yangzhou University 530 670 140 lower 185 Univ. of Houston 459 318 142 lower 186 Eindhoven U. Tech. 194 340 146 mid 187 Maastricht University 131 278 147 mid 188 NC State University 303 450 147 mid 189 Arizona State Univ. 213 360 147 mid 190 Reichman University 999 850 149 tail 191 Azerbaijan State Oil U. 1962 2115 153 tail 192 Keele University 563 410 153 lower 193 Princess Nourah U. 504 660 156 lower 194 Univ. of Salamanca 916 1073 157 lower 195 King Abdulaziz Univ. 399 243 157 lower 196 Swansea University 321 483 162 lower 197 Harbin Inst. Tech. 132 305 173 mid 198 Univ. of Würzburg 180 355 175 mid 199 Macquarie University 169 345 176 mid 200 Peking University 13 190 177 elite 201 Univ. of Connecticut 395 215 180 lower 202 TU Dresden 176 358 182 mid 203 Univ. of Warwick 123 305 182 strong 204 Ruhr U. Bochum 271 455 184 mid 205 Tech. U. Munich 27 213 186 strong 206 Lebanese American U. 268 455 187 lower 207 Univ. of S. Denmark 281 470 189 mid 208 Charité Berlin 91 280 189 strong 209 Univ. of Helsinki 104 295 191 strong 210 Claude Bernard U. Lyon 1 682 490 192 lower 211 Univ. of Tabuk 780 588 193 tail 212 Wageningen U.& Research 66 263 197 strong 213 Hyogo Medical Univ. 1303 1500 197 tail 214 Chiba University 1055 858 198 tail 215 Erasmus U. Rotterdam 107 305 198 strong 216 LUT University 301 500 199 mid 217 Univ. of Saskatchewan 397 600 203 lower 218 Nanjing U. Sci. & Tech. 622 418 205 lower 219 Graz U. of Tech. 633 840 207 lower 220 U. of Central Florida 478 685 207 lower 221 Univ. Sunshine Coast 578 785 207 tail 222 U. Tenaga Nasional 673 465 208 lower 223 Univ. of Genoa 410 200 210 lower 224 York University 422 635 213 lower 225 Univ. of Twente 191 405 214 mid 226 Duke University 28 243 215 strong 227 Univ. of Cologne 165 383 218 mid 228 UC Irvine 97 315 218 mid 229 Univ. of Bologna 130 350 220 mid 230 KTH 99 320 221 strong 231 Auburn University 644 870 226 lower 232 Yonsei University 86 320 234 strong 233 American U. Sharjah 560 795 235 lower 234 London South Bank U. 713 950 237 tail 235 Univ. of Valencia 543 303 241 lower 236 Fudan University 36 278 242 strong 237 Auckland U. Tech. 587 345 242 lower 238 Univ. of Tübingen 98 345 247 strong 239 Univ. of Johannesburg 388 635 247 lower 240 TU Dortmund 558 805 247 lower 241 Chinese U. Hong Kong 44 305 261 strong 242 Hebrew U. Jerusalem 261 525 264 mid 243 Prince Sultan Univ. 390 650 260 lower 244 Univ. of Haifa 607 875 268 lower 245 Ontario Tech Univ. 947 1215 268 tail 246 Izmir Inst. Tech. 1157 1425 268 tail 247 Lahore College Women U. 1417 1690 273 tail 248 Semmelweis University 273 555 282 lower 249 U. du Québec 541 828 287 lower 250 Sabancı University 393 685 292 mid 251 U. Sains Malaysia 412 705 293 mid 252 Univ. of Windsor 536 830 294 lower 253 Univ. of Iceland 521 223 299 lower 254 Yale University 10 310 300 elite 255 Univ. of Marburg 409 710 301 mid 256 Université Laval 413 715 302 lower 257 Univ. of Duisburg-Essen 312 620 308 mid 258 Univ. of Crete 649 340 309 tail 259 Natl. Yang Ming Chiao Tung U. 436 748 312 mid 260 Riphah Intl. Univ. 1038 1350 312 tail 261 Univ. Paris-Saclay 68 383 315 strong 262 Southern Cross Univ. 499 815 316 lower 263 U. Brunei Darussalam 374 690 316 mid 264 Univ. of Ilorin 1232 1550 318 tail 265 Dhofar University 677 995 318 tail 266 Univ. of Konstanz 283 603 320 mid 267 Aalto University 196 518 322 mid 268 Univ. of Kansas 375 698 323 mid 269 Univ. of Oregon 458 135 323 lower 270 Univ. of Sadat City 1145 1475 330 tail 271 SW Jiaotong University 836 505 331 lower 272 Indian Inst. Tech. Indore 573 905 332 lower 273 Univ. of Notre Dame 195 533 338 mid 274 Ohio State University 108 450 342 strong 275 Birla Inst. Tech. 1122 1465 343 tail 276 Univ. of Guelph 446 795 349 lower 277 Univ. of Tokyo 26 378 352 strong 278 Natl. Res. Nucl. U. MEPhI 702 1055 353 lower 279 George Washington U. 226 585 359 mid 280 Univ. of Reading 201 570 369 mid 281 Masaryk University 697 1068 371 lower 282 Univ. of Fribourg 423 800 377 lower 283 Arab Acad. Sci. Tech. 1020 1403 383 tail 284 Karunya Inst. Tech. 1205 1588 383 tail 285 Ajman University 411 795 384 lower 286 Univ. of Salford 851 1240 389 tail 287 Wenzhou University 872 480 392 tail 288 Kyoto University 61 455 394 strong 289 Kyushu University 319 720 401 mid 290 Hacettepe University 877 1285 408 tail 291 Kermanshah U. Med. Sci. 387 800 413 mid 292 Goldsmiths London 575 990 415 lower 293 Saarland University 522 950 428 lower 294 Univ. of Nicosia 514 945 431 lower 295 Zhejiang Chinese Med. U. 1451 1015 436 tail 296 San Diego State U. 1179 735 444 tail 297 UiT Arctic U. Norway 669 1115 446 lower 298 Swinburne U. Tech. 256 705 449 mid 299 Bilkent University 631 1083 452 lower 300 Univ. of Vienna 95 555 460 strong 301 Abu Dhabi University 233 695 462 mid 302 North South University 939 1405 466 tail 303 Natl. Taiwan Normal U. 570 1050 480 lower 304 Worcester Polytechnic Inst. 721 1205 484 tail 305 Victoria U. Wellington 434 925 491 mid 306 Zayed University 497 1000 503 lower 307 École Normale Sup. Lyon 315 830 515 mid 308 Univ. of Malakand 853 1368 515 tail 309 Purdue University 85 605 520 strong 310 Univ. of Osaka 152 680 528 strong 311 Federal U. Rio de Janeiro 723 1253 530 lower 312 Univ. of Luxembourg 265 813 548 lower 313 Edge Hill University 1092 1645 553 tail 314 Prince Mohammad Bin Fahd U. 352 905 553 mid 315 Tohoku University 105 660 555 strong 316 Univ. of Zagreb 1269 700 569 tail 317 Gazipur Agri. Univ. 935 1513 578 tail 318 Nanjing Tech Univ. 794 213 582 lower 319 Koç University 351 953 602 mid 320 La Trobe University 264 880 616 mid 321 Paderborn University 690 1308 618 lower 322 Indraprastha Inst. Info. Tech. 1169 1800 631 tail 323 Michigan State Univ. 106 758 652 mid 324 Tufts University 190 868 678 mid 325 Univ. of Klagenfurt 637 1305 668 lower 326 Chung Yuan Christian U. 1364 670 694 tail 327 Hangzhou Normal Univ. 1083 390 693 tail 328 Ain Shams University 904 1605 701 tail 329 Univ. of Colombo 1008 1760 752 tail 330 Indian Inst. Science 251 1005 754 mid 331 Univ. of Göttingen 122 895 773 strong 332 Kazan Federal Univ. 989 1790 801 tail 333 Mahatma Gandhi Univ. 567 1395 828 lower 334 George Mason Univ. 439 1320 881 lower 335 Istanbul Medipol Univ. 843 1733 890 tail 336 Centrale Nantes 621 1563 942 lower 337 Penn State 110 1095 985 strong 338 Univ. of Michigan 23 1005 982 strong 339 Univ. of Graz 585 1625 1040 lower 340 Univ. of Hagen 891 2000 1109 tail 341 ENTPE 635 1790 1155 lower 342 Lomonosov Moscow State U. 133 1400 1267 strong 343 Sciences Po 612 2120 1508 lower 344 MIPT 370 1903 1533 lower 345 Univ. of Milan 323 1880 1557 mid 346 Inst. Polytechnique de Paris 69 1700 1631 strong 347 Bauman Moscow State Tech. U. 340 2080 1740 mid 348 UT Austin 50 1813 1763 strong 349 Friedrich Schiller U. Jena 203 2070 1867 mid 350 Univ. of Toulouse 546 2425 1879 lower 351 Univ. of Cape Town 166 2200 2034 mid Figure 10: Error heatmap (tier × region). Cell values show mean AE. Darker cells indicate higher error. Cells with N<5N<5 are marked with an asterisk.