Paper deep dive
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
Deepak Kumar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/31/2026, 1:38:50 AM
Summary
SWE-PRBench is a benchmark for evaluating AI code review quality, consisting of 350 human-annotated pull requests. It tests 8 frontier models across three context configurations (diff-only, diff-with-file-content, and full-context). Results show that all models degrade monotonically as context increases, with a significant collapse in Type2_Contextual issue detection, suggesting that attention dilution is a primary constraint in AI code review.
Entities (5)
Relation Signals (3)
Deepak Kumar â authored â SWE-PRBench
confidence 100% · SWE-PRBench: Benchmarking AI Code Review Quality... Deepak Kumar
SWE-PRBench â contains â 350 pull requests
confidence 100% · We introduce SWE-PRBench, a benchmark of 350 pull requests
SWE-PRBench â hostedon â HuggingFace
confidence 100% · A public SWE-PRBench dataset on HuggingFace
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at kappa=0.75, 8 frontier models detect only 15-31% of human-flagged issues on the diff-only configuration, demonstrating that AI code review remains far below human expert performance despite strong results on code generation benchmarks. Pull requests are drawn from active open-source repositories, filtered from 700 candidates using a Repository Quality Score, and evaluated under three frozen context configurations: diff only (config_A), diff with file content (config_B), and full context (config_C), enabling systematic ablation of context provision strategies. All 8 models degrade monotonically from config_A to config_C, even when context is provided via structured semantic layers including AST-extracted function context and import graph resolution. The dominant mechanism is a collapse of Type2_Contextual issue detection at config_B, consistent with attention dilution in long contexts: a structured 2,000-token diff-with-summary prompt outperforms a 2,500-token full-context prompt enriched with execution context, behaviour mapping, and test signatures across all 8 models. The top four models are statistically indistinguishable (mean score 0.147-0.153) while a clear tier gap separates them from the remaining four (mean score <= 0.113). Dataset, contexts, annotations, and evaluation harness are released publicly.
Tags
Links
- Source: https://arxiv.org/abs/2603.26130v1
- Canonical: https://arxiv.org/abs/2603.26130v1
Trouble viewing inline? Open PDF directly â
Full Text
46,072 characters extracted from source content.
Expand or collapse full text
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback Deepak Kumar Independent Researcher deepak.kumar@foundryhq.ai (March 2026 â · arXiv preprint) Abstract We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at Îș=0.75Îș=0.75, 8 frontier models detect only 15â31% of human-flagged issues on the diff-only configuration, demonstrating that AI code review remains far below human expert performance despite strong results on code generation benchmarks. Pull requests are drawn from active open-source repositories, filtered from 700 candidates using a Repository Quality Score, and evaluated under three frozen context configurations : diff only (config_A), diff with file content (config_B), and full context (config_C) : enabling systematic ablation of context provision strategies. All 8 models degrade monotonically from config_A to config_C even when context is provided via structured semantic layers including AST-extracted function context and import graph resolution. The dominant mechanism is a collapse of Type2_Contextual issue detection at config_B, consistent with attention dilution in long contexts [18]: a structured 2,000-token diff-with-summary prompt outperforms a 2,500-token full-context prompt enriched with execution context, behaviour mapping, and test signatures across all 8 models. The top four models are statistically indistinguishable (sÂŻ s 0.147â0.153) while a clear tier gap separates them from the remaining four (sÂŻâ€0.113 s†0.113). Dataset, contexts, annotations, and evaluation harness are released publicly. Keywords. code review automation, large language models, software engineering benchmarks, pull request analysis, context-induced performance degradation, LLM-as-judge evaluation. 1 Introduction We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Existing benchmarks measure whether models can produce correct code [3, 4, 5]; none measure whether a model can evaluate proposed code changes as an expert reviewer would. We address this gap. Code review is a judgment task: the reviewer is presented with a diff, not a broken system; the goal is identifying and explaining problems in someone elseâs work, not generating a solution. Ground truth is distributed across human expert judgments rather than captured by pass/fail tests, and quality varies simultaneously across depth, actionability, and factual accuracy [1, 9]. A model can score highly on SWE-Bench while being unreliable as a code reviewer. Review bottlenecks are well-documented [2]; automating the evaluation task requires a benchmark that measures it directly. SWE-PRBench provides 350 pull requests selected from 700 candidates via a Repository Quality Score, evaluated under three frozen context configurations of increasing richness: diff only (config_A), diff with file content (config_B), and full context (config_C). Eight frontier models are evaluated using an LLM-as-judge framework validated at Îș=0.75Îș=0.75. Baseline results establish that no model detects more than 31% of human-flagged issues, and all 8 models degrade monotonically as context expands: a 2,000-token structured diff-with-summary prompt outperforms a 2,500-token full-context prompt across every model tested. We address the following research questions: (RQ1) What fraction of human-flagged issues can current LLMs detect, and how do models differ in hallucination rate? (RQ2) Does additional code context improve review quality, and in what direction does performance move? (RQ3) Is context degradation uniform across capability tiers and difficulty types? (RQ4) What categories of issues remain consistently beyond current model capability? This paper makes the following contributions. âą Dataset. 350 PRs with human-annotated ground truth across 6 languages and 3 difficulty types (Type1_Direct, Type2_Contextual, Type3_Latent). âą Context configurations. Three frozen, reproducible structures (config_A, config_B, config_C) enabling systematic context ablation. âą Evaluation protocol. A multi-dimensional scorer with LLM-as-judge classification and bipartite matching, validated at Îș=0.75Îș=0.75 and cross-validated under a second judge (Îș=0.616Îș=0.616). âą Finding. All 8 models degrade monotonically from config_A to config_C. The primary mechanism is attention dilution [18]: unstructured file content harms Type2_Contextual detection specifically, indicating that attention representation rather than content selection is the binding constraint for context-aware code review. âą Infrastructure. A public SWE-PRBench dataset on HuggingFace (https://huggingface.co/datasets/foundry-ai/swe-prbench), evaluation harness on GitHub (https://github.com/FoundryHQ-AI/swe-prbench), to support community benchmarking and extension. Table 1 summarises how SWE-PRBench differs from the closest existing code review datasets and evaluation frameworks. Table 1: Comparison of SWE-PRBench with related code review evaluation resources. Property CodeReviewer DeepCRCEval RovoDev SWE-PRBench (Ours) Primary contribution Model + dataset Eval metrics Production tool Dataset + protocol Ground truth source Synthetic pairs Generated CRR metric only Human reviewersâ Source links retained No No N/A Yes Difficulty taxonomy None None None 3 types Context configs None None None 3 frozen Issue detection eval No No Partial Yes Judge validated No No No Îș=0.75Îș=0.75 Public dataset Partial No No Yes â Ground truth in SWE-PRBench consists of review comments written by human engineers during the actual review process on merged pull requests, collected after the fact via GitHubâs review API. No comments are generated, synthesised, or modified during dataset construction. 2 Related Work 2.1 Code Generation and Repair Benchmarks SWE-Bench [3] tasks models with resolving GitHub issues by producing correct patches; top models score around 70% on the verified subset. SWE-Bench Pro [4] increases difficulty through GPL-licensed repositories and human-augmented specifications, with top models scoring approximately 23%. SWE Atlas [5] extends evaluation to deep codebase comprehension through agentic exploration. All three benchmarks measure code production. None measures code evaluation: the cognitive task of identifying problems in someone elseâs proposed work requires different capabilities [1] and is not predicted by generation benchmark scores. 2.2 Code Review Automation Automated code review has evolved from static analysis tools to pre-trained models. T5CR [6] and CodeReviewer [7] use T5-based architectures for review generation and code refinement; Tufano et al. [8] found that direct LLM use surpasses these on refinement but lags on review generation. LAURA [16] improves review generation by up to 38% IH-Score through retrieval-augmented context, demonstrating that context curation matters. DeepCRCEval [13] proposes LLM-based metrics to replace BLEU for evaluating generated review text. Our work is complementary: where DeepCRCEval evaluates the quality of generated review text given a fixed model, SWE-PRBench evaluates whether a model identifies the correct issues at all, a prerequisite question that precedes comment quality. Existing code review datasets. The dominant dataset resource is the CodeReviewer dataset [7], released alongside the CodeT5-based model. Prior model papers including Tufano et al. [6, 8] used training data not released as standalone evaluation benchmarks. All existing datasets have been shown to contain approximately 25â32% low-quality entries despite filtering [8]. Critically, no existing dataset retains PR context beyond the diff itself, source links are removed, and none is structured for evaluation of issue detection capability. SWE-PRBench addresses all three gaps: it retains full PR context, links to source repositories, and defines ground truth in terms of human-flagged issues rather than syntactic diff similarity. 2.3 LLM-as-Judge Methodology MT-Bench [10] established that GPT-4 as evaluator achieves over 80% agreement with human raters on open-ended tasks. G-Eval [11] extended this with chain-of-thought evaluation. JudgeBench [17] showed that many models perform near-randomly on challenging factual judgment tasks, motivating our choice of a high-capability fixed judge model validated against an explicit rubric (Îș=0.75Îș=0.75, Section 5.4) and cross-validated under a second judge (Îș=0.616Îș=0.616, Section 5.5). LAURAâs finding that curated retrieval-augmented context helps [16] is consistent with our result: context benefit depends on curation. Our config_B represents the uncurated baseline against which retrieval-based approaches should be measured. 3 The SWE-PRBench Dataset 3.1 Repository Selection and Quality Scoring Not all repositories have substantive review cultures. We define a Repository Quality Score (RQS) aggregating five components (Table 2) to select only repositories where human experts engage in technical discussion before merging. Repositories scoring below 60/100 are excluded. The median RQS of retained repositories is 74. Table 2: Repository Quality Score (RQS) components. Component What it measures Max pts Review culture Share of substantive human review comments 30 PR recency Merged PRs in last 90 days 25 Test quality Test files, CI presence, coverage tooling 20 PR volume Average monthly merged PRs over 6 months 15 Contamination Inverse of star count 10 Total 100 Review culture (max 30) is the primary signal: we sample the 30 most recent merged PRs within 90 days and score by the fraction of substantive human comments (human author, â„ 10 words, substance keywords or code references). PR recency ensures active maintenance. Test quality signals reliable CI infrastructure. PR volume distinguishes sustained from burst activity. Contamination penalises high-star repositories likely to appear in training corpora. 3.2 PR Collection and Filtering Ground truth for SWE-PRBench consists of review comments written by human engineers during the actual review process on merged pull requests. Comments are collected after the fact from GitHubâs review API; no comments are generated, synthesised, or modified during dataset construction. This distinguishes SWE-PRBench from prior code review datasets that derive ground truth from model-generated or heuristically-filtered comments [7, 6]. For each qualifying repository we collect merged pull requests using the GitHub GraphQL API [19] for PR metadata, reviews, and file-level information, and the GitHub REST API separately to fetch full unified diff patches (which GraphQL does not return). The collection window spans six months prior to the benchmark freeze date, with a minimum age of 30 days to ensure that review cycles have completed. We collect up to 30 PRs per repository, yielding an initial pool of approximately 3,000 raw pull requests. Raw pull requests are passed through a ten-stage filtering pipeline before entering the dataset. Stages filter: (1) merged-only status; (2) at least two substantive human comments; (3) at least one non-test file changed; (4) not documentation-only changes; (5) not automated dependency updates (Dependabot, Renovate); (6) not AI-dominated reviews, defined as more than 30% of comments originating from known AI review bots or matching AI-generated comment signatures; (7) diff parseable without API errors; (8) base commit accessible for file content retrieval; (9) repository still public at collection time; and (10) PR Review Value Score (RVS, Section 3.3) at or above 0.3. Bot and AI detection is explicit. The is_bot_username() function checks reviewer logins against a curated list of known AI review accounts (coderabbitai, github-copilot[bot], cursor-bugbot, qodo-merge-pro[bot], and 15 additional entries), and also matches accounts ending in [bot] or containing -bot. The is_ai_generated_comment() function detects AI-generated comment bodies via structural patterns: three or more Markdown H2 headers, three or more table lines, known AI review phrases, or length exceeding 2,000 characters with two or more headers. PRs where the AI comment ratio exceeds the 30% threshold are excluded from the dataset, ensuring that ground truth reflects genuine human engineering judgment rather than AI-assisted review [14]. After all filtering stages, the 700 candidate pull requests are scored with RQS and the 350 highest-quality PRs are retained as the final SWE-PRBench dataset. 3.3 PR Review Value Score Repository-level quality filtering via RQS is necessary but not sufficient: individual PR quality varies substantially even within high-RQS repositories. We define a PR Review Value Score (RVS) to quantify the ground-truth signal carried by each pull request and to weight its contribution to aggregate benchmark metrics. RVS is a weighted sum of five components, each normalised to [0,1][0,1]: RVS=0.25â ddepth+0.20â dcomplexity+0.20â ddiscussion+0.15â dtest+0.20â dbugRVS=0.25· d_depth+0.20· d_complexity+0.20· d_discussion+0.15· d_test+0.20· d_bug (1) where ddepthd_depth captures substantive comment count, thread count, and change-request events (normalised to 20); dcomplexityd_complexity reflects file count, lines changed, and cross-directory scope; ddiscussiond_discussion counts reply threads and unique reviewers; dtestd_test scores test modification signal (1.0 if tests are modified, 0.5 if added, 0 otherwise); and dbugd_bug scores bug-fix signal from title and label patterns. In the aggregate scoring formula, RVS contributes through a difficulty weight of logâĄ(human_comments+1) (human\_comments+1), ensuring that a PR with 25 expert-level review comments contributes more to aggregate metrics than a PR with three comments. 3.4 Difficulty Taxonomy We classify each pull request into one of three difficulty types based on where the evidence for a reviewable issue resides in the codebase [1]. Classification uses the is_in_diff field of human reviewer comments, cross-referenced against diff hunk line ranges parsed from the raw unified diff. Type1_Direct (34%). The issue is directly visible in the changed lines. A reviewer needs only the diff to identify it. Example: missing info < 0 error handling in a diff hunk that only checks info > 0. Type1 detection rate should be stable across all three context configurations. Type2_Contextual (40%). The issue requires understanding how changed code interacts with surrounding unchanged code in the same file. Example: a height-to-hash edge case where the bug is in the new logic, but understanding why it is wrong requires seeing the existing blockchain interface. Type2 detection rate should improve from config_A to config_B. Type3_Latent_Candidate (26%). The issue resides in files that import or depend on the changed files. Example: a Cython change that exposes a test teardown pattern issue in a file that uses the changed module. Type3 issues require the full context provided in config_C and, by design, test whether models can reason across file boundaries : a capability that current context windows only partially support. This taxonomy directly drives our ablation: if context provides signal rather than noise, Type2 detection should improve from config_A to config_B, and Type3 should only become accessible in config_C. Our results show the opposite: Type2 scores collapse at config_B across all models, establishing that the current context structure does not help contextual issues and actively harms performance on them. 3.5 Contamination Mitigation Model training corpora may include GitHub pull requests, creating a risk that benchmark PRs are partially memorised rather than genuinely reasoned over [3, 4]. We apply four complementary mitigations. First, all PRs are sourced from a six-month collection window ending at the benchmark freeze date, ensuring recency relative to the training cutoffs of evaluated models. Second, the RQS contamination component penalises high-star repositories, which are more likely to be crawled for pretraining. Third, we over-sample from GPL-licensed repositories following the approach of SWE-Bench Pro [4]. Fourth, and most importantly, Type3_Latent issues require cross-file reasoning that cannot be resolved by retrieving a memorised patch : a model must genuinely trace the dependency relationship at inference time. We embed each PR using a sentence-level encoder and compute cosine similarity against known benchmark repositories [3]; PRs exceeding a similarity threshold of 0.85 are excluded. We cannot fully eliminate contamination risk, and recommend that future evaluations gate on PRs merged after the training cutoff of each evaluated model. 3.6 Dataset Statistics Table 3 shows the construction funnel from raw collection to the final released dataset. Table 3: Dataset construction funnel. Stage Count Retention Raw PRs collected ⌠3,000 100% After 10-stage hard filter 700 23% After RVS â„ 0.35 350 12% Released (annotated) 350 â Table 4 summarises the final SWE-PRBench dataset. The 350 pull requests are sourced from 65 of the 100 repositories that passed the RQS threshold; the remaining 35 repositories contributed no PRs meeting the RVS â„0.35â„ 0.35 floor. Per-repository PR counts range from 1 to 30 (mean 5.4). The dataset is Python-dominant, reflecting the language distribution of high-RQS open-source repositories, with Python accounting for 242 PRs (69.1%), followed by JavaScript (37, 10.6%), Go (35, 10.0%), TypeScript (21, 6.0%), and Java (15, 4.3%). By PR type, feature PRs are the most common (214, 61.1%), followed by bug fixes (118, 33.7%), refactors (9), performance changes (8), and security fixes (1). The difficulty distribution reflects the classification method described in Section 3.4: 232 PRs (66.3%) are Type1_Direct, 75 (21.4%) are Type2_Contextual, and 43 (12.3%) are Type3_Latent_Candidate. The skew toward Type1 is expected : most review comments reference changed lines directly : and the 12.3% Type3 rate provides meaningful signal for testing cross-file reasoning capability. Table 5 shows the full cross-tabulation of difficulty by language. Pull request size ranges broadly. Lines added span 5 to 797 (mean 218.2, median 158), lines removed span 0 to 532 (mean 62.6, median 22), and files changed span 1 to 83 (mean 9.2, median 6). The upper bound of 797 added lines reflects the hard filter ceiling described in Section 3.2; the wide range ensures the benchmark covers both focused single-function patches and broader multi-file refactors. All PRs have an RVS score at or above the 0.35 threshold. The median RVS of the retained set is 0.52. All raw pull requests, rejection logs, RQS scores, RVS breakdowns, and difficulty classifications are released publicly alongside the benchmark to support reproducibility and dataset extension. Table 4: SWE-PRBench dataset statistics (350 PRs from 65 of 100 RQS-qualified repositories). Property Value Total PRs 350 Candidate pool (before RVS cut) 700 Repositories (RQS-qualified) 100 Repositories (contributing PRs) 65 PRs per contributing repo (min/mean/max) 1 / 5.4 / 30 Language distribution Python 242 (69.1%) JavaScript 37 (10.6%) Go 35 (10.0%) TypeScript 21 (6.0%) Java 15 (4.3%) PR type distribution Feature 214 (61.1%) Bug fix 118 (33.7%) Refactor 9 (2.6%) Performance 8 (2.3%) Security 1 (0.3%) Difficulty distribution Type1_Direct 232 (66.3%) Type2_Contextual 75 (21.4%) Type3_Latent 43 (12.3%) PR size Lines added (min/mean/median/max) 5 / 218.2 / 158 / 797 Lines removed (min/mean/median/max) 0 / 62.6 / 22 / 532 Files changed (min/mean/median/max) 1 / 9.2 / 6 / 83 RVS threshold â„0.35â„ 0.35 Median RVS (retained set) 0.52 Median RQS (retained repos) 74 Table 5: Difficulty type by language. Language Type1 Type2 Type3 Total Python 160 50 32 242 JavaScript 24 7 6 37 Go 24 8 3 35 TypeScript 13 7 1 21 Java 11 3 1 15 All 232 75 43 350 Each PR is stored as a JSON record containing metadata, scoring signals, reproducibility anchors (base and head commit SHAs), ground truth human review comments, and the raw unified diff patch. The full schema and a representative record are in Appendix A. 4 Context Configurations 4.1 The V1 Problem A naive collect-dump-truncate approach to context construction fails for three reasons: (1) token budgets are dominated by test files (60â70% on test-heavy PRs), leaving implementation files barely represented; (2) truncation at a fixed limit bisects diff hunks, producing syntactically incomplete fragments; (3) no explicit relationship between files is represented, so the agent cannot distinguish which unchanged file is related to which changed line. The V2 context builder addresses content selection through AST-based function extraction, import graph resolution, and behaviour mapping. As the results in Section 6 show, performance degrades monotonically even with V2, indicating that attention representation rather than content selection is the binding constraint for context-aware code review evaluation. 4.2 Three Fixed Configurations Three configurations are frozen at benchmark publication. Table 6 summarises them. config_A mirrors the information in a GitHub PR email notification. config_B mirrors the GitHub PR web view. config_C mirrors a reviewer working in a full IDE with the repository checked out. No configuration may be changed after results are reported, as they define the experimental conditions. Table 6: The three frozen context configurations. Budgets reflect actual measured token counts from the pipeline; the three configs differ in layer composition, not in token volume (range: 2,000â2,500 tokens). Config Layers Analogue Budget config_A Task focus (L0), Key changes summary (L1), Diff (L2), Minimal metadata (L6) GitHub PR email notification 2,000 tokens config_B + Execution context (L3), Behaviour mapping (L4) GitHub PR web view with file context 2,200 tokens config_C + Test signatures (L5) Reviewer with test suite access 2,500 tokens 4.3 V2 Improvements Over V1 Four targeted improvements in V2 address the V1 failure modes. âą Key changes summary (200 tokens). An LLM-generated summary of principal changes inserted before the diff. This was the highest-return-on-investment change, measurably reducing hallucination rate even in config_A. âą Behaviour mapping layer (100â150 tokens). For Go and TypeScript PRs with configuration-to-implementation propagation chains, V2 injects an explicit mapping of environment variable bindings and propagation paths, addressing the Type2/Type3 detection gap for non-Python languages. âą Hard no-truncation rule. If adding the next file would require mid-structure truncation, that file is dropped entirely and its budget allocation rolls over. Every included fragment is syntactically complete. âą Test noise reduction. Test files are included in body-stripped form: only def, assert, class, and @fixture lines are retained, targeting a ratio of 40% implementation diff, 30% execution context, 20% test signatures, and 10% metadata. Unlike V1âs collect-dump-truncate approach, V2 addresses the content selection problem. The persistence of the A>>B>>C pattern in V2 results therefore indicates that attention representation, not content selection, is the binding constraint. Full layer specification details are in Appendix B. 4.4 Reproducibility All pre-built contexts are released on HuggingFace as frozen artifacts at pipeline version v0.4.1 (https://huggingface.co/datasets/foundry-ai/swe-prbench). The was_truncated field is always false for V2 contexts by the no-truncation invariant. 5 The SWE-PRBench Evaluation Protocol The evaluation harness reads frozen context fixtures and makes independent agent and judge calls per record; the pipeline and all outputs are released at version v0.4.1. Agent and judge calls are fully decoupled: judge prompts can be updated and the full evaluation re-run without rebuilding contexts. 5.1 Agent Setup Each model receives a fixed system prompt (â 200 tokens) instructing it to identify only issues traceable to a specific line in the provided diff or context, to avoid inventing behaviour not visible in the shown code, and to return output as a valid JSON array with severity labels P0 (production bug), P1 (fix before merge), or P2 (credible risk), targeting 4â6 issues per PR. We evaluate 8 frontier models: Claude Haiku 4.5, Claude Sonnet 4.6, GPT-4o, GPT-4o-mini, DeepSeek V3, Mistral Large 3, Mistral Small, and Llama 3.3 70B (via Groq). Models are called via provider APIs at temperature 0 with independent per-record requests. If JSON extraction fails after one retry, the record is hard-zeroed (overall_score = 0); judge parse fallbacks halve the score (val *= 0.5). 5.2 Judge Design and Classification Rubric Agent comments are classified by a fixed judge model (GPT-5.2 in the final run; an earlier 20-PR validation used Claude Sonnet 4.6, with the same Îș=0.75Îș=0.75 result confirming judge consistency). The judge assigns each comment one of three labels: CONFIRMED (same underlying issue as a human comment, same file/area, same type of code change required); PLAUSIBLE (factually correct observation about visible code, but not in ground truth, since human reviews are not exhaustive [1]); or FABRICATED (references code not in context, or makes factually incorrect claims). PLAUSIBLE is not penalised as hallucination. Full criteria are published in RUBRIC.md alongside the dataset (Appendix C). A two-step consistency constraint is enforced: if agent comment X is CONFIRMED against human comment H, then H is marked CAUGHT. We enforce 1-to-1 bipartite matching using all-mpnet-base-v2 sentence embeddings to prevent duplicate credit; pairs below affinity 0.30 are discarded. The 0.30 threshold was calibrated against the rubric validation set; values below 0.30 produced no judge-CONFIRMED pairs that also passed blind human review. Matching and affinity formula details are in Appendix C. 5.3 Scoring Formula The overall score for a single pull request is: s=maxâĄ(0,minâĄ(1, 0.40â r+0.25â p+0.15â a+0.10â qâ+0.05â eâ0.25â hâ0.15â Ïâ0.10â Ï))s= \! (0,\, \! (1,\;0.40· r+0.25· p+0.15· a+0.10· q^*+0.05· e-0.25· h-0.15·Ï-0.10Â·Ï ) ) (2) where r = recall, p = precision, a = semantic alignment of matched pairs, qâq^* = actionability (discounted 80% when no comments are confirmed), e = efficiency (confirmed / total agent comments), h = hallucination rate (FABRICATED fraction), Ï = redundancy penalty, and ÏÏ = excess plausible penalty above threshold Ï=0.70Ï=0.70. The aggregate benchmark score is the difficulty-weighted mean: S=âisiâwi/âiwiS= _is_iw_i/ _iw_i where wi=logâĄ(human_commentsi+1)w_i= (human\_comments_i+1). 5.4 Rubric Validation (Îș=0.75Îș=0.75) We validate that the judge operationalises the published rubric by sampling 30 judge classifications (10 per category, stratified) and applying the rubric independently using a blind template that omits the judgeâs verdict. Cohenâs Kappa between judge labels and rubric-derived labels: Îș=0.75Îș=0.75 (substantial agreement, N=30N=30, exact agreement 83.3%). 5.5 Cross-Judge Validation To confirm the A>>B>>C ordering is not a judge artefact, we reran the judge step on a stratified 10-PR sample (batch_03, GPT-4o-mini agent outputs, N=127N=127 comments) using Claude Sonnet 4.6. Overall inter-judge agreement: 78.7% (Îș=0.616Îș=0.616). Both judges independently produce the same A>>B>>C ordering (Table 7). Table 7: Cross-judge agreement by context configuration (N=127N=127 comments, 10-PR sample, GPT-4o-mini agent). Both judges reproduce A>>B>>C. Config N Agreement GPT-5.2 proxy Sonnet proxy config_A 46 78.3% â-0.035 â-0.033 config_B 40 72.5% â-0.063 â-0.045 config_C 41 85.4% â-0.066 â-0.070 Overall 127 78.7% (Îș=0.616Îș=0.616) direction: A>>B>>C confirmed The lowest agreement is at config_B (72.5%), consistent with the Type2_Contextual collapse: agent comments generated with file content are more specific and therefore harder to classify unambiguously. The highest agreement is at config_C (85.4%), where increased hallucination makes FABRICATED classifications more clear-cut. 6 Baseline Evaluation on SWE-PRBench 6.1 Experimental Setup We evaluate on a 100-PR stratified sample: Type1_Direct (40 PRs), Type2_Contextual (40), and Type3_Latent (20). Type2 and Type3 are deliberately over-sampled relative to their 21% and 13% dataset prevalence to ensure statistical power for per-difficulty analysis; per-difficulty scores are therefore reported separately and the weighted aggregate sÂŻ s uses the actual 100-PR counts. 6.2 Leaderboard Results Table 8 establishes that SWE-PRBench is non-trivial: no model detects more than 31% of human-flagged issues on any configuration, and the mean detection rate across all 8 models on config_A is approximately 26%. These results quantify the gap between current AI code review capability and human expert performance. Figure 1 shows 95% bootstrap confidence intervals (N=300N=300, B=10,000B=10,000). The top four models (Haiku, Sonnet, DeepSeek, Mistral Large; sÂŻ s 0.147â0.153) are not statistically distinguishable, while the rank-4/rank-5 gap (0.147 vs. 0.113) is statistically significant. Table 8: SWE-PRBench leaderboard: 100-PR sample, 8 models, sorted by mean difficulty-weighted composite sÂŻ s across all configurations. Parse failures in footnotes. Best per column in bold. Model sAs_A sBs_B sCs_C sÂŻ s DRA FPRavg Claude Haiku 4.5⥠0.172 0.129 0.126 0.153 0.306 0.346 Claude Sonnet 4.6â 0.190 0.135 0.122 0.152 0.297 0.227 DeepSeek V3 0.181 0.132 0.118 0.150 0.312 0.315 Mistral Large 3 0.170 0.125 0.135 0.147 0.305 0.353 GPT-4o 0.134 0.110 0.090 0.113 0.220 0.193 GPT-4o-mini 0.110 0.095 0.093 0.108 0.210 0.353 Mistral Small 0.131 0.080 0.091 0.106 0.257 0.251 Llama 3.3 70B§ 0.088 0.071 0.065 0.079 0.223 0.417 â 1 parse failure (hard-zeroed). ⥠11 parse failures (hard-zeroed; removing them raises sÂŻ s by â 0.006, no rank change). § 30 judge parse fallbacks (score Ă0.5Ă 0.5). Figure 1: SWE-PRBench leaderboard with 95% bootstrap CIs. Top four models cluster (sÂŻ s 0.147â0.153); clear tier gap at rank 4/5. DeepSeek V3 achieves the highest raw detection on config_A (DR = 0.312) and outperforms all OpenAI models at approximately $0.28/M input tokens versus $2.50 for GPT-4o [21]: Tier 1 performance at roughly 9Ă lower cost. GPT-4o achieves the lowest hallucination rate (FPR = 0.193) but its detection rate lags all upper-tier competitors. 6.3 Context Configuration Behaviour Figure 2: Composite score by model and context configuration. All 8 models degrade Aâ ; the Aâ drop is the largest step. All 8 models degrade monotonically from config_A to config_C (Figure 2). Critically, the three configurations differ in layer composition, not in token volume: config_A and config_C differ by only 500 tokens (2,000 vs. 2,500). The A>>B>>C degradation therefore cannot be attributed to context length alone. Rather, it implicates the specific content added: execution context and behaviour mapping (config_B) and test signatures (config_C) each introduce additional signal that models cannot reliably integrate, even within a modest token budget. The V2 context builder uses AST-based function extraction and import graph resolution, ruling out content selection as the cause. The failure is in attention representation: once relevant context is provided as a flat token sequence alongside the diff, models cannot reliably direct attention to the changed lines. The Type2_Contextual collapse is the primary diagnostic: Sonnet drops from 0.22 to 0.10 on contextual issues when execution context is added at config_B, and DeepSeek drops from 0.20 to 0.10, consistent with the attention dilution documented by Liu et al. [18]. Table 9 shows the full difficulty-by-configuration breakdown. Figure 3: Difficulty Ă config heatmap for three models. The Type2 collapse at config_B is visible across all three. Table 9: Composite score by difficulty and config for three representative models. Collapse: >>50% drop from config_A to config_B. Model Difficulty A B C Trend Claude Sonnet 4.6 Type1 0.18 0.19 0.14 stable Type2 0.22 0.10 0.09 collapse Type3 0.17 0.13 0.14 stable DeepSeek V3 Type1 0.21 0.19 0.16 gradual Type2 0.20 0.10 0.10 collapse Type3 0.12 0.12 0.11 stable GPT-4o-mini Type1 0.15 0.13 0.15 stable Type2 0.09 0.08 0.06 gradual Type3 0.10 0.10 0.11 stable 6.4 Hallucination Profiles Figure 4: Mean hallucination rate (FPR) across config_A/B/C. GPT-4o is most precise (0.193); Llama fabricates at 0.417. Figure 4 shows a clear precision-recall tradeoff: Haiku and DeepSeek achieve higher detection but at higher hallucination cost; Sonnet and GPT-4o produce more reliable comments with lower FPR. The composite scorer correctly captures this: Sonnet ranks second despite lower raw DR than Haiku, because the hallucination and redundancy penalties offset Haikuâs recall advantage. Batch ordering is stable across all 10 batches, confirming the ranking is not a sample artefact. Practical deployment guidelines. (1) Use diff-only context (config_A): adding file content does not help and actively harms performance across all tested models. (2) Prioritise models by hallucination rate: GPT-4o (FPR = 0.193) and Sonnet (FPR = 0.227) produce more reliable comments than higher-recall alternatives, reducing triage burden. (3) DeepSeek V3 provides Tier 1 detection at roughly 9Ă lower API cost than GPT-4o with acceptable precision. 6.5 Benchmark Properties Table 10: Summary of benchmark properties and research question answers. RQ Question Answer Evidence 1 Detection rate and FPR 15â31% on config_A; FPR 0.193â0.417 Table 8 2 Context direction All models degrade A>>B>>C Figure 2 3 Uniformity Universal direction; Type2 drives the Aâ drop Table 9 4 Hard categories Type2 collapses at config_B; Type3 near-zero Figure 3 7 Dataset Analysis Table 5 shows the full difficulty-by-language distribution. Python has the highest absolute Type3 count (32 of 43, 74.4%), consistent with Pythonâs dynamic import patterns. Go and TypeScript have the highest proportional Type2 rates (22.9% and 33.3%), consistent with interface contracts and type propagation creating contextual issues. PR size spans 5â797 added lines (mean 218.2) and 1â83 files changed (mean 9.2), covering both focused patches and broad refactors. A SWE-PRBench task is hard in three structurally distinct ways: Type1 tasks require identifying a fault in changed lines without hallucinating non-existent bugs; Type2 tasks require understanding changed code in relation to surrounding context, but the current flat-token representation makes this harder with more context; Type3 tasks require cross-file dependency reasoning from import structure alone. The AI ceiling is clear: the best model (Haiku, DRA = 0.306) detects roughly one in three direct issues. If a second human reviewer achieves 50â70% recall [1], the gap to human performance is approximately 20â40 percentage points. 8 Limitations Scope of the context degradation finding. The V2 context builder uses AST extraction and import graph resolution, not raw file dumping. The degradation therefore implicates attention representation rather than content selection. Context provision strategies encoding relevance at the token level, such as explicit changed-vs-unchanged boundary markers or interleaved annotations, may produce different outcomes and are a direct priority for future work. Language and sample scope. The dataset is Python-dominant (69.1%) and evaluated on a 100-PR sample from the 350-PR full corpus. Results for non-Python languages and per-difficulty sub-group statistics should be interpreted with this in mind. Full 350-PR evaluation is deferred to the camera-ready version. Evaluation methodology. The primary judge is GPT-5.2. Cross-judge validation confirms the A>>B>>C direction under Sonnet 4.6 (Îș=0.616Îș=0.616), but absolute score values differ between judges. The potential for mild judge-family bias cannot be excluded without cross-family validation on the full leaderboard. No human performance baseline is available; the 20â40 percentage-point gap estimate to human recall is derived from prior literature [1]. Contamination. Despite recency filtering, RQS contamination weighting, and GPL preference, we cannot fully exclude the possibility that some benchmark PRs appeared in model pretraining. Type3 issues require cross-file reasoning that cannot be recalled from a memorised patch, providing partial contamination resistance for the hardest difficulty tier. 9 Conclusion We introduced SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. The dataset, pre-built context artifacts and evaluation harness are released publicly: âą Dataset: https://huggingface.co/datasets/foundry-ai/swe-prbench âą Evaluation harness: https://github.com/FoundryHQ-AI/swe-prbench Baseline results on 8 frontier models establish that no model detects more than 31% of human-flagged issues on any configuration, confirming that AI code review remains far from human expert performance. The context configuration ablation reveals that even structured semantic context â built via AST extraction, import graph resolution, and behaviour mapping â degrades performance relative to diff-only prompts. This implicates attention representation rather than content selection as the binding constraint: models cannot reliably distinguish changed from unchanged lines once both appear in a flat token sequence. We invite the community to submit results to the SWE-PRBench leaderboard, extend the dataset to additional languages and domains, and develop context representation strategies â such as explicit changed-vs-unchanged boundary markers or retrieval-augmented context â that close the gap between current AI performance and human expert review quality. Reproducibility Note. Scores reported in this paper reflect pipeline version v0.4.1 with GPT-5.2 as judge at temperature =0=0. Frontier model APIs do not guarantee full determinism at temperature =0=0, so minor score variation across independent runs is expected. The two-tier ranking structure and A>>B>>C ordering are stable across runs and confirmed by cross-judge validation (Îș=0.616Îș=0.616, Section 5.5). References [1] A. Bacchelli and C. Bird, âExpectations, outcomes, and challenges of modern code review,â in Proc. ICSE, 2013, p. 712â721. [2] J. Corbet, âKernel development: Who does it, how they do it, and what it means,â LWN.net, Jan. 2016. [Online]. Available: https://lwn.net/Articles/671120/ [3] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, âSWE-bench: Can language models resolve real-world GitHub issues?â in Proc. ICLR, 2024. [4] Scale AI, âSWE-Bench Pro: A harder benchmark for software engineering agents,â Technical Report, Scale AI, 2025. [5] Scale AI, âSWE Atlas: Codebase question answering for software engineering,â Technical Report, Scale AI, 2026. [6] M. Tufano, D. Drain, A. Svyatkovskiy, S. K. Deng, and N. Sundaresan, âUsing pre-trained models to boost code review automation,â in Proc. ICSE, 2022, p. 1â12. [7] Z. Li, S. Lu, S. Guo, N. Duan, S. Jannu, G. Jenks, D. Majumder, J. Green, A. Svyatkovskiy, S. Fu, and N. Sundaresan, âAutomating code review activities by large-scale pre-training,â in Proc. ESEC/FSE, 2022, p. 1035â1047. [8] R. Tufano, S. Gualandi, M. Aghajani, M. Monperrus, and A. Bacchelli, âCode review automation: Strengths and weaknesses of the state of the art,â IEEE Transactions on Software Engineering, vol. 50, no. 5, 2024. [9] O. Kononenko, O. Baysal, and M. W. Godfrey, âCode review quality: How developers see it,â in Proc. ICSE, 2016, p. 1028â1038. [10] L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, âJudging LLM-as-a-judge with MT-Bench and chatbot arena,â in Proc. NeurIPS, 2023. [11] Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, âG-Eval: NLG evaluation using GPT-4 with better human alignment,â in Proc. EMNLP, 2023, p. 2511â2522. [12] J. Cohen, âA coefficient of agreement for nominal scales,â Educational and Psychological Measurement, vol. 20, no. 1, p. 37â46, 1960. [13] J. Lu, D. Li, B. Yu, and L. Zhang, âDeepCRCEval: Revisiting the evaluation of code review comment generation,â arXiv preprint arXiv:2412.18291, 2025. [14] M. Golzadeh, A. Decan, D. Legay, and T. Mens, âA ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments,â Journal of Systems and Software, vol. 175, 2021. [15] A. Bosu, M. Greiler, and C. Bird, âCharacteristics of useful code reviews: An empirical study at Microsoft,â in Proc. MSR, 2015, p. 146â156. [16] Y. Zhang, Y. Zhang, Z. Sun, Y. Jiang, and H. Liu, âLAURA: Enhancing code review generation with context-enriched retrieval-augmented LLM,â arXiv preprint arXiv:2512.01356, 2024. [17] R. Tan, J. Fu, T. Pang, L. Du, M. Lin, and C. Gao, âJudgeBench: A benchmark for evaluating LLM-based judges,â arXiv preprint arXiv:2410.12784, 2024. [18] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, âLost in the middle: How language models use long contexts,â Transactions of the Association for Computational Linguistics, vol. 12, p. 157â173, 2024. [19] GitHub, âGitHub GraphQL API,â 2024. [Online]. Available: https://docs.github.com/en/graphql [20] GitHub, âGitHub REST API: Comparing commits,â 2024. [Online]. Available: https://docs.github.com/en/rest/commits/commits [21] OpenAI, âOpenAI API pricing,â 2024. [Online]. Available: https://openai.com/api/pricing/ Appendix A PR Record Schema Each SWE-PRBench pull request is stored as a JSON record with five field categories. Metadata: task_id, repo, pr_number, language, pr_type, difficulty, merged_at. Scoring signals: rvs_score and five component scores (review_depth, code_complexity, discussion_signal, test_change_signal, bug_fix_signal), plus lines_added, lines_removed, files_changed. Reproducibility anchors: base_commit and head_commit SHA hashes pinning the exact file state at review time. Ground truth: human_review_comments array with per-comment fields for author, body, path, line, diff hunk, and reply status; plus num_substantive_comments, num_unique_reviewers, has_requested_changes, ai_comments_removed. Diff: diff_patch as the raw unified diff retrieved from the GitHub REST API; agent_input is populated by the context builder at evaluation time. A representative record (dask/dask PR #12221) is available in the HuggingFace dataset card. This PR carries RVS = 0.334, placing it below the 0.35 inclusion threshold; it is provided as a boundary illustration of the quality cut, not as a dataset member. Appendix B Context Builder Layer Specification The V2 context builder assembles each configuration through a layered pipeline with budget rollover (unused tokens flow to subsequent layers). Layer 1a (metadata, 200 tokens fixed): PR title, repository, difficulty type, language, file count, truncated description; never includes RVS score, ground truth, or post-merge data. Layer 1b (key changes summary, 200 tokens): LLM-generated natural-language summary of principal changes inserted before the diff. Layer 2 (diff, 800â1,200 tokens): per-file hunks sorted by importance =0.5Ălines_changed+0.3Ăcomment_density+0.2Ăfunction_changes=0.5Ălines\_changed+0.3Ăcomment\_density+0.2Ăfunction\_changes; production files before test files; at most 4 files. Layer 3 (changed files, 400â800 tokens): AST extraction for Python (get_innermost_enclosing_nodes plus relevant imports), cluster-based line windows for other languages; always uses base_commit, never head_commit; excludes test files. Layer 4a (related files, 100â300 tokens): import graph resolution for Python; GitHub code search for others; empty with not_applicable_type1 flag for Type1 PRs. Layer 4b (test files, 150â300 tokens): body-stripped, retaining def, async def, class, assert, and @fixture lines only; Go/TypeScript PRs also receive a behaviour mapping layer (100â150 tokens) making config-to-implementation propagation chains explicit. Appendix C Evaluation Protocol Detail CONFIRMED criteria (all must hold): (1) the agent identifies a specific issue in the code; (2) a human reviewer identified the same underlying issue; (3) the issue concerns the same file or functional area; (4) the agent concern would lead to the same kind of code change as the human concern. PLAUSIBLE: grounded in visible code, factually correct, reasonable engineering concern, but not matched to any human comment. FABRICATED (any sufficient): references code not in context; makes factually incorrect claims; describes a non-existent bug; invents method signatures or variable names. Matching detail. Similarity uses all-mpnet-base-v2 sentence embeddings. Stage 1: judge-CONFIRMED pairs are sorted by combined actionability-and-similarity signal; the top-scoring 1-to-1 assignment is selected greedily. Stage 2: unmatched non-FABRICATED comments are evaluated against remaining human comments via pairwise affinity =cosineâ(ea,eh)+0.15â â[same file]+ÎŽline=cosine(e_a,e_h)+0.15·1[same file]+ _line, clamped to [0,1][0,1]. Pairs below 0.30 are discarded. The threshold was calibrated against the 30-sample rubric validation set; values below 0.30 produced no judge-CONFIRMED pairs that also passed blind human review. Agent comments unmatched after both stages and not FABRICATED are reclassified to PLAUSIBLE.