Paper deep dive
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
Sihong Wu, Yiling Ma, Yilun Zhao, Tiansheng Hu, Owen Jiang, Manasi Patwardhan, Arman Cohan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 1:05:02 AM
Summary
RbtAct is a framework that improves the actionability of AI-generated peer review feedback by using author rebuttals as implicit supervision. The authors introduce RMR-75K, a large-scale dataset mapping review segments to rebuttal responses, and employ a training pipeline involving supervised fine-tuning followed by Direct Preference Optimization (DPO) to align model outputs with feedback that historically leads to concrete author revisions.
Entities (5)
Relation Signals (3)
RbtAct â optimizes â Llama-3.1-8B-Instruct
confidence 95% · We then train the Llama-3.1-8B-Instruct model with supervised fine-tuning... followed by preference optimization
RbtAct â utilizes â RMR-75K
confidence 95% · We propose RBTACT... We also build a large dataset named RMR-75K
RMR-75K â contains â Review-Rebuttal Mapping
confidence 90% · RMR-75K (Review-Map-Rebuttal) which maps review and rebuttal at the segment level
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are increasingly used across the scientific workflow, including to draft peer-review reports. However, many AI-generated reviews are superficial and insufficiently actionable, leaving authors without concrete, implementable guidance and motivating the gap this work addresses. We propose RbtAct, which targets actionable review feedback generation and places existing peer review rebuttal at the center of learning. Rebuttals show which reviewer comments led to concrete revisions or specific plans, and which were only defended. Building on this insight, we leverage rebuttal as implicit supervision to directly optimize a feedback generator for actionability. To support this objective, we propose a new task called perspective-conditioned segment-level review feedback generation, in which the model is required to produce a single focused comment based on the complete paper and a specified perspective such as experiments and writing. We also build a large dataset named RMR-75K that maps review segments to the rebuttal segments that address them, with perspective labels and impact categories that order author uptake. We then train the Llama-3.1-8B-Instruct model with supervised fine-tuning on review segments followed by preference optimization using rebuttal derived pairs. Experiments with human experts and LLM-as-a-judge show consistent gains in actionability and specificity over strong baselines while maintaining grounding and relevance.
Tags
Links
- Source: https://arxiv.org/abs/2603.09723v1
- Canonical: https://arxiv.org/abs/2603.09723v1
Trouble viewing inline? Open PDF directly â
Full Text
81,901 characters extracted from source content.
Expand or collapse full text
RBTACT: Rebuttal as Supervision for Actionable Review Feedback Generation Sihong Wu Y Yiling Ma Y Yilun Zhao Y * Tiansheng Hu N Owen Jiang Y Manasi Patwardhan T Arman Cohan Y Y Yale University N New York University T TCS Research Abstract Large language models (LLMs) are increasingly used across the scientific workflow, including to draft peer-review reports. However, many AI-generated reviews are superficial and insufficiently actionable, leaving authors without concrete, implementable guidance and motivating the gap this work addresses. We propose RBTACT, which targets actionable review feedback generation and places existing peer review rebuttal at the center of learning. Rebuttals show which reviewer comments led to concrete revisions or specific plans, and which were only defended. Building on this insight, we leverage rebuttal as implicit supervision to directly optimize a feedback generator for actionability. To support this objective, we propose a new task called perspective-conditioned segment-level review feedback generation, in which the model is required to produce a single focused comment based on the complete paper and a specified perspective such as experiments and writing. We also build a large dataset named RMR-75K that maps review segments to the rebuttal segments that address them, with perspective labels and impact categories that order author uptake. We then train the Llama-3.1-8B-Instruct model with supervised fine-tuning on review segments followed by preference optimization using rebuttal derived pairs. Experiments with human experts and LLM-as-a-judge show consistent gains in actionability and specificity over strong baselines while maintaining grounding and relevance. Data: RMR-75KCode: RbtAct 1 Introduction LLMs are increasingly used in scientific research, including assisting with scientific writing and peer reviews [Zhao et al., 2025, Zhu et al., 2025a]. Early work explored whether LLMs can help draft peer reviews or support reviewers through prompting methods [Robertson, 2023, Hosseini and Horbach, 2023, Xu et al., 2025]. Subsequent research moved from prompting to fine-tuning methods [Gao et al., 2024, Weng et al., 2025] with multi-agent coordination [Jin et al., 2024, Tan et al., 2024, Zhu et al., 2025b]. Prior studies show that while LLMs can draft fluent reviews, they often miss specific issues, show shallow analysis, and produce generic phrasing, so their feedback does not reliably act as actionable guidance [Liu and Shah, 2023, Shin et al., 2025]. At the same time, the peer review process itself contains a rich supervision signal that scrutinizes the scientific merits of a work and allows researchers to improve their research. This is often done through author rebuttals in the peer review process, where authors either commit to concrete changes or defer action on certain reviewer comments [Kennard et al., 2022]. We argue that rebuttals are an underutilized source of implicit human feedback for learning what kinds of comments actually trigger revision. 1 In this work, we propose RBTACT, utilizing rebuttals as an implicit preference signal and using them to directly optimize review generation for actionability. Concretely, we derive pairwise preferences from rebuttal outcomes and apply preference optimization so that the model favors comments that elicited author action [Rafailov et al., 2024]. This turns rebuttal from an object of analysis [Zhang et al., 2025] into a supervision source for training a model. Typically, full conference reviews mix strengths, weaknesses, and questions across multiple aspects. Authors mainly respond to weaknesses and questions, which vary across perspectives such as experiments, novelty, writ- ing [Ghosal et al., 2022]. Treating a full review as one unit makes actionability difficult to evaluate because author reactions target only parts of the review. We therefore decompose reviews into key-point segments and study segment-level review feedback generation for a given perspective: given the full paper and a target perspective (e.g., * Correspondence to: Yilun Zhao (yilun.zhao@yale.edu) 1 Reviews and rebuttals can be noisy or inconsistent, but we hypothesize that large-scale use provides sufficient signal to improve model training. 1 arXiv:2603.09723v1 [cs.CL] 10 Mar 2026 Raw D a t a Collection OpenReview.net Strong X We a k ICLR PDF API Gateway Full Manuscript Reviewer's Review Data collected from OpenReview - ICLR2024 Quality Control Structural filter Substance filter Official Comment Comment Thank you for your review. The quick brown fox jumps o v e r the lazy dog. Lorem ipsum dolor sit a m e t , ... 0 . 6 I am n o t able to parse Drop reports unable t o split into atomic part Official Review We a k n e s s 1.Not clear 2.Language needs improvment 3.Looks l i k e a homework Coverage filter Confidence filter â A v Drop review segment w i t h no corresponding rebuttal span The best matching p a i r is ... Confidence score: 0.1 D r o p review- rebuttal pair with low confidence We clarify that... Author's Rebuttal Lightweight classifier flags segments with vague requests H u m a n verification Official Review Summary Authors describe a benchmark for... Strengths 1 . Well-written and We l l - o rg a n i z e We a k n e s s 1 . U n c l e a r o n the d e c i s i o n criteria 2.No report o n g r o u p differences... Questions 1.What is t h e cutoff? 2.What is the metric?... Official Comment Comment T h a n k you for your review. W h a t is the cutoff? We set the cutoff as..... What is the metric? We use... Data Construction Review Segmentation âą Target actionable content from review Atomic Question Atomic Rebuttal âą M e t h o d 1: Split comment by original enumeration M e t h o d 2: LLM segment rebuttal into atomic items Review-Rebuttal Mapping A= M e t h o d 1: What is the cutoff? U t i l i z i n g explicit anchor We set the cutoff as.... A = A= M e t h o d 2: LLM m a t c h e r with e n f o r c d one-to-one alignment The best matching p a i r is .. Confidence: ... Figure 1: Overview of our construction pipeline for RMR-75K. Experiments), the model produces one focused comment. This design narrows scope, promotes specificity, and enables precise supervision by aligning each review segment with the rebuttal segment that addresses it. Prior work connects reviews to downstream author behavior in different ways. ARIES [DâArcy et al., 2024a] links review comments to paper edits, enabling edit prediction from feedback but not aligning comments to rebuttal text. DISAPERE [Kennard et al., 2022] annotates discourse relations between review and rebuttal at the sentence level, but at a smaller scale and without perspective labels. Our setting complements both: we segment reviews into key points, align each point to the corresponding rebuttal span, and attach a perspective label and an impact category that captures the authorâs reaction, such as a concrete revision performed, a planned revision, or a defense without changes. These impact categories make actionability concrete and induce pairwise preferences in which comments leading to revisions outrank those that yield plans or defenses. To formulate preference pairs, for each paper and perspective, we collect all review segments aligned to rebuttal spans; whenever two segments share the same paper and perspective but have different impact categories, we form a pair. Compared with previous datasets, our resource targets segment-level review feedback generation and provides rebuttal-anchored supervision that reflects what authors actually did or committed to do. We first fine-tune Llama-3.1-8B-Instruct on perspective-conditioned review segments to establish a strong baseline. We then apply preference-optimization through rebuttal-derived DPO [Rafailov et al., 2024] to optimize the model for actionability. This pipeline treats rebuttal as a natural reward model to serve as a signal for actionable review generation, targeting the gap identified by prior evaluations that LLM-generated reviews are often generic and not sufficiently tied to revision [Hosseini and Horbach, 2023, Chamoun et al., 2024, Sadallah et al., 2025]. Experiments show that our model produces review feedback that is more actionable and specific than competitive baselines under both human and LLM-as-a-judge evaluations. We evaluate on a test set built from ICLR 2025 and the improvements hold across multiple perspectives. Comparisons cover our RBTACT, an SFT-only variant, and larger prompted LLMs such as Llama-3.1-70B and GPT-5-chat. RBTACT achieves the highest actionability in both studies: 3.46 out of 5 in human evaluation and 3.38 out of 5 in LLM-as-a-judge evaluation, maintaining parity on groundedness and relevance, with fine-grained analyses across seven perspectives and pairwise win rates indicating consistent gains. We summarize our contributions as follows: âąWe present the framework RBTACT that first utilizes author rebuttals as an implicit supervision and applies preference optimization to directly optimize a feedback generator for actionability. âą We release a large-scale dataset named RMR-75K, which contains 75,542 examples. Each example consists of (i) a review segment, (i) an associated perspective on the review, (i) an author response to the review, and (iv) an annotated impact category indicating the reviewâs actionability. âąWe propose an effective training pipeline that yields consistent gains in actionability and other aspects over competitive baselines under human and LLM-as-a-judge evaluations. 2 2 Related Work Peer Review Datasets. Early corpora such as PeerRead and NLPeer collect manuscripts, decisions, and textual reviews to support large-scale analysis of peer review [Kang et al., 2018, Dycke et al., 2023]. More recent large- scale resources support fine-tuning reviewer models, including ReviewMT, Review-5K and DeepReview-13K [Tan et al., 2024, Zhu et al., 2025b, Weng et al., 2025]. Some other resources incorporate rebuttal to study review: PRRCA links reviews with rebuttal counter-arguments for meta-review generation [Wu et al., 2022]; MOPRD aggregates multi-disciplinary open reviews with rebuttal letters and editorial signals [Lin et al., 2023]; Re 2 scales multi-turn review and rebuttal discussions across many venues [Zhang et al., 2025]. To analyze reviewârebuttal links, DISAPERE adds sentence-level discourse annotations over 506 review and rebuttal pairs [Kennard et al., 2022]; JitsuPeer further aligns reviewârebuttal sentences with rebuttal action types for attitude-grounded rebuttal generation [Purkayastha et al., 2023]. However, these links are defined at the sentence level and remain small-scale. In contrast, we construct large-scale point-to-point mappings that align each review point to its corresponding rebuttal segment and add perspective and impact labels, using them to supervise review feedback generation. Review Generation.Early efforts prompted LLMs to write full reviews, revealing limited but non-trivial useful- ness and notable failure modes [Yuan et al., 2021, Liu and Shah, 2023, Robertson, 2023]. Later systems improved specificity through fine-tuning and structured pipelines, and most recently multi-agent approaches that assign roles to reduce generic feedback and increase comment quality [Jin et al., 2024, Zhu et al., 2025b, Zou et al., 2026]. Another related thread focuses on actionable reviewing, concerning whether feedback induces concrete changes or commitments. Most prior approaches operationalize actionability using reviewer-centric signals. ReAct labels reviews for actionability and intent to support detection and triage [Choudhary et al., 2021]; ARIES links review comments to implemented edits, enabling analysis of which feedback leads to revisions [DâArcy et al., 2024a]; MARG explicitly frames the task as generating actionable peer review with specialized agents and reports reduced generic comments [DâArcy et al., 2024b]. In contrast, our method leverages rebuttal-anchored signals to ground actionability in authorsâ responses and actions, facilitating actionable review feedback generation. 3 Review and Rebuttal Mapping To train models that generate actionable review comments using rebuttal as supervision, we need fine-grained mappings that connect each review point to the specific rebuttal that addresses it. Such segment-level alignment lets us attach a perspective label and an impact signal indicating whether the comment led to a concrete revision, a specific plan, or a defense without changes. Existing resources are not sufficient: DISAPERE [Kennard et al., 2022] provides sentence-level discourse annotations but lacks our segment-level labels and is much smaller (506 pairs); ARIES [DâArcy et al., 2024a] is related but targets authorsâ edits; other datasets either omit rebuttal or do not treat it as a training signal. Therefore we introduce our dataset RMR-75K (Review-Map-Rebuttal) which maps review and rebuttal at the segment level, and is much bigger (about 150Ă more than DISAPERE). 3.1 ReviewâRebuttal Source Data Collection We curate a corpus with papers, reviews and rebuttals from ICLR 2024 on OpenReview by querying the official API for each submissionâs manuscript, full reviews, and the corresponding author responses. For downstream text processing, we convert the PDF manuscripts to Markdown with MinerU [Wang et al., 2024]. For each paper we keep (i) the full manuscript text; (i) every reviewerâs complete review and free-form comments; (i) the author rebuttal thread. When a rebuttal appears in multiple messages, we concatenate them in chronological order to obtain a single response document per paper. 3.2 Dataset Construction We construct a point-to-point mapping between reviewer key points and the specific portion of the rebuttal that addresses each point, named RMR-75K, shown in Figure 1. Dataset statistics are summarized in Table 1. Notation. Let paperphave reviewer setJ p . A reviewerjâJ p writes a reviewR p j , which we decompose into a sequence of weakness/question segmentsR p j,1:K j =r p j,1 ,...,r p j,K j . The concatenated rebuttal text forpisB p , which we further split into candidate response spans B p 1:T =b p 1 ,...,b p T . The goal is to construct a mapping A(p) = (r p j,k , b p t , Ëc p j,k,t ) (j,k,t)âI p , whereËc p j,k,t â [0, 1]is a confidence score and we enforce a one-to-one mapping: eachr p j,k and eachb p t appears in at most one pair. 3 StatisticValue Total mappings75,542 Total papers4,825 Avg. reviewers per paper3.44 Avg. mappings per paper15.66 Distinct reviews16,583 Avg. mappings per review4.56 Avg. confidence score0.9268 ConferenceICLR 2024 Table 1: Summary statistics for the Review-Rebuttal-Mapping-75k dataset. Review Segmentation.We target only actionable content by extracting Weaknesses and Questions because author rebuttals mainly respond to these two parts. When reviewers already enumerate items (e.g., â1/2/3â, âW1/W2/. . . â or bullet lists), we take those spans directly. Otherwise, we prompt GPT-5 [OpenAI, 2025] to splitR p j into atomic critique unitsr p j,k that each express a single concern. The prompt is shown in Figure 4. ReviewâRebuttal Mapping. We perform two-stage alignment. First, a heuristic pass links items using explicit anchors (e.g., âW1â, quoted phrases, or reviewerâs specific references), yielding high precision pairs. Second, an LLM matcher conducts span-level semantic linking: given a segmentr p j,k and the candidate spansb p t , it selects the best-supportedb p t and returns a calibrated confidenceËc p j,k,t . We enforce one-to-one alignment by greedy matching in descending Ëc and discard ties. This realizes a paper-specific map Map : R p , B p â (r p j,k , b p t ) (j,k,t) . The prompt is shown in Figure 5. 3.3 Cleaning and Quality Control Because our training uses rebuttal-derived impact categories as preference signals for actionability, supervision quality hinges on precise reviewârebuttal alignment. We therefore keep only confidently aligned pairs via the filters below and verify quality with targeted human checks. Structural Filter. We drop reviews with no visible itemization and for which the segmenter cannot produce stable atomic units. Coverage Filter. During alignment, if a review segmentr p j,k receives no plausible rebuttal span (noËcâ„ Ï), we remove it. Confidence Filter. We set a thresholdÏand keep only pairs withËcâ„ Ï; we also prune pairs where the matched rebuttal span merely restates the question without addressing it. Substance Filter. We filter out review segments that do not semantically raise any substantive issue or recommen- dation. Human Verification. We sample 60 papers to cover different perspectives. From these papers, we include all review segments that passed our filters and met the confidence threshold, yielding 944 mapped segments. Two trained annotators independently map each review segment to the most specific rebuttal span (or mark âNo Responseâ), following the same one-to-one constraint and span granularity as our pipeline. A mapping is counted as correct if the predicted rebuttal span overlaps the gold span with token-level IoUâ„ 0.5. We compute span-level Precision/Recall/F1 for our automatic mapping against the adjudicated gold, and report CohenâsÎșbetween the two annotators before adjudication. Results are in Table 2. Our mapping achieves high span-level accuracy (F1 = 0.91) with substantial IAA (Îș = 0.80), indicating that the pipeline provides reliable supervision for downstream training. SettingPrecisionRecallF1Cohenâs Îș Filtered set0.930.900.910.80 Table 2: Gold-standard validation of reviewârebuttal span mappings. Îș is inter-annotator agreement (IAA). 4 Methodology 4.1 Task Definition Given a paperp, a perspectivesand the whole paper text as context, the model generates a single focused review segment y that raises weaknesses or questions of paper p from perspective s. 4 Label setDefinition Perspective of Review Segment ExperimentsSetup, baselines, ablations, datasets EvaluationMetrics, analysis, claims vs. results ReproducibilityMissing code/details, reproducibility info NoveltyOriginality, relation to prior work TheoryAssumptions, derivations, proofs WritingClarity, grammar, readability PresentationFigures, tables, organization Impact Category of Rebuttal Segment CRPConcrete Revision Performed SRPSpecific Revision Plan VCRVague Commitment to Revise DWCDefend Without Change DRFDeflect or reframe, no change Evaluation Experiments Novelty Presentation Reproducibility Theory Writing 0 20 40 60 80 100 Percentage (%) CRPSRPVCRDWC DRF Figure 2: Left: brief summary of perspective labels for review segments and impact categories for rebuttal segments. Right: normalized (100%) impact category composition by perspective. Review Perspective Labels. Each mapped review segment receives one label from a 7-category taxonomy: Ex- periments, Writing, Presentation, Theory, Novelty, Reproducibility, and Evaluation. We assign labels automatically with a rubric-based prompt to GPT-5 that asks for the single best label and a short rationale, where prompt is shown in Appendix B. To verify quality, two trained annotators conducted stratified spot checks over papers and confidence bins. They independently reviewed a held-out sample and then resolved disagreements. The automatic labels matched human judgment on most cases (accuracy about 92%; CohenâsÎș = 0.81). The remaining confusions were mainly Writing vs. Presentation and Experiments vs. Evaluation. Moreover, the detailed definition of perspective labels is shown in Table 7. Rebuttal Impact Category as Actionability Signal. For each mapped rebuttal segment, we annotate an impact category that reflects the authorâs concrete action in response to the review by GPT-5: (1) Concrete Revision Performed (CRP), (2) Specific revision plan (SRP), (3) Vague Commitment to Revise (VCR), (4) Defend Without Change (DWC), and (5) Deflect/Reframe (DRF). The definition is shown in Figure 2, and more detailed definition is shown in Table 7. The automatic labels matched the adjudicated human labels on most cases (accuracy 89%), with inter-annotator agreementÎș = 0.79. They correspond to varying degrees of modifications made by the authors, reflecting the reviewâs level of actionability. The categories leverage author reactions in rebuttal segments, reflecting how much revision or defense occurred, to capture the actionability of reviewer comments and their impact on rebuttals. After labeling, the distribution of each impact category and perspective in RMR-75K is shown in Figure 2. We detail the specific label process and the distribution of each category in Appendix B. 4.2 Training Dataset Construction SFT Data. We construct REVIEWSEG-SFT-13K, a supervised corpus collecting pairsD SFT = (x i ,y i ) N i=1 . Each input isx i = (p,s), wherepis the paper content andsis the target perspective (e.g., Experiments, Writing); the targety â i is the gold review sentence selected from the mapped reviewer segmentsR p j,1:K j . Dataset size and coverage are reported in Table 3. There are13,300pairs spanning4,637papers, balanced to include1,900instances per perspective. Preference Data. From the reviewârebuttal alignmentsA(p)in §3.2, we build REVIEWPREF-DPO-22K as preference triples(x,y w ,y â ), wherex = (p,s)matches the SFT input andy w ,y â are two review segments drawn from the same paperpand perspectives. Winners follow the rebuttal impact order CRP Âż SRP Âż VCR Âż DWC Âż DRF, which reflects increasing author uptake and revision degree in rebuttals. We compare only strictly ordered labels and stratify by impact gap to modulate difficulty: large (CRP vs. DWC, DRF), medium (SRP vs. DWC, DRF, CRP vs. VCR), and small (CRP vs. SRP, VCR vs. DWC, DRF, DWC vs. DRF). To keep the signal robust, we balance pair counts across papers and perspectives, never mix different papers or perspectives within a pair, and cap how often any single segment can appear so no segment dominates the pool. As summarized in Table 3, the preference set contains 21,822 pairs from 4,825 papers with 3,117 pairs per perspective in average. Moreover, we quantify the difficulty stratification in Table 4. This distribution exposes the model to a spectrum of preference margins while keeping pairs within the same paper and perspective. 5 StatisticSFT-13KDPO-22K # pairs13,30021,822 # papers4,6374,825 Per-perspective1,900 eachavg. 3,117 each Avg. paper length22,15221,798 Avg. output length62Ch: 65, Re: 63 ConferenceICLR 2024ICLR 2024 Extra test pairs (â10%)13302180 Table 3: Statistics for SFT dataset and preference dataset for DPO. The unit of length is tokens. âChâ denotes Chosen and âReâ denotes Rejected. TierWinnerLoser#PairsPercentages Large impact gap (Easy) CRPDWC9,88745.3 CRPDRF3811.7 Subtotal10,26847.1 Medium impact gap SRPDWC4,96922.8 SRPDRF2481.1 CRPVCR1,9048.7 Subtotal7,12132.6 Small impact gap (Hard) CRPSRP2,48411.4 VCRDWC1,4236.5 VCRDRF700.3 DWCDRF4562.1 Subtotal6,43320.3 Total21,822100.0 Table 4: Difficulty-stratified breakdown of preference pairs in REVIEWPREF-DPO-22K. Pairs are constructed only within the same paper and perspective and use strictly ordered rebuttal-impact labels. 4.3 Policy Optimization Direct Preference Optimization.Direct Preference Optimization (DPO) learns a policy from pairwise preferences without fitting a separate reward model [Rafailov et al., 2024]. LetÏ Îž be the trainable policy andÏ ref a fixed reference policy (the SFT model). Under a BradleyâTerry preference model, DPO maximizes the probability that the policy prefersy w overy â by matching the optimal policyâs log-density ratio relative toÏ ref . The loss for a batch of preference triples is as follows: Letâ Ξ,ref (x,y) = logÏ Îž (y|x)â logÏ ref (y|x). We sample preference triples (x,y w ,y â ) from the datasetD pref . The DPO loss is L DPO (Ξ) =âE (x,y w ,y â ) h logÏ ÎČ [ â Ξ,ref (x,y w )â â Ξ,ref (x,y â ) ] i (1) whereÎČ> 0controls the sharpness of the preference. Intuitively, the policy is encouraged to increase likelihood on comments that led to higher-impact author actions (CRP, SRP) while decreasing it on comments that were defended or deflected (DWC, DRF). Stabilization. We keepÏ ref frozen at the SFT checkpoint and optionally mix in a small fraction ofL SFT on positive samples to prevent drift on perspective control when the context is long. The full objective is L(Ξ) = L DPO (Ξ) + λE (x,y w ) [â logÏ Îž (y w |x)], with a small λ. In practice, we set λ as 0.1. 4.4 Training Details All experiments run on NVIDIA H200 141GB GPUs with bf16 compute. We use the base model Llama-3.1-8B- Instruct [Meta AI, 2024]. We train RBTACT-SFT model on dataset REVIEWSEG-SFT-13K. We run3epochs (4,989steps) with learning rate1.0Ă10 â4 and a cosine scheduler. Total training time on H200 is approximately120 hours. To obtain RBTACT, we then utilize dataset REVIEWPREF-DPO-22K for further DPO training. The DPO policyÏ Îž is initialized from the SFT checkpoint and trained for2epochs (2,728steps) using the BradleyâTerry 6 DPO loss, keeping the referenceÏ ref frozen. We use learning rate1.0Ă10 â5 and a cosine scheduler. Total DPO training time on H200 is approximately203hours. The additional training details including training configs and other optimizations are shown in Appendix C. 5 Experiments 5.1 Baselines We compare RBTACT against three baseline types under identical inputs, prompts and decoding. Fine-tuned (SFT-only). As a trained baseline, we fine-tune Llama-3.1-8B-Instruct on ReviewSeg-SFT-13k, getting RBTACT-SFT, using the same prompt template and output format. Prompted LLMs.We query API models and run competitive open-source models in zero-shot mode to produce a single review segment for the requested perspective: GPT-5-chat [OpenAI, 2025], DeepSeek-V3.2 [DeepSeek-AI, 2025], Llama-3.1-70B [Meta AI, 2024] and Qwen-3-32B [Team, 2025]. The prompt is shown in Appendix D.4. Other Methods.We additionally compare against three task-adapted methods that are widely used around review feedback generation: (i) MARG [DâArcy et al., 2024b], a multi-agent prompting framework for scientific review generation; (i) LimGen [Faizullah et al., 2024], a limitations-focused generation setup that targets suggestive weaknesses; and (i) DeepReviewer-14B [Zhu et al., 2025b], a review generation model to produce comprehensive reviews. For each method, we adapt its protocol to our setting (paper + requested perspectiveâa single review feedback) and normalize the final output to our unified segment format. The details are shown in D.2. 5.2 Evaluation Protocol Evaluation Dataset.We construct a test set from a subset of ICLR 2025 papers using the same pipeline as in §3.2 to reduce data contamination, ensuring that none are resubmissions of ICLR 2024 papers. We sample 700 papers, stratified by perspective with 100 papers per perspective across the seven perspectives. Each paper is paired with one perspective labeled by annotators and one human review segment, which serves as the gold reference. Human Evaluation Setup.We evaluate a subset of the evaluation dataset: 50 papers sampled across perspectives. For each paper, annotators view the title, relevant content, and the target perspective. Three PhD-level or senior graduate annotators (each withâ„2 completed reviews at major ML venues) rate nine anonymized model outputs per paper on a 1â5 scale for Actionability, Specificity, Groundedness, Relevance, and Helpfulness, following a written rubric with positive and negative indicators and anchor examples. Outputs are shown as three order-randomized pairs to mitigate position bias. Scores are averaged over annotators. Full instructions are reported in Appendix D.3. LLM-as-a-Judge Evaluation Setup. We evaluate on 105 papers from the evaluation dataset. Each model produces one segment per paper, yielding 105 segments per model and 1350 total across nine models. The judge model (GPT-5-chat) sees the paper context, the target perspective, and one or two candidate segments [Zheng et al., 2023]. For pointwise scoring, it assigns 1â5 on the same five dimensions as the human study. For pairwise comparison, it scores both candidates and then chooses which is more Actionable with a brief rationale. The exact prompts are in Figure 15 and Figure 16. HumanLLM-as-a-Judge SystemAction. Spec. Ground. Rel. Help.Action. Spec. Ground. Rel. Help. Ours RBTACT (ours)3.464.084.304.76 4.263.383.704.054.82 3.74 RBTACT-SFT3.284.014.164.70 4.243.183.593.944.72 3.66 LLMs GPT-5-chat3.384.044.354.98 4.473.283.664.124.95 3.78 DeepSeek-V3.23.153.984.224.88 4.283.133.564.004.86 3.70 Llama-3.1-70B3.223.954.184.65 4.153.113.543.964.74 3.53 Qwen-3-32B3.063.784.124.58 4.123.033.363.904.68 3.32 Other Methods MARG3.203.874.154.72 4.183.193.473.914.69 3.51 DeepReviewer-14B3.273.964.284.75 4.213.233.484.034.74 3.58 LimGen3.143.924.084.64 4.053.083.383.884.54 3.39 Table 5: Pointwise ratings on five quality dimensions. Left: Human. Right: LLM-as-a-judge. Higher is better. 7 5.3 Experiment Results Human Evaluation Results. As shown in Table 5, RBTACT attains the highest Actionability while remaining competitive on other dimensions. LLM-as-a-judge Pointwise Results.The judge reproduces the human trend: RBTACT scores highest on Action- ability and Specificity, while remaining close to strong LLM baselines on other dimensions (Table 5). The prompt is shown in Figure 15. LLM-as-a-judge Pairwise Results. We compare Actionability via pairwise judgments on the same paper and perspective. Table 6 reports the percentage of wins for each model. To visualize patterns across perspectives, Figure 17 shows heatmaps where each cell is the win rate of the row model over the column model. Overall, RBTACT attains the highest average pairwise win rate and leads in most perspectives, followed by GPT-5-chat. The pairwise prompt is in Figure 16. Actionability (Pairwise Win Rate %) WinnersRBTACT GPT-5 DeepSeek Llama-3.1 DeepReviewer MARG Qwen-3 RBTACT-SFT LimGen RBTACTâ57.163.861.965.768.666.765.776.2 GPT-5-chat42.9â55.257.160.062.960.054.371.4 DeepSeek-V3.236.244.8â54.357.159.058.153.368.6 Llama-3.1-70B38.142.945.7â52.455.256.251.466.7 DeepReviewer-14b34.340.042.947.6â51.453.350.564.8 MARG31.437.141.044.848.6â50.549.561.9 Qwen-3-32B33.340.041.943.846.749.5â52.458.1 RBTACT-SFT34.345.746.748.647.650.547.6â56.2 LimGen23.828.631.433.335.238.141.943.8â Table 6: LLM-as-a-judge pairwise win rates (%) on Actionability HumanâLLM Agreement. We quantify alignment between human judgments and the LLM judge along three axes. (i) Model-level rank correlation: ranking models by pointwise means yields high agreement for Actionability (SpearmanâsÏ=0.94, KendallâsÏ b =0.87), with only a minor swap between mid-ranked baselines. (i) Item-level correlation: across matched paperâmodel cells (20 papersĂ9 models), human vs. LLM pointwise scores show moderate positive correlation on Actionability (Spearmanâs Ï=0.52). 5.4 Automatic Evaluation We evaluate all 700 instances from evaluation dataset with four metrics: BLEU@4 [Papineni et al., 2002], ROUGE- L sum (F1) [Lin, 2004], METEOR [Banerjee and Lavie, 2005], and chrF [Popovi Ì c, 2015]. Details are shown in Appendix D.5. Our models are competitive with or stronger than the baselines on most overlap-based metrics. In particular, RBTACT achieves the best ROUGE-L sum and METEOR, while RBTACT-SFT performs similarly and attains the highest BLEU@4. 5.5 Case Study Here we show some case studies to demonstrate why our model RBTACT is more actionable than other baselines in most cases from different perspectives. They are shown in Appendix E. 5.6 Summary Both human and LLM-as-a-judge evaluations show that our RBTACT improves the practical value of review segments. Automatic evaluation also shows the strong capability of RBTACT. Gains concentrate on actionability and specificity with parity on groundedness, relevance, and Helpfulness. Importantly, our 8B-scale RBTACT remains competitive with much larger 32â70B and proprietary models on actionability (and often specificity), suggesting the practical value of rebuttal-supervised training. 6 Conclusion We study actionable review feedback generation by placing rebuttal at the center of learning. Our framework, RBTACT, uses rebuttals as implicit supervision and frames the task as segment-level generation from a given perspective with explicit mappings from each review segment to its addressing rebuttal span. We release RMR- 75K with perspective labels and impact categories that utilize actionability. An effective pipeline of supervised fine-tuning followed by preference optimization on impact ordered pairs yields consistent gains in actionability and specificity while maintaining grounding and relevance against strong baselines under both human evaluation and LLM-as-a-judge protocols. These results show rebuttal signals are a practical form of human feedback for 8 producing targeted, implementable guidance. We release data and code to spur research on rebuttal driven learning and better evaluation of actionability. Limitations Our approach relies on rebuttal as an implicit supervision signal for actionability, which is informative but imperfect. Rebuttals reflect short horizon author uptake during review, not long-term implementation, and can include strategic promises or deferrals. The dataset focuses on venues with public rebuttals, mainly computer science communities that use OpenReview, so generalization to journals, non-English venues, and other fields remains uncertain. Our model can generate precise yet infeasible suggestions and our current setup does not include rigorous verification against the manuscript, code, or data artifacts. References Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu, Ronan Le Bras, Yixin Liu, Xiangru Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao, Hannaneh Hajishirzi, Doug Downey, and Arman Cohan. Sciarena: An open evaluation platform for non-verifiable scientific literature-grounded tasks. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=am6R85mnc. Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. DeepReview: Improving LLM-based paper review with human-like deep thinking process. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mo- hammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 29330â29355, Vienna, Austria, July 2025a. Associa- tion for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.1420. URL https://aclanthology.org/2025.acl-long.1420/. Zachary Robertson. Gpt4 is slightly helpful for peer-review assistance: A pilot study, 2023. URLhttps: //arxiv.org/abs/2307.05492. M. Hosseini and S. P. J. M. Horbach. Fighting reviewer fatigue or amplifying bias? considerations and recom- mendations for use of chatgpt and other large language models in scholarly peer review. Research Integrity and Peer Review, 8(4):1â12, 2023. doi: 10.1186/s41073-023-00133-5. URLhttps://doi.org/10.1186/ s41073-023-00133-5. Zhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Can LLMs identify critical limitations within scientific research? a systematic evaluation on AI research papers. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20652â20706, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025. acl-long.1009. URL https://aclanthology.org/2025.acl-long.1009/. Zhaolin Gao, Kiant Ì e Brantley, and Thorsten Joachims. Reviewer2: Optimizing review generation through prompt generation, 2024. URL https://arxiv.org/abs/2402.10886. Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. Cycleresearcher: Improving automated research via automated review, 2025. URLhttps://arxiv.org/ abs/2411.00816. Yiqiao Jin, Qinlin Zhao, Yiyang Wang, Hao Chen, Kaijie Zhu, Yijia Xiao, and Jindong Wang. AgentReview: Exploring peer review dynamics with LLM agents. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1208â1226, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/ v1/2024.emnlp-main.70. URL https://aclanthology.org/2024.emnlp-main.70/. Cheng Tan, Dongxin Lyu, Siyuan Li, Zhangyang Gao, Jingxuan Wei, Siqi Ma, Zicheng Liu, and Stan Z. Li. Peer review as a multi-turn and long-context dialogue with role-based interactions, 2024. URLhttps://arxiv. org/abs/2406.05688. Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. Deepreview: Improving llm-based paper review with human-like deep thinking process, 2025b. URL https://arxiv.org/abs/2503.08569. Ryan Liu and Nihar B. Shah. Reviewergpt? an exploratory study on using large language models for paper reviewing, 2023. URL https://arxiv.org/abs/2306.00622. 9 Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, and Juho Kim. Mind the blind spots: A focus-level evaluation framework for llm reviews, 2025. URL https://arxiv.org/abs/2502.17086. Neha Nayak Kennard, Tim OâGorman, Rajarshi Das, Akshay Sharma, Chhandak Bagchi, Matthew Clinton, Pranay Kumar Yelugam, Hamed Zamani, and Andrew McCallum. DISAPERE: A dataset for discourse structure in peer review discussions. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1234â1249, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.89. URLhttps://aclanthology.org/ 2022.naacl-main.89/. Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2024. URLhttps://arxiv.org/ abs/2305.18290. Daoze Zhang, Zhijian Bao, Sihang Du, Zhiyi Zhao, Kuangling Zhang, Dezheng Bao, and Yang Yang. Re 2 : A consistency-ensured dataset for full-stage peer review and multi-turn rebuttal discussions, 2025. URLhttps: //arxiv.org/abs/2505.07920. Tirthankar Ghosal, Shubham Kumar, Prasenjit Kumar Bharti, and Asif Ekbal. Peer review analyze: A novel benchmark resource for computational analysis of peer reviews. PLOS ONE, 17(1):e0259238, 2022. doi: 10.1371/journal.pone.0259238. Mike DâArcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, and Doug Downey. Aries: A corpus of scientific paper edits made in response to peer reviews, 2024a. URLhttps://arxiv.org/abs/ 2306.12587. Eric Chamoun, Michael Schlichtkrull, and Andreas Vlachos. Automated focused feedback generation for scientific writing assistance. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 9742â9763, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.580. URLhttps://aclanthology.org/ 2024.findings-acl.580/. Abdelrahman Sadallah, Tim Baumg Ì artner, Iryna Gurevych, and Ted Briscoe. The good, the bad and the constructive: Automatically measuring peer reviewâs utility for authors, 2025. URLhttps://arxiv.org/abs/2509. 04484. Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. A dataset of peer reviews (PeerRead): Collection, insights and NLP applications. In Marilyn Walker, Heng Ji, and Amanda Stent, editors, Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1647â1661, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1149. URL https://aclanthology.org/N18-1149/. Nils Dycke, Ilia Kuznetsov, and Iryna Gurevych. NLPeer: A unified resource for the computational study of peer review. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5049â5073, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.277. URL https://aclanthology.org/2023.acl-long.277/. Po-Cheng Wu, An-Zi Yen, Hen-Hsen Huang, and Hsin-Hsi Chen. Incorporating peer reviews and rebuttal counter- arguments for meta-review generation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, CIKM â22, page 2189â2198, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450392365. doi: 10.1145/3511808.3557360. URLhttps://doi.org/10.1145/ 3511808.3557360. Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi. Moprd: A multidisciplinary open peer review dataset. Neural Computing and Applications, 35(34):24191â24206, September 2023. ISSN 1433-3058. doi: 10.1007/s00521-023-08891-5. URL http://dx.doi.org/10.1007/s00521-023-08891-5. Sukannya Purkayastha, Anne Lauscher, and Iryna Gurevych. Exploring jiu-jitsu argumentation for writing peer review rebuttals, 2023. URL https://arxiv.org/abs/2311.03998. Weizhe Yuan, Pengfei Liu, and Graham Neubig. Can we automate scientific reviewing?, 2021. URLhttps: //arxiv.org/abs/2102.00176. 10 Zhuoyang Zou, Abolfazl Ansari, Delvin Ce Zhang, Dongwon Lee, and Wenpeng Yin. Diagpaper: Diagnosing valid and specific weaknesses in scientific papers via multi-agent reasoning, 2026. URLhttps://arxiv.org/ abs/2601.07611. Gautam Choudhary, Natwar Modani, and Nitish Maurya. ReAct: A Review Comment Dataset for Actionability (and more), page 336â343. Springer International Publishing, 2021. ISBN 9783030915605. doi: 10.1007/ 978-3-030-91560-524. URL http://dx.doi.org/10.1007/978-3-030-91560-5_24. Mike DâArcy, Tom Hope, Larry Birnbaum, and Doug Downey. Marg: Multi-agent review generation for scientific papers, 2024b. URL https://arxiv.org/abs/2401.04259. Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, Fukai Shang, Bo Zhang, Liqun Wei, Zhihao Sui, Wei Li, Botian Shi, Yu Qiao, Dahua Lin, and Conghui He. Mineru: An open-source solution for precise document content extraction, 2024. URLhttps://arxiv.org/abs/ 2409.18839. OpenAI. Gpt-5. https://openai.com/gpt-5/, 2025. Accessed: 2025-10-07. Meta AI. Introducing llama 3.1: Our most capable models to date.https://ai.meta.com/blog/ meta-llama-3-1/, 2024. Accessed: 2025-10-07. DeepSeek-AI.Introducing deepseek-v3.2-exp.https://api-docs.deepseek.com/news/ news250929, 2025. Accessed: 2025-10-07. Qwen Team. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. URLhttps://arxiv.org/ abs/2505.09388. Abdur Rahman Bin Md Faizullah, Ashok Urlana, and Rahul Mishra. Limgen: Probing the llms for generating suggestive limitations of research papers, 2024. URL https://arxiv.org/abs/2403.15529. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311â318, Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics. doi: 10.3115/1073083.1073135. URLhttps: //aclanthology.org/P02-1040/. Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74â81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URLhttps: //aclanthology.org/W04-1013/. Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss, editors, Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65â72, Ann Arbor, Michigan, June 2005. Association for Computational Linguistics. URLhttps: //aclanthology.org/W05-0909/. Maja Popovi Ì c. chrF: character n-gram F-score for automatic MT evaluation. In Ond Ë rej Bojar, Rajan Chatterjee, Christian Federmann, Barry Haddow, Chris Hokamp, Matthias Huck, Varvara Logacheva, and Pavel Pecina, editors, Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392â395, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/W15-3049. URLhttps: //aclanthology.org/W15-3049/. 11 "review_point": "id": "P1", "content": "Tool making is an interesting idea. However, one weakness in this direction is that we could have many external tools in real-world scenarios and the capability of tool-making should be able to explore new tools that do not have before. Therefore, to further improve the feasibility of the proposed method, it will be better to explore the capability of tool-making to make new tools when also considering the existence of external tools.", "rebuttal_response": "id": "R1", "content": "Regarding weakness 1, we agree exploring the tool maker's ability to create entirely new tools beyond existing APIs would be an exciting direction for future work. Our initial study focused on a simpler closed-loop setting, but we plan to investigate more open-ended tool making scenarios as well.", "confidence_score": 0.95 Reviewer Author Rebuttal Review-Rebuttal -Mapping-75k Dataset (2 samples) "review_point": "id": "P3", "content": "This paper mainly utilized the GPT-based models for evaluation, do you try any other open-source LLMs?", "rebuttal_response": "id": "R3", "content": "Q1) We have not yet evaluated our approach on other open-source LLMs. Testing on other open-source LLMs could illuminate generalizability. However, as our results show, even GPT-4 sometimes fails, so advanced LLMs may be needed for robust tool making. But expanding to other model families is a helpful suggestion for future work.", "confidence_score": 0.95 Figure 3: A mapping example of our Review-Rebuttal-Mapping-75k dataset. The review and rebuttal are from the paper titled âLarge Language Models as Tool Makersâ. A Review and Rebuttal Mapping Details A.1 Example of Mapping Here is an example in our Review-Rebuttal-Mapping-75k dataset in Figure 3. We sample two mappings in the list. A.2 Prompts for Review Segment and Mapping The prompt for review segmentation is shown in Figure 4. The prompt for mapping between review segment and rebuttal segment is shown in Figure 5. B Label Details We discuss details of our labeling of perspectives and impact categories here. B.1 Detailed Definition of Perspective labels and Impact Categories The detailed definition of perspective labels and rebuttalâs impact categories are shown in Table 7. B.2 Perspective and Impact Category Distribution We analyze the seven-perspectives and impact category distribution in our Review-Rebuttal-Mapping-75k dataset, which is shown in Table 8. B.3 Prompt for labeling perspectives The prompt for labeling which perspective a review segment belongs to is shown in Figure 6. B.4 Prompt for labeling impact categories The prompt for labeling which impact category a rebuttal segment belongs to is shown in Figure 7. C Additional Training Details All training details are shown in Table 9. 12 [System Prompt] You are a professional academic review text analysis assistant. Your task is to segment a complete âWeaknesses & Questionsâ section from an academic paper review into independent, specific points. You must follow these rules: 1. Each point should be an independent, specific issue or weakness 2. Preserve the core meaning of the original text without adding or removing information 3. Maintain existing numbering structures if present (e.g., 1., 2., W1, W2, etc.) 4. Handle various formatting styles including: - Numbered lists (1., 2., 3.) - Letter prefixes (W1, W2, Q1, Q2) - Markdown bullet points (-, *, +) - Section headers (## Weaknesses, ## Questions) 5. If no clear numbering exists, logically segment based on content structure 6. Each point should contain sufficient context to be understood independently 7. Preserve the original language and terminology used by the reviewer [User Prompt] Please segment the following Weaknesses & Questions text into independent points: weaknesses&questions text IMPORTANT: Regardless of the input format (bullet points, numbered lists, paragraphs, etc.), you MUST output in this exact format: Point 1: [Complete content of the first weakness point] Point 2: [Complete content of the second weakness point] Point 3: [Complete content of the third weakness point] ... Rules: - Use exactly âPoint X:â where X is a number starting from 1 - Include ALL weakness points from the input, donât skip any - Each point should be complete and independently understandable - Preserve the original meaning and wording as much as possible - If the input has bullet points (-, *, +) or numbered lists (1., 2.), convert them to Point format - If the input has long paragraphs, break them into logical points Figure 4: Prompt used to segment the weaknesses and questions parts of the review into segments. Framework and hardware. We train usingLLaMA-Factoryon NVIDIA H200 GPUs (141 GB) withbf16 compute. We apply LoRA adapters to attention and MLP projections (q,k,v,o,gate,up,down) on top of the Llama-3.1-8B-Instruct base model. All runs use per-device batch size1with gradient accumulation as specified below. Memory and speed optimizations.Direct Preference Optimization (DPO) requires evaluating both the chosen and the rejected responses under the policy and a frozen reference model, which increases memory use compared with SFT. To enable a32k token context in DPO, we combine: (i) FlashAttention-2 for efficient attention; (i) DeepSpeed ZeRO-2 with gradient checkpointing; and (i) a 4-bit quantized frozen reference model. This configuration reduces activation and parameter memory, allowing full paper context while maintaining throughput. Schedules.SFT trains for3epochs (= 4,989steps) with learning rate1.0Ă10 â4 and a cosine scheduler (warmup ratio0.10). DPO initializes from the SFT checkpoint and trains for 2 epochs (= 2,728steps) with learning rate 1.0Ă10 â5 and a cosine scheduler (warmup ratio0.05). Wall-clock time on H200 isâ 120hours for SFT andâ 203 hours for DPO. D Evaluation Details D.1 Baseline Prompt Figure 8 shows the prompt for different LLM baselines generating review segments. 13 Label setDefinition Perspective of Review Segment ExperimentsConcerns about experimental setup and design: missing or insufficient experiments, weak or absent ablations, unfair/weak baselines, unclear dataset descriptions or splits, problematic hyperparameter/seed choices, compute budgets, or training details that undermine empirical support. EvaluationHow results are measured, analyzed, and interpreted: inappropriate/missing metrics, lack of statistical testing or error bars, insufficient analysis (e.g., no variance/sensitivity), unfair comparisons, or inconsisten- cies between reported numbers and claims. ReproducibilityAbility to reproduce results: missing implementation details or pseudo-code, absent code/data/links, unspecified hyperparameters, unclear preprocessing, seeds, or hardware; insufficient instructions to replicate tables/figures. NoveltyOriginality and relation to prior work: overlap with existing methods, incremental contributions, unclear differentiation, missing or superficial positioning against closely related literature. TheoryTheoretical correctness and justification: flawed or unstated assumptions, gaps in proofs, incorrect derivations, mismatches between theorems and algorithms, or theory not supporting the claimed guarantees. WritingClarity and readability: grammar/style issues, ambiguous phrasing, undefined symbols/terms, confusing explanations, or organization that impedes understanding at the sentence/paragraph level. PresentationFigures/tables/organization: unclear plots or legends, poor formatting, misplaced or redundant content, and overall paper structure that makes the narrative hard to follow. Impact Level of Rebuttal Segment CRPConcrete Revision Performed: authors point to specific changes or verifiable artifacts already added (new text/sections, new experiments/tables/figures, updated numbers, released code/data). Cues: âWe added/updated . . . â, âSection X rewritten . . . â, âNew ablation in Sec. . . . shows . . . â, âCode/data at . . . â. SRPSpecific Revision Plan: authors commit to concrete future edits with where/what to change, but not yet implemented. Cues: âWe will add an ablation in Sec. X . . . â, âWe will redraw Fig. . . . â, âWe will clarify definitions in §. . . â. VCRVague Commitment to Revise: promises to improve without actionable details (no locations, artifacts, or timelines). Cues: âWe will revise accordingly.â, âWe will improve clarity/writing.â DWCDefend Without Change: argues the paper already addresses the point; no edits proposed. Cues: âAlready covered in Sec. . . . â, âSetup is standard.â, âClaim stands.â DRFDeflect/Reframe: shifts responsibility or reframes the issue; no change offered. Cues: âReviewer misinterprets . . . â, âOut of scope . . . â, âReviewer phrasing is incorrect.â Table 7: Summary of perspective labels for review segments and impact labels for rebuttal segments in §4.1. PerspectiveCRPSRPVCRDWCDRF Evaluation (11,257)4,766 / 42.3%903 / 8.0%171 / 1.5%5,249 / 46.6%168 / 1.5% Experiments (25,160)12,059 / 47.9%2,272 / 9.0%401 / 1.6%9,833 / 39.1%595 / 2.4% Novelty (8,585)2,828 / 32.9%872 / 10.2%185 / 2.2%4,578 / 53.3%122 / 1.4% Presentation (4,776)2,894 / 60.6%803 / 16.8%256 / 5.4%784 / 16.4%39 / 0.8% Reproducibility (4,402)2,009 / 45.6%465 / 10.6%120 / 2.7%1,747 / 39.7%61 / 1.4% Theory (12,822)4,253 / 33.2%1,110 / 8.7%282 / 2.2%6,859 / 53.5%318 / 2.5% Writing (8,540)4,693 / 55.0%1,149 / 13.5%631 / 7.4%1,997 / 23.4%70 / 0.8% Overall33,502 / 44.3%7,574 / 10.0%2,046 / 2.7%31,047 / 41.1%1,373 / 1.8% Table 8: Label distribution by perspective (counts / %). D.2 Baseline Adaptation Details All baselines of âOther Methodsâ receive the same paper text and the same requested perspective, and are required to output exactly one review segment following our unified format (Appendix D.4). When a baseline naturally produces longer outputs (e.g., multi-section reviews, multiple bullets, or agent traces), we apply a deterministic post-processing rule to extract one segment, described below. MARG.MARG [DâArcy et al., 2024b] generates review feedback via multiple LLM instances (agents) that each read a portion of the paper and then aggregate intermediate discussion into final comments. To adapt MARG to our single-segment, perspective-conditioned setting, we inject the requested perspective into the leaderâs instruction so that the final synthesized comment targets the requested perspective. We output only the leaderâs final synthesized comment and discard intermediate agent messages. If multiple bullets are produced in the final synthesis, we take the first bullet as the single segment. LimGen. LimGen [Faizullah et al., 2024] focuses on generating limitations with suggestions for improvement. We adapt LimGen by prompting it to generate limitations conditioned on the requested perspective. If the output 14 SettingSFTDPO FrameworkLLaMA-FactoryLLaMA-Factory Base modelLlama-3.1-8B-InstructLlama-3.1-8B-Instruct Precisionbf16bf16 GPUH200 141 GBH200 141 GB Context limit (cutoff len)32,76832,768 Generation limit (maxnewtokens)512512 DatasetREVIEWSEG-SFT-13KREVIEWPREF-DPO-22K ObjectiveToken-level CEDPO (sigmoid) Epochs / Steps3 / 4,9892 / 2,728 LR / Scheduler1eâ4 / cosine1eâ5 / cosine Warmup ratio0.100.05 Batch (per-devĂ accum)1Ă 81Ă 16 Eff. batch816 LoRA targetq,k,v,o,gate,up,downq,k,v,o,gate,up,down LoRA rank / α / dropout8 / 16 / 0.0516 / 16 / 0.05 Grad. checkpointingOffOn Reference modelN/AFrozen SFT Ref. quantizationN/A4-bit DeepSpeed (ZeRO-2)OptionalOn FA2 (FlashAttention-2)OnOn Wall-clock timeâ 120 hâ 203 h Table 9: Key training configuration for SFT and DPO. âEff. batchâ is per-device batch sizeĂaccumulation steps. DeepSpeed denotes ZeRO-2; FA2 denotes FlashAttention-2. contains multiple limitation items, we deterministically select the first item (including its paired suggestion when present) as the single segment. We only normalize surface formatting (e.g., removing numbering or headings) and do not add new content. DeepReviewer-14B. DeepReviewer-14B [Zhu et al., 2025b] is trained to produce comprehensive, structured reviews. To make it comparable in our setting, we prompt DeepReviewer-14B to produce a single perspective- specific comment in our segment format. If the model outputs a full multi-section review despite the instruction, we extract the subsection that best matches the requested perspective using a fixed heading-based mapping (e.g., Exper- iments/Empirical Evaluationâexperiments; Writing/Presentation/Clarityâclarity; Impact/Significance/Novelty âcontribution/impact), and then take the first paragraph in that subsection as the segment. We run DeepReviewer- 14B in pure generation mode without external tools to keep inputs comparable. D.3 Human Expert Evaluation Protocol This appendix describes the interface and rubric we used to evaluate review segments and actionable rebuttal guidance. D.3.1 Interface We collect two judgments per comparison on the page shown in Figure 9: (i) a pairwise preference between two anonymized candidate reviews (A vs. B vs. Tie), and (i) per-candidate 1â5 ratings on five dimensions: Actionability, Specificity, Groundedness, Relevance, and Helpfulness. Anchors are1 = Very poor,2 = Poor, 3 = Fair, 4 = Good, 5 = Excellent. Raters judge only using evidence in the paper. They should avoid hallucinations and must not penalize a section for missing information if the same information exists elsewhere in the paper. D.3.2 Dimension Rubrics (Concise) 1. Actionability: Clear next steps with parameters or acceptance criteria as shown in Figure 10. 2. Specificity: Pinpoints exact sections, figures, metrics, or settings as shown in Figure 11. 3. Groundedness:Supported by the paper with explicit references or numbers as shown in Figure 12. 4. Relevance: Aligned with the target perspective and main contributions as shown in Figure 13. 5. Helpfulness: Clear, constructive guidance that improves the paper as shown in Figure 14. 15 D.4 LLM-as-a-Judge Evaluation Here we discuss some details of LLM-as-a-judge evaluation. The prompt for point-wise evaluation is Figure 15. The prompt for pairwise evaluation on Actionability is Figure 16. To visualize pairwise evaluation results across seven different perspectives, Figure 17 shows heatmaps where each cell is the win rate of the row model over the column model. D.5 Automatic Evaluation Besides human and LLM-as-a-Judge evaluation, we also evaluate our method and different baselines by some metrics. Results are shown in Table 10. As mentioned, we evaluate all 700 test instances from evaluation dataset with four metrics: BLEU@4 [Papineni et al., 2002], ROUGE-L sum (F1) [Lin, 2004], METEOR [Banerjee and Lavie, 2005], and chrF [Popovi Ì c, 2015]. Overall, our models perform competitively against baselines. RBTACT achieves the best ROUGE-L sum (12.64) and METEOR (11.65), while RBTACT-SFT attains the highest BLEU@4 (14.93). Both RBTACT variants outperform large proprietary and open models on these three metrics. On chrF, GPT-5-chat is strongest, and our models are close to each other (18.51 vs. 18.57). Comparing RBTACT with RBTACT-SFT, differences are small. RBTACT slightly improves over SFT on ROUGE- L sum and METEOR. BLEU@4 is close to RBTACT-SFT. These trends indicate that rebuttal optimization refines phrasing but does not drastically change surface overlap. ModelBLEU@4 ROUGE-L sum METEOR chrF RBTACT (Ours)14.6212.6411.6518.57 RBTACT-SFT14.9312.3311.5318.51 GPT-5-chat11.1710.119.9624.90 DeepSeek-V3.210.199.768.4917.78 Llama-3.1-70B10.488.768.2716.42 Qwen-3-32B9.728.588.1417.18 DeepReview12.4011.3010.2019.60 MARG10.959.108.6018.10 LimGen10.908.207.9017.40 Table 10: Automatic evaluation on 700 test instances. Values are reported as percentages. E Case Study Here we show some case studies to demonstrate why our model RBTACT is more actionable than other baselines in most cases from different perspectives. They are shown in Figure 18. F Additional Analyses F.1 Severity and Paper-Strength Analyses. We further analyze whether rebuttal-derived supervision over-emphasizes minor issues and whether gains vary with paper strength. We studies (i) issue severity using a perspective-based proxy and shows that our training signal substantially covers major issues that are often defended, and (i) the relationship between paper strength (OpenReview ratings) and actionability improvements. F.1.1 Are we biased toward minor issues? A severity proxy. We use a lightweight proxy for issue severity based on the review perspective labels. We treat major issues as Experiments, Evaluation, Theory, Novelty, Reproducibilityand minor issues asWriting, Presentation. Table 11 aggregates the impact-label counts by this proxy. As a result, major issues are indeed defended more often (DWC 45.4% vs. 20.9%), which matches the intuition that authors push back on higher-stakes critiques. However, rebuttal supervision is not dominated by minor issues: over half of major-issue mappings correspond to specific revisions (CRP+SRP = 50.7%). Moreover, our preference pairs are constructed within the same paper and perspective and balanced across perspectives (§4.2), which reduces the risk that training is driven primarily by easy-to-fix minor edits. 16 Impact label Major (Ev/Ex/Th/No/Re) Minor (Wr/Pr) CRP25,915 (41.6%) 7,587 (57.0%) SRP5,622 (9.0%) 1,952 (14.7%) VCR1,159 (1.9%)887 (6.7%) DWC28,266 (45.4%) 2,781 (20.9%) DRF1,264 (2.0%)109 (0.8%) Table 11: Impact-label distribution by a perspective-based severity proxy. Major issues are more frequently defended (DWC), but still exhibit substantial author uptake: CRP+SRP is 50.7% for major issues and 71.6% for minor issues. Bucket (by mean rating) n RBTACT Baseline â (ours â base) Weak (bottom 1/3)353.353.08+0.27 Medium (middle 1/3)353.333.11+0.22 Strong (top 1/3)353.463.27+0.19 Table 12: Actionability by paper-strength bucket. F.1.2 Do gains correlate with paper strength? Setup. We operationalize paper strength using the mean reviewer rating on OpenReview (averaged across reviewers for the same submission). We split papers into three buckets (Weak / Medium / Strong) by tertiles of the mean rating. On the 105-paper LLM-as-a-judge set, we compute the average Actionability score per bucket. Interpretation.The trend suggests that while stronger papers already elicit reasonably actionable feedback even from strong baselines, weaker papers benefit more from rebuttal-anchored supervision. This is aligned with our training signal: rebuttals often make explicit which critiques lead to concrete revisions versus defenses, which helps the model prioritize actionable, fixable issues. 17 [System Prompt] You are a professional academic review analysis assistant. Your task is to perform precise one-to-one mapping between review weakness points and author rebuttal responses. Guidelines for mapping: 1. Carefully analyze the rebuttal text to identify which sections respond to specific weaknesses 2. Look for explicit references (W1, W2, Point 1, etc.) or implicit topical connections 3. Extract the complete response content that addresses each weakness 4. Assign confidence scores (0-1) based on the clarity and directness of the mapping 5. Mark as âNo Responseâ if a weakness is not addressed in the rebuttal 6. Be conservative with confidence scores - only use high scores (Âż0.8) when the mapping is very clear 7. Preserve the exact wording from the rebuttal when extracting responses CRITICAL RULE - NO SHORTCUTS OR REFERENCES: You must NEVER use summarizing phrases or references like â[Same content as W2 response]â, â[Similar to above]â, â[As mentioned earlier]â, etc. Always copy the complete, verbatim text from the rebuttal for each weakness point, even if the same rebuttal section addresses multiple weaknesses. Repetition is required and expected - do not try to avoid it. [User Prompt] Please map each weakness point in ÂĄweakness pointsÂż to its corresponding rebuttal response from ÂĄrebuttaltextÂż: ÂĄweakness pointsÂż weaknesseslist ÂĄ/weaknesspointsÂż ÂĄrebuttal textÂż rebuttal text ÂĄ/rebuttal textÂż MANDATORY REQUIREMENTS: - Output a mapping line for EVERY weakness point (W1, W2, W3, ... up to Wlen(weaknesspoints)) - Use exactly âW[number] -Âż R[number]:â or âW[number] -Âż No Responseâ format - Include confidence score in parentheses for every mapping - Do not skip any weakness numbers - If you cannot find a rebuttal response, write âNo Responseâ instead of omitting the line CRITICAL: You MUST provide a mapping for EVERY weakness point listed above. Do not skip any weakness points. ABSOLUTELY CRITICAL - NO SUMMARIZATION OR SHORTCUTS FOR REBUTTAL CONTENT: - ALWAYS copy the COMPLETE, FULL, VERBATIM text from the rebuttal for each weakness, even if content is identical to previous responses - NEVER use summarizing or abbreviated phrases like â[Same content as W2 response]â or â[Similar to above]â or â[Full content identical to W5 response]â - always provide the complete original text - Do NOT abbreviate, summarize, or reference other responses - If the same rebuttal section addresses multiple weaknesses, copy the ENTIRE text in full for each relevant weakness - NEVER add any commentary, explanation, or meta-text beyond the actual rebuttal content ÂĄoutput formatÂż (use EXACTLY this format): W1 -Âż R1: [Specific rebuttal content addressing W1] (Confidence: 0.x) W2 -Âż R2: [Specific rebuttal content addressing W2] (Confidence: 0.x) W3 -Âż R3: [Specific rebuttal content addressing W3] (Confidence: 0.x) ...continue for ALL weakness points... If no rebuttal response exists for a weakness: Wx -Âż No Response (Confidence: 1.0) ÂĄ/output formatÂż Figure 5: Prompt used for mapping review segments with rebuttal segments. 18 [System Prompt] You are an expert in academic paper review classification. Your task is to identify from which perspective the reviewer is raising concerns or questions about the paper. [User Prompt] The following review point is a weakness or question raised by a reviewer during peer review. Please classify this review point based on the PERSPECTIVE from which the reviewer is critiquing or questioning the paper: 1. **Experiments**: The reviewer is questioning experimental **setup and design**. This includes missing or insufficient experiments, lack of ablation studies, weak baseline comparisons, unclear descriptions of datasets, or issues with hyperparameter selection. 2. **Writing**: The reviewer is concerned about writing quality - grammar, clarity, readability, ambiguous phrasing, typos, missing definitions of symbols/terms, unclear explanations of concepts. 3. **Presentation**: The reviewer is critiquing presentation and organization - figures, tables, and organization issues, unclear plots, missing legends, poor formatting, misplaced content, overall paper structure making it hard to follow. 4. **Theory**: The reviewer is questioning theoretical aspects - incorrect mathematical derivations, flawed assumptions, weak theoretical justification, missing proofs, inconsistency between claims and formulas. 5. **Novelty**: The reviewer is questioning novelty and originality - lack of novelty or originality, overlap with prior work, incremental contribution, insufficient differentiation from existing methods. 6. **Reproducibility**: The reviewer is concerned about reproducibility - missing implementation details, absent code or pseudo-code, hyperparameters not specified, insufficient information to reproduce results. 7. **Evaluation**: The reviewer is concerned about how the experimental results are **measured, analyzed, and interpreted**. This includes the use of inappropriate or missing evaluation metrics, insufficient analysis of results, or inconsistencies between reported results and the paperâs claims. 8. **Miscellaneous**: Content that is not a direct review point (weaknesses, questions, suggestions) about the paper. This includes polite remarks, Summative or transitional comments, summaries of the paperâs or reviewâs content, or irrelevant text. Please analyze the following review point and identify from which perspective the reviewer is raising their concern. Respond with ONLY the category name (Experiments, Writing, Presentation, Theory, Novelty, Reproducibility, Evaluation, Miscellaneous). Review point to classify: weakness content Perspective: Figure 6: Prompt used to label which perspective a review segment belongs to. 19 [System Prompt] You are a precise, deterministic classifier for rebuttal responses. [User Prompt] Return only a compact JSON object: âimpactâ: âÂĄoneof:[CRP,SRP,VCR,DWC,DRF]Âżâ. Categories: - CRP: Concrete revision already made or concrete, verifiable artifact provided (new text/sections, new experiments/ta- bles/figures, code/data links, numbers). Cues: âWe added/updated...â, âSection X rewritten...â, âNew ablation in Sec. ... shows ...â, âCode/data at ...â. - SRP: Specific revision plan committed, but not yet implemented; where/what to revise is specific. Cues: âWe will add ablation in Sec. X...â, âWe will redraw Fig. ...â, âWe will clarify definitions in §...â. - VCR: Vague promise to revise; no concrete actions, locations, or artifacts. Cues: âWe will revise accordingly.â, âWe will improve writing/clarity.â. - DWC: Defend current paper as-is; no new change proposed. Cues: âAlready covered in Sec. ...â, âSetup is standardâ, âClaim standsâ. - DRF: Shift issue to reviewer or avoid underlying point; no change offered. Cues: âReviewer misinterprets ...â, âOut of scope ...â, âReviewer phrasing is incorrectâ. Figure 7: Prompt used to label which impact category a rebuttal segment belongs to. [System Prompt] You are a professional reviewer. Provide a comment such as weakness, question or suggestion on the given paper in 1 to 3 sentences. [User Prompt] Request: From the perspective of x, provide a comment on the following paper. <PAPERCONTEXT> Figure 8: Prompt used to generate review segments by different LLM baselines. 20 Figure 9: The interface of our human expert evaluation. The page contains 2 tasks: pairwise preference and per-candidate 1â5 ratings. 21 Figure 10: Comparison guidelines for the âActionabilityâ criterion. Figure 11: Comparison guidelines for the âSpecificityâ criterion. 22 Figure 12: Comparison guidelines for the âGroundednessâ criterion. Figure 13: Comparison guidelines for the âRelevanceâ criterion. 23 Figure 14: Comparison guidelines for the âHelpfulnessâ criterion. 24 [System Prompt] You are an expert in evaluating peer review quality. Your task is to assess a peer review comment from multiple dimensions and provide scores (1-5) with detailed reasoning for each dimension. [User Prompt] Please evaluate the following peer review comment based on the scoring rubric provided. Scoring Rubric ### Actionability (1-5) 1. Very poor: No concrete next step. Vague remarks like âimprove experiments.â 2. Poor: A possible step is implied but not described. No criteria for success. 3. Fair: At least one concrete suggestion, but incomplete or underspecified. 4. Good: Clear, feasible steps with some parameters or success criteria. 5. Excellent: A short plan with steps, locations in the paper, parameters or tests, and what outcome would address the issue. ### Specificity (1-5) 1. Very poor: Generic template text that could apply to any paper. 2. Poor: Mentions broad areas but no details. 3. Fair: Refers to a section, figure, dataset, or claim but stays broad. 4. Good: Points to exact sections, figures, metrics, or settings. 5. Excellent: Pinpoints precise passages or numbers and names exact variables, metrics, or ablation locations. ### Groundedness (1-5) 1. Very poor: Speculative, incorrect, or contradicted by the paper. 2. Poor: Weak link to the paper; no verifiable reference. 3. Fair: Partly grounded with at least one reference to paper content. 4. Good: Well supported with references to specific content. 5. Excellent: Strongly supported with exact identifiers or numbers from the paper (for example âTable 2 shows 71.3 vs 71.1 and the claim of a large gain is not supportedâ). ### Relevance (1-5) 1. Very poor: Off topic relative to the target perspective or the main paper issues. 2. Poor: Mostly off topic with minor relevant content. 3. Fair: Partially aligned. Mixes relevant and irrelevant feedback. 4. Good: Mostly aligned with the target perspective. 5. Excellent: Fully aligned with the target perspective and the paperâs main contributions. ### Helpfulness (1-5) 1. Very poor: Unclear, hostile, or not useful. 2. Poor: Slightly useful but confusing or impractical. 3. Fair: Some useful content, needs refinement to be actionable. 4. Good: Clear, constructive, and practically useful. 5. Excellent: Directly helps the authors improve the paper with minimal ambiguity. **Paper Content:** paper content **Review Perspective:** perspective **Review Comment to Evaluate:** reviewtext Please provide scores (1-5) for each dimension along with your reasoning. Be critical and precise in your evaluation. You MUST respond with a valid JSON object in the following format (no markdown code blocks, just raw JSON): âactionabilityscoreâ: ÂĄ1-5Âż, âactionabilityreasoningâ: âÂĄbrief explanationÂżâ, ...... Figure 15: Prompt used for point-wise evaluation for LLM-as-a-judge. 25 [System Prompt] You are an impartial judge comparing the actionability of two peer-review segments. Actionability means the feedback gives concrete, specific, and feasible guidance that authors can directly implement. Prefer segments that: -specify what to change (methods, experiments, analyses, writing), -localize where to change (section/figure/table/scope), -propose how to change (procedures, metrics, datasets, ablations, edits), -include verifiable artifacts or acceptance criteria (e.g., code/data, new experiments, numbers to report). Output JSON schema:âwinnerâ: âAâ â âBâ, âjustificationâ: â1â2 sentences citing the most decisive actionable cues.â [User Prompt] Task: Choose the more actionable review segment for the specified perspective. Remember: no ties. Perspective: <PERSPECTIVE> Paper context: <PAPER CONTEXT> Segment A: <REVIEWSEGMENTA> Segment B: <REVIEW SEGMENTB> Figure 16: Prompt used for pairwise evaluation for LLM-as-a-judge on Actionability. 26 RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 57.163.861.965.768.666.765.776.2 42.955.257.157.163.858.154.371.4 36.244.853.357.161.058.150.567.6 38.142.946.752.455.255.249.566.7 34.342.942.947.651.453.350.564.8 31.436.239.044.848.650.549.561.0 33.341.941.944.846.749.545.759.0 34.345.749.550.549.550.554.359.0 23.828.632.433.335.239.041.041.0 Overall RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 40.060.060.066.773.366.793.366.7 60.053.353.360.073.360.053.393.3 40.046.753.373.373.353.353.360.0 40.046.746.753.366.766.753.366.7 33.340.026.746.740.053.360.066.7 26.726.726.733.360.053.353.360.0 33.340.046.733.346.746.753.353.3 6.746.746.746.740.046.746.740.0 33.36.740.033.333.340.046.760.0 Experiments RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 60.073.380.0100.066.766.753.366.7 40.060.040.033.380.060.060.073.3 26.740.040.046.760.046.733.373.3 20.060.060.046.746.746.753.373.3 0.066.753.353.380.040.053.353.3 33.320.040.053.320.053.320.066.7 33.340.053.353.360.046.726.786.7 46.740.066.746.746.780.073.366.7 33.326.726.726.746.733.313.333.3 Evaluation RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 53.386.740.060.073.353.360.080.0 46.733.366.780.046.760.053.353.3 13.366.766.753.346.766.746.766.7 60.033.333.360.046.753.353.360.0 40.020.046.740.066.760.046.766.7 26.753.353.353.333.346.746.753.3 46.740.033.346.740.053.340.060.0 40.046.753.346.753.353.360.060.0 20.046.733.340.033.346.740.040.0 Reproducibility RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 53.353.360.073.373.366.766.766.7 46.760.046.753.373.346.766.780.0 46.740.053.360.066.766.746.773.3 40.053.346.753.353.353.340.080.0 26.746.740.046.746.753.340.066.7 26.726.733.346.753.346.766.760.0 33.353.333.346.746.753.353.353.3 33.333.353.360.060.033.346.760.0 33.320.026.720.033.340.046.740.0 Novelty RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 80.053.366.746.766.780.060.0100.0 20.053.366.753.353.360.046.780.0 46.746.746.766.753.360.046.753.3 33.333.353.353.340.060.040.060.0 53.346.733.346.740.040.053.373.3 33.346.746.760.060.046.753.366.7 20.040.040.040.060.053.333.346.7 40.053.353.360.046.746.766.753.3 0.020.046.740.026.733.353.346.7 Theory RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 53.366.766.753.366.766.760.073.3 46.753.360.060.060.060.046.753.3 33.346.753.340.053.366.766.773.3 33.340.046.740.066.753.360.060.0 46.740.060.060.046.753.353.353.3 33.340.046.733.353.346.746.760.0 33.340.033.346.746.753.340.040.0 40.053.333.340.046.753.360.053.3 26.746.726.740.046.740.060.046.7 Writing RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen RbtAct GPT-5-chat DeepSeek-V3.2 Llama-3.1-70B DeepReview-14b MARG Qwen-3-32B RbtAct-SFT LimGen 60.053.360.060.060.066.766.780.0 40.073.366.760.060.060.053.366.7 46.726.760.060.073.346.760.073.3 40.033.340.060.066.753.346.766.7 40.040.040.040.040.073.346.773.3 40.040.026.733.360.060.060.060.0 33.340.053.346.726.740.073.373.3 33.346.740.053.353.340.026.780.0 20.033.326.733.326.740.026.720.0 Presentation 020406080100 Win rate (%) Figure 17: Pairwise win rates heatmaps by perspective (row beats column) of §5.3 27 Case Study of Why RBTACT Reviews More Actionable RBTACT Review Feedback: The reported gains may be entangled with augmentation choices (MixUp/CutMix) and unquantified run-to-run vari- ance, which leaves robustness in §5.2 not directly stress- tested; to disambiguate these factors, re-train without Mix- Up/CutMix under a fixed seed with three independent trials, report mean±std Top-1 in Table 3, and add a Corrupted ImageNet check aligned with the claims. Qwen-3-32b Review Feedback: The current training description under-specifies key reg- ularization and sizing decisions, making it unclear how sensitive results are to setup and randomness. It would help to include a compact ablation that revisits weight decay and batch size. Analysis: It specifies exact factors to remove, run counts, where to report (Table 3), and a concrete robustness checkâclear steps and success criteria. RBTACT Review Feedback: In Figs. 2-3, small axis labels, palettes that are not color- blind-safe, missing y-axis units, and cramped legends re- duce interpretability; revising the plots with adequately sized labels, an OKLCH-based color-blind-safe palette, ex- plicit y-axis units, legends moved below the panels, and captions define each metric and sample size would make the figures verifiable. Llama-3.1-70b Review Feedback: Figure readability is limited by low apparent resolution, small typography, weak color separability, and inconsistent styling, with legend placement adding clutter. You may consider increasing DPI and font sizes, adopting more dis- tinguishable colors under a unified style, clarifying captions in figure. 2. Analysis: It names specific figures, thresholds, concrete palette guidance, and exact caption contentâeasy to implement and verify. RBTACT Review Feedback: Extend Sec. 4.2: add baselines and SOTA with identical prompts; report macro-F1 and calibration with 95% CIs via paired bootstrap over papers. Include error taxonomy in Appendix. MARG Review Feedback: Broaden evaluation to more datasets and compare against stronger models. Include confidence intervals where appro- priate and a brief qualitative error analysis. Analysis: It pins down section, comparators, metrics, CI method, and deliverables, giving precise implementation and acceptance checks. Figure 18: Case study comparing review feedback on Actionability from Experiment, Presentation and Evaluation perspectives. 28