Paper deep dive
PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.23794v1
- Canonical: https://arxiv.org/abs/2607.23794v1
Trouble viewing inline? Open PDF directly →
Full Text
67,927 characters extracted from source content.
Expand or collapse full text
IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 20261 PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Member, IEEE , and Dan Hu Abstract — Pathological diagnosis is inherently multi- scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely de- veloped under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reason- ing. Moreover, naively constructed visual question answer- ing (VQA) tasks may be susceptible to text-only or su- perficial visual shortcuts, leading to unreliable assess- ments of visual understanding. To address these limita- tions, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for se- mantic reasoning questions and a Structure-controlled Dis- tractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diag- nostic paths across multiple magnification levels. Build- ing on the semantic reasoning set, PathScale-R1 is opti- mized through Difficulty-driven Reasoning Distillation su- pervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which en- courages the use of evidence across magnifications. Exten- sive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effec- tive transfer to conventional single-scale pathology VQA. Our code is available at PathScale-R1. Index Terms— Vision-language model, Reinforcement Learning, Reasoning, Pathological Image Analysis. This work was supported by the Ministry of Education Tier 1 grant, Singapore (24-1250-P0001), and the Ministry of Education Tier 2 grant, Singapore (T2EP20224-0028). This work was powered by the UnPuzzle & PuzzleCloud Platform (https://puzzlelogic.com/unpuzzle) and sup- ported by PuzzleLogic Pte Ltd, Singapore. ChiPhanandTianyiZhangcontributedequallytothis work. Corresponding Author: Sudong Wang (e-mail: sudong- wang@puzzlelogic.com), Yueming Jin (e-mail: ymjin@nus.edu.sg) and Dan Hu (e-mail: hudan@fjmu.edu.cn). Chi Phan, Tianyi Zhang, and Yueming Jin are with Department of Electrical and Computer Engineering, National University of Singa- pore, Singapore 117417 (e-mails:chiphan, zhangtianyi@u.nus.edu; ymjin@nus.edu.sg). Yufeng Wu, Linghan Cai, Zeyu Liu, and Sudong Wang are with Puz- zleLogic Pte Ltd, Singapore 229594 (e-mails:yufengwu, linghancai, zeyuliu, sudongwang@puzzlelogic.com). Qiaochu Xue and Yueming Jin are with Department of Biomedical Engineering, National University of Singapore, Singapore 117417, Sin- gapore (e-mails: e1352520@u.nus.edu; ymjin@nus.edu.sg). Jiajie Zhang and Dan Hu are with Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital, Fuzhou, China (e-mails: naili816@126.com; hudan@fjmu.edu.cn). I. INTRODUCTION P ATHOLOGY is the gold standard for cancer diagnosis, providing essential evidence for disease characterization, treatment planning, and prognosis [1]. Due to the hierarchical organization and heterogeneity of tissue morphology, patho- logical interpretation requires intensive visual analysis and substantial domain expertise. Recent developments in large language models and multimodal learning have motivated pathology-focused Vision-Language Models (VLMs) [2]–[8], which show promising capabilities in pathological under- standing, visual question answering (VQA), and diagnostic support. Despite these substantial advancements, there is still a noticeable gap between the underlying design of current models and the practical workflow of pathological assessment. Real-world pathological diagnosis is fundamentally a cross- scale reasoning process, in which pathologists systematically navigate across different magnifications to synthesize a final decision [9], [10], [11]. Slide review typically begins at low power to assess global tissue architecture and lesion distri- bution, proceeds to intermediate magnifications to examine localized tissue organization, and finally uses high magnifi- cation to verify cellular morphology and fine-grained diag- nostic features [12], [13]. These observations are highly in- terdependent, with macroscopic findings directing subsequent zoom-in decisions and microscopic evidence validating earlier structural impressions. Ultimately, a diagnosis is synthesized by merging this architectural, structural, and cellular evidence across scales. Therefore, a clinically meaningful pathology VLM should not only recognize findings at one specific scale, but also needs to connect cross-scale evidence to support a coherent diagnostic conclusion. However, existing pathology VLMs [2]–[8] are predomi- nantly developed and evaluated in single-scale settings, which do not explicitly capture how diagnostically related evidence is connected across magnifications (Fig. 1). Patch-level VQA benchmarks such as PathVQA [14], PathMMU [15], and Path- Bench [16] contain pathology images from diverse sources and magnification levels, but each question is typically associated with only a single ROI. They therefore assess recognition and reasoning within an isolated field of view, rather than whether a model can connect low-power tissue architecture, intermediate-scale tissue organization, and high-power cellular morphology from the same diagnostic trajectory. At the whole- slide image (WSI) level, benchmarks such as WSI-VQA [17], arXiv:2607.23794v1 [cs.CV] 26 Jul 2026 2IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2026 Cross-scale Visual Grounding Ground truth (D) Question: The image above shows a 40× magnified patch extracted from the 10× patch below. Each option shows the same 10× patch with a different masked region. Which masked region corresponds to the location of the 40× patch? Mid-mag Options: Masked Low-mag 1Masked Low-mag 2Masked Low-mag 3Masked Low-mag 4 (A) (B) (C) (D) Task: Origin Grounding Text- only infer “This question cannot be answered without the images because it requires a direct visual comparison of textures, structures, and landmarks between the two magnification levels. Without the visual data, there is no text or coordinate information available to deduce the correct answer. Please upload the images!” Text- only infer Cross-scale Semantic Reasoning Low- mag Mid- mag High-mag Task: Localization “This question can't be answered correctly without the images because all four options describe plausible, [...]. Without seeing the exact coordinates of the cells, specifically whether they are hugging the airways (peribronchiolar) [..] there is no biological "rule" to favor one location over the others.” Question: Which option best characterizes the compartmental context of the structures shown in this set of images? (A) Cells are predominantly within peribronchiolar connective tissue, with partial (B) Cells are predominantly within alveolar airspaces (intra-alveolar), with partial (C) Cells are predominantly within the alveolar septal interstitium, with partial (D) Cells are predominantly within perivascular connective tissue, with partial Options: Ground truth: (C) (C) Our PathScale-VQA: Shortcut-Resistant Cross-Scale Pathological VQA Question: Based on the cellular morphology, which layer of stratified squamous epithelium is most prevalent on the left side of the image? (A) Spinous layer (B) Basal layer Ground truth: (A)Source: PathMMU (C) Granular layer (D) Keratin layer (A) Single-scale VQA Question: (A) Crowded epithelial cell populations with irregular nuclei reduce intercellular spacing, producing architectural distortion at low magnification (B) Stromal collagen deposition compresses otherwise orderly epithelial cells, altering the low-magnification outline (C) Superficial epithelial flattening leads to apparent irregularity without corresponding nuclear changes (D) Muscular layer thickening alters surface epithelial alignment despite preserved nuclear morphology (A) Ground truth: Which explanation best connects the low-magnification pattern and the high-magnification atypical nuclear findings in this case? (B) Cross-Scale VQA with Text-Only Shortcut Solution Gemini 3 Pro Text- only infer “The correct option is A. I can answer this without the images because pathology follows a strict cause-and-effect logic: "distorted" or "thickened" pattern at low magnification must be physically caused by "crowded" and "atypical" cells at high magnification. Option A is the only choice that provides this biological link, [...]” High- mag collapse of adjacent alveolar spaces collapse of adjacent alveolar spaces collapse of adjacent alveolar spaces collapse of adjacent alveolar spaces Fig. 1. Comparison of Pathology VQA designs. (A) Single-scale VQA evaluates isolated local recognition. (B) Naively constructed cross-scale VQA can still be solved from textual cues alone, revealing text-only shortcut risks. (C) Our PathScale-VQA uses expert-verified diagnostic paths to construct shortcut-resistant semantic reasoning and visual grounding tasks that require cross-scale visual evidence. (mag: magnification) SlideBench [7], and WSI-Bench [8] evaluate slide-level un- derstanding from gigapixel images. Existing WSI VLMs [5], [7], [8] commonly make such inputs computationally tractable by aggregating patch embeddings at a specific magnification into a compact slide-level representation. Such representa- tions support slide-level prediction, yet still provide a limited assessment of how findings across scales interact. Under these single-scale framings, models may learn scale-specific visual recognition, but receive little supervision for the hier- archical reasoning process that links low-power architectural context to high-power cellular evidence. Consequently, they may perform well on existing benchmarks while struggling in tasks that require cross-scale understanding, leading to a critical misalignment with clinical practice. Addressing this limitation requires moving beyond the current single-scale framing toward a cross-scale reasoning paradigm centered on recognizing and integrating evidence across magnifications. Constructing cross-scale reasoning VQA tasks, however, encounters another critical challenge: the susceptibility to text- based shortcut solutions. In clinical pathology, prior knowl- edge and textual clinical context can guide interpretation, but assessments of tissue architecture, local organization, and cellular morphology are grounded in the visual evidence. Yet recent studies [18]–[20] have shown that VLMs can achieve high benchmark performance by exploiting linguistic cues, weak distractors, or dataset-specific artifacts, rather than rely- ing on the provided images. In medical VQA, such shortcuts may also yield fluent and seemingly image-grounded explana- tions despite limited visual dependence [16], [21]. This risk is particularly relevant in cross-scale pathology VQA, where magnification-specific terminology and stereotyped biomedical associations in the answer choices may reveal the correct response without requiring integration across scales (Fig. 1). Such leakage can create an “illusion of visual understand- ing” [21], making performance metrics unreliable for mea- suring model capabilities, which is particularly concerning in critical domains such as pathology. Therefore, constructing a reliable cross-scale reasoning setting requires not only collect- ing integrated multi-magnification images, but also actively suppressing shortcut solutions to enforce genuine reliance on the visual evidence. To address these challenges, we introduce a novel bench- mark and training framework for shortcut-resistant cross-scale pathology reasoning (Fig. 2). At the data level, we shift the basic unit of supervision and evaluation from an isolated ROI to a diagnostic path, defined as a pathologist-verified trajectory connecting clinically corresponding regions at 10×, 40×, and 200× within the same WSI. To ensure that benchmark performance reflects genuine use of visual evidence, we for- mulate two complementary task categories, each paired with a targeted curation strategy: adversarial text-only screening for textual options, and structure-controlled distractor sam- pling for image options. At the model level, we develop PathScale-R1, a pathology VLM trained to integrate diagnostic evidence across scales through a two-stage framework. The first stage performs Difficulty-driven Reasoning Distillation, transferring structured cross-scale rationales on samples that are challenging for the base model. The second stage applies reinforcement learning with a novel Scale-aware Reasoning Structure reward, encouraging the model to produce coherent reasoning that incorporates evidence from multiple magnifi- cations. Extensive experiments demonstrate that PathScale-R1 substantially improves cross-scale semantic reasoning while maintaining strong performance on conventional single-image pathology VQA benchmarks. An overview of the dataset and benchmark performance is provided in Fig. 3. The main contributions of this work are as follows: • We introduce a cross-scale formulation of pathological image analysis, in which the diagnostic path across magnifications serves as the basic unit of supervision PHAN et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING3 and evaluation. Our curated cross-scale VQA benchmark, PathScale-VQA, comprises 10,373 questions grounded in 1,368 pathologist-verified diagnostic paths designed to reflect the practical workflow of pathological assessment. • We develop a shortcut-resistant VQA curation pipeline that combines adversarial text-only screening with structure-controlled distractor sampling, reducing linguis- tic leakage and superficial visual cues for more reliable evaluation of cross-scale visual understanding. • We propose PathScale-R1, a two-stage optimization framework combining Difficulty-driven Reasoning Dis- tillation with reinforcement learning guided by the novel Scale-aware Reasoning Structure Reward, enabling more effective integration of evidence across scales. • Extensive experiments demonstrate that PathScale-R1 achieves state-of-the-art performance on the cross-scale benchmark and strong transferability to outperform the evaluated baselines on conventional single-scale pathol- ogy VQA. Our findings also further reveal that fine- grained cross-scale visual grounding is an important yet underdeveloped capability of current VLMs. This work substantially extends our preliminary MICCAI 2026 version [22] in four aspects. First, we expand the original benchmark from 4,685 to 10,373 questions, with substan- tially more annotated WSIs and diagnostic paths, thereby improving data coverage and clinical grounding. Second, we introduce five visual grounding tasks that reflect key operations in pathological assessment, together with structure-controlled distractors for a more rigorous evaluation. Third, we extend the optimization strategy to a two-stage framework consist- ing of difficulty-driven supervised fine-tuning with distilled rationales, followed by reinforcement learning guided by a novel Scale-aware Reasoning Structure Reward. Finally, we broaden the evaluation through additional model comparisons, ablation analyses, qualitative studies, and transfer experiments on conventional single-image pathology VQA. Together, these extensions establish a more comprehensive benchmark and training framework for advancing cross-scale reasoning in pathological image analysis. I. RELATED WORK A. Vision-Language Models for Pathology The rapid development of general-purpose VLMs, such as LLaVA [23], Qwen2.5-VL [24], Qwen3-VL [25], and InternVL3.5 [26], has substantially improved multimodal un- derstanding and visual reasoning. Medical VLMs, including LLaVA-Med [27], Lingshu [28], and HuatuoGPT [29], fur- ther adapt these capabilities to clinical images and biomed- ical question answering. Building on these developments, pathology-specific VLMs have emerged to support morpholog- ical interpretation, diagnostic support, and pathology-oriented reasoning. Patch-level models, such as Quilt-LLaVA [3], CLOVER [2], PathAsst [30], PathLens [6], and Patho-R1 [4], have demonstrated strong performance in pathology VQA, local feature recognition, and diagnostic reasoning. WSI- level models, including WSI-LLaVA [8], SlideChat [7], and PathReasoner-R1 [5], extend this scope by aggregating in- formation from multiple sampled regions to support slide- level understanding and reasoning over spatially distributed evidence. Despite these advances, explicit supervision for linking clinically corresponding evidence across magnification levels remains limited. B. Pathology Visual Question Answering Benchmarks Pathology VQA benchmarks have progressively expanded in scale, task diversity, and clinical relevance. Early patch- based datasets, such as PathVQA [14], established the evalua- tion of natural-language question answering over histopathol- ogy images, while later benchmarks including Quilt-VQA [3] and PathMMU [15] introduced larger and more diverse ques- tion sets for assessing morphological recognition and di- agnostic knowledge. WSI-level benchmarks, such as WSI- VQA [17], SlideBench [7], and WSI-Bench [8], further ex- tend evaluation to whole-slide understanding and slide-level question answering. Although these benchmarks cover a wide range of pathology tasks, dedicated evaluation of whether models can associate, verify, and integrate corresponding evidence across successive magnifications remains limited. Moreover, recent studies have shown that VQA benchmarks may be susceptible to linguistic priors, weak distractors, and annotation artifacts [18], [19], [21], allowing models to exploit non-visual shortcuts. These gaps motivate PathScale-VQA, which evaluates cross-scale reasoning over clinically linked multi-magnification diagnostic paths with shortcut-resistant benchmark construction. I. METHODOLOGY A. Cross-scale Diagnostic Path Construction To instantiate the principle of cross-scale evidence, we define the diagnostic path as the foundational evidence unit of our dataset. As illustrated in Fig. 2(A), a diagnostic path consists of a set of 10×, 40×, 200× ROIs from the same WSI. This structure reflects the clinical workflow of pathology diagnosis: low-power views provide architectural context, intermediate views reveal local tissue organization, and high-power views confirm cellular morphology. We con- struct these diagnostic paths from H&E-stained WSIs obtained from The Cancer Genome Atlas (TCGA) [31]. Clinical ex- perts are therefore involved at the earliest stage of our data construction, rather than serving only in post hoc validation as in prior benchmarks [7], [8], [15]. For each WSI, three junior pathologists independently inspect the slide to construct diagnostic paths. Each annotator first selects a diagnostically informative 10× ROI and then identifies clinically relevant subregions at 40× and 200×, forming a progressive zoom- in trajectory across magnifications. These candidate paths are subsequently reviewed by a senior pathologist to ensure that each trajectory corresponds to a clinically meaningful diag- nostic process. For each validated path, pathologists provide two levels of annotation: scale-specific captions describing the pathological features in each ROI and the rationale for each zoom-in decision. We then use GPT-5.2 [32] to draft 4IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2026 AB CD Cross-scale Visual Grounding Supervised Fine-Tuning A B D C 10x 200x 40x glandular arch. with irregular outlines... Text-only Adversarial Inference Scale-specific Feature Decomposition D B A C Advantage Estimation R1 R2 R3 Rn Cross-scale Semantic Reasoning VQA Generation Correspondence Confirmation Explanation Diagnosis Localization Scale Anticipation Scale Bridging Origin Grounding Zoom-Out Zoom-In Cross-scale Task Formulation Dim-specific visual-dependent Semantic VQA Constraint Refinement Path Registry Construction 10x caption 40x caption 200x caption Cross-scale description Same-path Anchoring Teacher Model Reasoning Distillation Long Reasoning Traces Pretrained Pathology VLM SFT on Hard VQA with Distilled CoT Student Model ... Hard VQA set (low empirical success rate) Difficulty Filtering Semantic Reasoning VQA Normalized & Structured Chain-of-Thought PathScale- SFT Stage-1: CoT Distillation & Supervised Fine-TuningStage-2: Reinforcement Learning Optimization Reward Function Policy Model (SFT init) Reference Model PathScale-R1 Group Sampling Policy Gradient Update Low description similarity Within-slide sampling for Spatial Alignment Cross-slide sampling for Structural Consistency Low mutual IoU Low IoU w/ GT Sufficient tissue Cross-scale VQA constraints QA candidates Leakage Analysis Can be answered correctly without visual input? Which aspects of the question or options permit text-only infer? Tissue Processing Question Anchor Question Anchor Cross-scale Evidence Construction KL div Accuracy Reward Format Reward Scale-aware Reasoning Structure Reward (A) (C) (B) Yes No ... Adversarial Text-only Screening Loop Structure-controlled Distractor Sampling Shortcut-Resistant VQA Curation Fig. 2. Overview of the proposed cross-scale benchmark construction and model optimization framework. (A) Expert-verified diagnostic paths link clinically relevant 10×, 40×, and 200× ROIs from the same WSI, providing scale-specific captions and cross-scale evidence anchors. (B) From these paths, we construct cross-scale semantic reasoning and visual grounding tasks, with adversarial text-only screening to reduce language shortcuts and structure-controlled distractor sampling to reduce superficial visual shortcuts. (C) PathScale-R1 is optimized by difficulty- driven reasoning distillation followed by reinforcement learning with accuracy, format, and scale-aware reasoning structure rewards. a cross-scale description from the ROI captions and zoom- in rationales. This description is refined and confirmed by pathologists to ensure that it explicitly connects low-power architecture, intermediate tissue organization, and high-power cellular morphology. The resulting expert-verified diagnostic paths and annotations serve as high-quality evidence anchors for subsequent VQA curation. B. Task Formulation & Shortcut-resistant VQA curation Based on the expert-verified diagnostic paths, we construct two complementary VQA schemes: cross-scale semantic rea- soning and cross-scale visual grounding, as illustrated in Fig. 1(C). The former evaluates whether models can integrate pathological evidence across magnifications, while the latter assesses visual correspondence between ROIs along a diagnos- tic trajectory. Because textual and image answer options are vulnerable to different shortcuts, we apply separate curation strategies to reduce text-only leakage and superficial visual cues, respectively. 1) Cross-scale Semantic Reasoning: In the semantic rea- soning scheme, each sample is formulated as a multiple-choice VQA with textual answer options and a clinically linked set of ROIs from one diagnostic path. Each question is designed to require evidence from at least two magnifications, so the model must connect observations across scales instead of relying on single-view recognition. To cover the major reasoning steps involved in pathological diagnosis, we define five semantic task types. Correspondence requires the model to determine whether findings observed at different magnifications repre- sent the same pathological process. Confirmation requires the model to determine whether higher-magnification evidence supports or contradicts a hypothesis formed at a lower mag- nification. Localization requires the model to identify where diagnostically relevant evidence appears within the cross- scale image set. Explanation requires the model to connect observations across magnifications to justify a conclusion. Diagnosis requires the model to determine the final clinical interpretation supported by the combined cross-scale evidence. Together, these tasks provide a comprehensive evaluation of cross-scale diagnostic reasoning, from evidence association to final interpretation. Adversarial Text-only Screening Loop. Since semantic rea- soning uses textual answer choices, the central curation chal- lenge is to prevent the correct answer from being inferred from linguistic cues alone. To reduce this risk, we propose an adversarial generate-screen-revise loop for data curation (Fig. 2B). Specifically, we first decompose each diagnostic path into scale-specific feature sets, separating findings visible at 10×, 40×, and 200×. Candidate questions are then gener- ated from these decomposed feature sets under task-specific constraints tailored to each reasoning dimension. After gener- ation, we subsequently remove all images and present only the question and textual options to strong closed-source LLMs, Gemini 3 Pro [33] and Qwen3-Max [34]. If either model predicts the correct answer, the sample is flagged as potentially solvable without visual evidence. For each flagged sample, we analyze the adversary’s rationale to identify possible sources of leakage, including overly specific correct options, implausible distractors, imbalanced option granularity, answer- position bias, and diagnostic priors embedded in the question wording. We then revise the question, options, or generation PHAN et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING5 85 93 90 70 90 30 30 30 30 40 70 70 70 80 45 Esophagus (402) Thyroid (491) MedVLThinker-7B OctoMed-7B QoQ-Med-VL-7B Lingshu-7B LLaVA-Med-7B HuatuoGPT-7B InternVL3.5-8B Qwen3-VL-8B MiMo-VL-7B Qwen2.5-VL-7B Patho-R1-7B Quilt-LLaVA CLOVER HealthGPT-8B PathScale-R1 (A) Brain (996) Bone (344) Bladder (501) Stomach (504) Kidney (514) Head & Neck (603) Colorectum (196) Uterus (1407) Breast (2317) Hepatobiliary (1041) Lung (499) Organ Distribution (VQAs) 10,373 VQAs 1,368 cross-scale diagnostic paths 373 WSIs PathScale-VQA 14 organs (B) Cross scale Single scale Lymphoid (558) Correspondence Confirmation Localization Explanation Diagnosis Zoom-In Zoom-Out Origin Grounding Scale Bridging Scale Anticipation PubMed SocialPath EduContent Atlas PathCLS Fig. 3.Dataset statistics and benchmark performance. (A) PathScale-VQA component statistics and organ distribution. (B) Task-wise performance of representative VLMs across single-scale and proposed cross-scale VQA benchmark. constraints accordingly, and re-evaluate the revised sample under the same setting. This iterative process turns semantic VQA construction from one-pass generation into leakage- aware curation. Finally, the revised VQAs are reviewed by pathologists to verify that the designated answer is clinically correct, the required evidence is observable in the correspond- ing ROIs, and the question depends on findings from at least two magnifications. This process yields a semantic VQA set with reduced text-only solvability and stronger reliance on cross-scale visual evidence. 2) Cross-scale Visual Grounding: While semantic reasoning evaluates cross-scale interpretation through textual answer choices, visual grounding evaluates whether a model can re- cover the visual relationships that support such interpretation. We formulate this scheme as image-option MCQs derived from expert-verified diagnostic paths. Each sample contains an anchor and multiple visual alternatives, with the correct option defined by the corresponding cross-scale trajectory. By replacing textual options with visual alternatives, this scheme reduces reliance on diagnostic label priors and directly probes whether models can align, localize, and complete cross- scale visual trajectories. To capture the core visual operations involved in practical pathology interpretation, we define five task types: Zoom-In Correspondence and Zoom-Out Corre- spondence identify corresponding higher- and lower-power views, respectively; Origin Grounding localizes a crop within its parent ROI; Scale Bridging recovers the missing interme- diate view between 10× and 200×; and Scale Anticipation selects the high-power appearance most consistent with lower- power evidence. Together, these tasks complement semantic reasoning by assessing the visual alignment and trajectory- level consistency underlying cross-scale diagnosis. Structure-controlled Distractor Sampling. Although image- option questions can avoid text-only leakage, they remain vulnerable to a different form of shortcut. Random distrac- tors may be distinguishable by staining variation, scanner- specific appearance, tissue coverage, background content, im- age sharpness, or crop artifacts. To address this risk, we adopt a structure-controlled distractor sampling pipeline based on the expert-verified diagnostic paths (Fig. 2B). First, we organize annotated ROIs into a path-indexed visual registry that records the diagnostic path ID per WSI, cross-scale ROI links, and parent-child relationships. For each grounding question, the anchor and correct option are drawn from the same diagnostic path, ensuring that correctness corresponds to an expert-validated trajectory. Distractors are then sampled according to the task objective. For spatial alignment tasks, including Zoom-In, Zoom-Out Correspondence, and Origin Grounding, distractors are drawn from other diagnostic paths within the same WSI. This controls slide-level appearance factors while requiring the model to identify the correct spatial correspondence or parent-child mapping. Candidate distractors are further constrained to valid tissue regions based on tissue segmentation and low IoU with the ground-truth region. For trajectory completion tasks, including Scale Bridging and Scale Anticipation, the focus is on morphological consistency across scales. Distractors are therefore sampled from different WSIs since the objective is to select the view that best completes or follows a diagnostic trajectory. We also filter distractors whose image descriptions are highly compatible with the anchor, reducing ambiguity while preserving visually challenging alternatives. This design thus provides a visual- focused VQA set to evaluate whether VLMs can correctly match corresponding regions and features across scales. C. Cross-scale Semantic Reasoning Optimization Building on the curated cross-scale semantic reasoning set, we develop PathScale-R1 through the two-stage optimization framework illustrated in Fig. 2(C). The first stage performs supervised fine-tuning (SFT) on difficult questions with dis- tilled reasoning traces, providing supervision for cross-scale evidence synthesis. The second stage applies reinforcement learning (RL) to improve answer correctness while encourag- ing the response to follow the visual evidence structure. 1) SFT via Difficulty-DrivenReasoningDistillation: Difficulty-based Question Selection. Motivated by recent findings that model complex reasoning ability can emerge through small, high-quality examples acting as ”cognitive templates” [35], [36], we adopt a difficulty-driven filtering strategy. By identifying questions with consistently low success rates across repeated inference, this strategy helps concentrate supervision on cross-scale reasoning patterns that remain challenging for the base model. Formally, for each sample (I,q,O,a ∗ ) ∈ D sem , where I denotes the input image set, q is the question, O is the answer option set, 6IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2026 and a ∗ is the expert-verified answer, we query the baseline pathological VLM, Patho-R1 [4], for K attempts. We use the number of correct predictions as an empirical estimate of model solvability and define the hard subset as: D hard =(I,q,O,a ∗ )∈D sem | c base (I,q,O) < τK, (1) where c base (I,q,O) denotes the number of correct responses among K attempts, and τ ∈ [0, 1] is a difficulty threshold controlling the maximum empirical success rate allowed for inclusion. This criterion retains samples on which the baseline model has a low success rate, thereby concentrating super- vised reasoning distillation on examples that better expose the model’s current weaknesses. Reasoning Construction via Distillation. For each retained hard question, we use a strong open-source teacher model, Qwen3.5-397B [37], to generate candidate reasoning traces. Since raw teacher outputs often contain redundant self- verification and repeated speculation, they are unsuitable for directly training a compact student VLM. Directly imitating such responses may dilute the core diagnostic logic and induce unstable reasoning. To address this, we normalize the teacher- generated traces into a compact and standardized format. The normalized response preserves the final answer and essential diagnostic logic, while removing redundancy and unsupported speculation. Based on manual inspection of retained high- quality traces, we organized the rationale into structured chain- of-thought rationales that include the task focus, image-wise visual analysis, key visual evidence, confirmed and uncertain findings, option evaluation, cross-scale conclusion, and final answer. This process yields a high-quality reasoning corpus that logically analyzes cross-scale evidence to reach diagnos- tically meaningful conclusions. SFT with Distilled Rationales. Using this hard subset and the normalized teacher rationales, we fine-tune the student model via the autoregressive supervised fine-tuning objective. The input consists of the cross-scale image set, question, and options, while the target is the distilled rationale followed by the final answer. This stage structurally teaches the model to analyze visual evidence across scales to reach a correct conclusion. 2) RL Optimization with Scale-aware Reasoning Structure Reward: After SFT establishes a structured cross-scale reason- ing pattern, we further optimize the model with RL to improve answer correctness and reinforce reasoning across scales. Reward Design. We formulate the reward as outcome- verifiable optimization under scale-aware reasoning con- straints. For a generated response o, expert-verified answer a ∗ , and cross-scale image set I =I 1 ,...,I N , the reward is defined as: R(o) = λ acc R acc (o,a ∗ )+λ srsr R srsr (o,I)+λ fmt R fmt (o), (2) where R acc (o,a ∗ ) = 1 [ˆa(o) = a ∗ ] ensures answer correct- ness, R fmt (o) = 1 [o∈F], ensures format compliance with <think> and <answer> tags. Although SFT provides image-wise reasoning demonstrations, the model may still collapse a multi-image input into a brief global description, without explicitly distinguishing the evidence contributed by individual images. This behavior is undesirable for cross-scale pathological reasoning, where different magnifications provide complementary morphological information. We therefore in- troduce a Scale-aware Reasoning Structure Reward (R srsr ) to reinforce explicit image-wise analysis before cross-scale synthesis. Let o ana denote the analysis segment extracted from the generated response. We apply a deterministic structure ex- tractor φ(·) to identify the ordered sequence of image indices explicitly referenced in the analysis: φ(o ana ) = (m 1 ,...,m L ). Here, the sequence is determined by the first occurrence of each image reference in the generated analysis. Given the expected scale sequence S(I) = (1,...,N), the structure reward is defined as R srsr (o,I) = ( 1, φ(o ana ) =S(I), 0, otherwise. (3) This reward is assigned only when the generated reasoning accounts for all provided visual contexts in their intended scale order. As a result, R srsr penalizes hallucinated visual references, or inconsistent cross-scale organization, thereby encouraging more complete and organized image-wise anal- ysis. Policy Optimization. We optimize the model using stan- dard Group Relative Policy Optimization (GRPO) [38] to maximize the proposed reward. The final optimized model, PathScale-R1, combines difficulty-driven reasoning distillation with scale-aware reward optimization to improve cross-scale pathological evidence integration. IV. EXPERIMENTS AND RESULTS A. Experiment Setting Datasets. Our curated PathScale-VQA contains 10,373 cross- scale VQA samples from 1,368 expert-verified diagnostic paths extracted from 373 TCGA WSIs, spanning fourteen organ sites (Fig. 3A). The semantic reasoning subset contains 6,030 samples, of which 3,680 are used for training and valida- tion, and 2,350 are used for testing. All splits are constructed at the patient-wise WSI level. The visual grounding subset contains 4,343 samples and is used exclusively for evaluation. We also evaluate on PathMMU [15], a public single-scale pathology VQA benchmark with 10,387 questions, to assess transferability beyond the proposed cross-scale setting. Implementation Details. PathScale-R1 uses Patho-R1-7B [4] as its backbone, which was trained on large-scale pathology data but focuses only on single-scale analysis. For difficulty filtering, we set K= 16, τ = 0.25. During SFT, we freeze the vision tower and train for 5 epochs with a learning rate of 1e-4. For RL, the policy is initialized from the SFT model and optimized with GRPO for 700 steps with 16 sampled responses per prompt. The reward weights are set to λ acc = 0.75, λ srsr = 0.20, and λ fmt = 0.05. To provide a comprehensive and fair comparison, we compare with recent general-purpose VLMs [24]–[26], [39], medical VLMs [27]–[29], [40]–[43], and pathology-specific VLMs [2]–[4] with parameter scales comparable to our 7B model. All models are evaluated using the same prompt template, image order, decoding settings, and answer-parsing procedure. Accuracy (%) is used as the evaluation metric for all VQA tasks. PHAN et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING7 TABLE I OVERALL RESULTS OF MODELS ON CROSS-SCALE SEMANTIC REASONING. THE BEST PERFORMANCE IN EACH COLUMN IS HIGHLIGHTED IN BOLD, THE SECOND-BEST IS UNDERLINED . Model OverallCorr.Conf.Loca.Expl.Diag. (2350)(470)(470)(470)(470)(470) General VLM Qwen2.5-VL-7B [24]55.9647.2348.3077.0239.5767.66 MiMo-VL-7B [39]47.4536.1743.6254.4738.3064.68 Qwen3-VL-8B [25]59.8348.3061.2875.7436.1777.66 InternVL3.5-8B [26]64.1353.1969.5778.9444.4774.47 Medical VLM LLaVA-Med-7B [27]22.2127.2315.9623.1921.9122.77 HuatuoGPT-7B [29] 50.1335.7438.5174.6839.1562.55 QoQ-Med-VL-7B [40] 52.9846.3841.4974.6833.8368.51 Lingshu-7B [28]56.2639.5754.8971.9141.0673.83 MedVLThinker-7B [41] 50.7241.2840.8571.2839.3660.85 OctoMed-7B [42] 59.6236.6064.0480.8543.8372.77 HealthGPT-8B [43]57.2830.2158.9478.5144.4774.26 Pathological VLM Quilt-LLaVA [3]38.8539.7932.1348.5138.0935.74 CLOVER [2]56.7738.5160.4378.7239.7966.38 Patho-R1 [4]50.2131.4934.6867.2344.0473.62 PathScale-R183.3280.8591.9188.7265.7489.36 B. Performance on Cross-Scale Semantic Reasoning We first evaluate the models on cross-scale semantic reason- ing, which directly measures their ability to integrate patholog- ical evidence across magnifications. Table I reports the overall accuracy and task-wise performance across five reasoning dimensions. PathScale-R1 achieves the best overall accuracy of 83.32%, substantially outperforming the evaluated general- purpose, medical, and pathology-specific VLMs. Compared with the strongest general-domain model, InternVL3.5-8B, PathScale-R1 improves the overall accuracy by 19.19%. It also exceeds the strongest medical baseline, OctoMed-7B, and the strongest pathology-specific baseline, CLOVER, by 23.70% and 26.55%, respectively. Relative to the Patho-R1 back- bone, PathScale-R1 improves overall accuracy from 50.21% to 83.32%, with consistent gains across all five reasoning dimensions. The particularly large gains on Correspondence and Confirmation indicate an improved ability to associate observations across magnifications and determine whether evidence at one scale supports findings observed at another. The results demonstrate that difficulty-driven reasoning distil- lation and scale-aware RL strengthen the cross-scale evidence integration targeted by the proposed framework. C. Transfer Capability to Single-Scale VQA Although the proposed framework is developed for multi- image cross-scale reasoning, its learned diagnostic reasoning strategy may also benefit conventional pathology tasks involv- ing a single image. We therefore examine PathScale-R1 on the PathMMU [15] validation and test sets, with the results reported in Table I. PathScale-R1 achieves the best overall performance across all evaluation splits, obtaining 63.6% accuracy on the validation set, 70.1% on the test-tiny set, and 64.5% on the full test set. Compared with the Patho-R1 backbone, these results correspond to improvements of 1.3%, 5.3%, and 1.9%, respectively. The gains are also consistent across the different data sources in PathMMU, outperform- ing strong general-domain, medical-domain, and pathology- domain baselines. These results suggest that learning from diagnostic paths appears to improve how the model identifies and organizes pathological evidence, which remains useful even when only one scale is available. Thus, cross-scale reasoning supervision not only improves multi-image evidence integration but also strengthens pathology understanding on conventional single-scale VQA. D. Effectiveness of Shortcut-Resistant VQA Curation To verify the effectiveness of our shortcut-resistant curation, we evaluate whether the PathScale-VQA-Semantic questions genuinely require visual and cross-scale evidence through a progressive image removal setting. To reduce model-specific bias and avoid floor effects from weak models, we select the five strongest-performing models: PathScale-R1, InternVL3.5- 8B, Qwen3-VL-8B, OctoMed-7B, and HealthGPT-8B. For each question, we retain the original question and answer options while progressively removing views from the diag- nostic path. The results are averaged across the five models under each condition. As shown in Table I, average accuracy decreases from 64.83% with all images to 55.35% after re- moving one image and 47.52% when only one image remains. This consistent degradation indicates that the questions depend on complementary evidence across magnifications rather than isolated visual recognition. In the text-only setting, accuracy further falls to 27.79%. These results show that textual priors alone are insufficient to reproduce full-image performance, confirming strong dependence on cross-scale visual evidence. E. Ablation Analysis of the Key Components We conduct ablations to examine the contributions of difficulty-driven SFT and the scale-aware reasoning structure reward in RL optimization. All ablation experiments are conducted on the same backbone, Patho-R1-7B, to ensure a controlled comparison. The results are presented in Table IV. Effectiveness of Difficulty-driven SFT. Using distilled ra- tionales from difficult samples substantially improves cross- scale semantic reasoning, increasing the average accuracy from 50.21% to 71.36%. The largest gains occur in Corre- spondence and Confirmation, which improve from 31.49% to 58.51% and from 34.68% to 83.62%, respectively. This indicates that targeted supervision on repeatedly failed cases efficiently addresses the backbone’s weaknesses in associating and verifying evidence across scales. The distilled rationales provide compact cognitive templates for associating patho- logical evidence across magnifications, enabling the model to better follow structured diagnostic logic. Effectiveness of RL Optimization. Applying RL after SFT fur- ther improves the average accuracy from 71.36% to 80.97%. The largest gains are observed in Correspondence and Expla- nation, with improvements of 21.06% and 12.55%, respec- tively. This suggests that SFT establishes the initial cross- scale reasoning pattern, whereas subsequent outcome-based optimization improves the model’s ability to apply this pattern 8IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2026 TABLE I OVERALL RESULTS OF MODELS ON THE PATHMMU. THE BEST PERFORMANCE IS HIGHLIGHTED IN BOLD, THE SECOND-BEST IS UNDERLINED . Model OverallPubMedSocialPathEduContentAtlasPathCLS ValTest-TinyTestTest-TinyTestTest-TinyTestTest-TinyTestTest-TinyTestTest-TinyTest (710)(1156)(8521)(281)(2787)(235)(1620)(255)(1683)(208)(799)(177)(1632) General VLM Qwen2.5-VL-7B [24]44.348.244.453.450.550.047.759.648.948.646.829.428.0 MiMo-VL-7B [39]36.937.635.848.840.245.441.342.041.936.140.315.815.5 Qwen3-VL-8B [25]49.453.952.963.756.858.354.453.754.951.056.742.941.9 InternVL3.5-8B [26]59.063.959.770.563.962.065.269.461.669.267.240.840.6 Medical VLM LLaVA-Med-7B [27]17.522.222.724.625.123.421.125.123.617.824.520.319.2 HuatuoGPT-7B [29]43.243.841.249.147.752.346.052.946.144.746.919.819.2 QoQ-Med-VL-7B [40]45.947.446.455.551.351.949.552.250.644.746.132.834.4 Lingshu-7B [28]51.956.053.760.557.661.658.169.458.758.162.530.431.6 MedVLThinker-7B [41]45.849.946.156.250.354.650.255.749.751.453.231.627.3 OctoMed-7B [42]53.059.455.869.461.562.556.965.156.956.760.143.543.5 HealthGPT-8B [43]53.863.460.069.864.668.564.668.262.070.267.040.141.7 Pathological VLM Quilt-LLaVA [3]33.832.331.930.334.226.834.635.733.341.833.327.124.0 CLOVER [2]53.160.656.668.059.867.660.568.662.262.064.136.736.6 Patho-R1 [4]62.364.862.668.764.963.965.771.066.178.473.541.842.7 PathScale-R163.670.164.573.068.073.667.178.068.080.375.245.844.1 TABLE I IMAGE-ABLATION ANALYSIS ON THE CROSS-SCALE SEMANTIC REASONING TEST SET. VALUES IN PARENTHESES DENOTE ACCURACY DROPS RELATIVE TO USING ALL IMAGES. Corresp.Confirm.Localiz.Explana.Diagnos.AVG Full images49.8369.1580.5546.9477.7064.83 Drop 1 image 44.2853.3669.4043.0366.6755.35 (↓5.55)(↓15.79)(↓11.15)(↓3.91)(↓11.03)(↓9.48) Drop 2 images 33.4246.9360.2938.6858.2847.52 (↓16.41)(↓22.22)(↓20.26)(↓8.26)(↓19.42)(↓17.31) Drop 3 images (Text-only) 17.4027.8634.4820.8938.3027.79 (↓32.43)(↓41.29)(↓46.07)(↓26.05)(↓39.40)(↓37.04) consistently and produce diagnostically correct conclusions. This supports our two-stage design: SFT first teaches the model how to reason over cross-scale evidence, while RL further aligns the generated reasoning and final answer with the desired diagnostic outcome. Effectiveness of the Scale-aware Reasoning Structure Re- ward. Introducing R srsr further increases the average accuracy from 80.97% to 83.32%. Improvements are observed across all task categories, ranging from 1.28% in Correspondence to 2.78% in Confirmation, with gains of 2.76, 2.34, and 2.55% in Localization, Explanation, and Diagnosis, respec- tively. The broadly distributed gains indicate that the reward provides a general benefit across different cross-scale tasks. This improvement indicates that explicitly rewarding the use of cross-scale evidence complements answer-level supervision and promotes more reliable evidence organization. F. Qualitative Comparison To qualitatively evaluate the reasoning behavior of our approach, we present representative cases compar- ing PathScale-R1 with the second-best-performing model, InternVL3.5-8B, and its Patho-R1 backbone. As shown in TABLE IV COMPONENT ABLATION OF DIFFICULTY-DRIVEN SFT AND RL OPTIMIZATION. THE BEST PERFORMANCE IN EACH COLUMN IS BOLD. Corresp.Confirm.Localiz.Explana.Diagnos.AVG Baseline31.4934.6867.2344.0473.6250.21 + SFT58.5183.6281.7050.8582.1371.36 + SFT + RL (w/oR srsr )79.5789.1385.9663.4086.8180.97 + SFT + RL (w/R srsr )80.8591.9188.7265.7489.3683.32 Fig. 4, the cross-scale diagnosis case requires weighing ev- idence from all three magnifications before reaching the fi- nal interpretation. Both baselines produce fluent pathological descriptions, yet each fixes on a locally salient impression and concludes a combined hepatocellular–cholangiocarcinoma without verifying whether ductal differentiation is actually supported at high power. PathScale-R1 instead proceeds through the images in low-to-high magnification order, records what each scale contributes, and explicitly notes the absence of convincing ductal differentiation before rejecting that option. This example illustrates that our model produces a more coherent cross-scale interpretation by relating evidence across views rather than relying on isolated findings or diagnostic priors. Fig. 5 further shows that this structured reasoning behavior transfers to single-scale VQA. PathScale-R1 correctly identifies the irregular and complex glandular architecture and provides an option-level justification grounded in the observed morphology, whereas the baselines rely on coarse visual im- pressions. Taken together, these cases support the quantitative results, demonstrating stronger cross-scale evidence integra- tion and more reliable morphology-based reasoning in both cross-scale and single-scale settings. G. Cross-Scale Visual Grounding: Performance and Remaining Challenges We further evaluate models on the evaluation-only cross- scale visual grounding, which tests whether a model can PHAN et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING9 <think> Let's analyze the images step by step. Image 1 shows a large mass with a distinct boundary from the surrounding tissue. Image 2 shows a region with dense cellular area and some areas of necrosis or hemorrhage. Image 3 shows cells with pleomorphic nuclei and prominent nucleoli, typical of malignant cells. Option A is less likely because hepatocellular adenomas typically do not show such a prominent desmoplastic stroma. Option B: While there is a desmoplastic stroma, the presence of hepatocytic differentiation makes this less likely. Option C fits well with the images. Option D: While there is a fibrotic background, the presence of duct-forming areas suggests more than just hepatocellular carcinoma. <answer> C </answer> INCORRECT <think> The image shows a tumor with two distinct components: hepatocytic differentiation and ductal differentiation. The background stroma is fibrotic, with a mixed inflammatory infiltrate. These features align with combined hepatocellular– cholangiocarcinoma, which requires both hepatocytic and cholangiocytic elements. Option C correctly identifies this mixed morphology. Option A lacks ductal component, Option B lacks hepatocytic features, and Option D describes a background but not the biphenotypic tumor. </think> <answer> C </answer> <think> The user wants me to analyze pathology features shown in the given images. Image analysis: - Image 1: Low-power shows a large, nodular liver lesion with a thick fibrous capsule/cirrhotic- appearing background. The lesion is composed of broad trabeculae/nests of eosinophilic cells separated by fibrous septa, with hemorrhage and inflammatory infiltrates. - Image 2: Medium-power highlights a large dilated vascular space (likely a sinusoid/cavernous space) surrounded by a dense inflammatory infiltrate and fibrous stroma. The lining cells are large with abundant eosinophilic cytoplasm, supporting a malignant hepatocytic component. - Image 3: High-power shows large polygonal cells with abundant eosinophilic cytoplasm, vesicular nuclei, and prominent nucleoli. There are also smaller cells with darker nuclei, consistent with a mixed population. The cells are arranged in trabeculae/nests, supporting hepatocytic differentiation. The key visual claim being tested is whether the lesion is a hepatocellular carcinoma arising in a fibrotic/cirrhotic background versus [...] Option Evaluation: - Option A: Incorrect. Hepatocellular adenoma would lack malignant cytology and a prominent inflammatory infiltrate with fibrous septa. - Option B: Incorrect. Intrahepatic cholangiocarcinoma would show glandular/ductal structures and a prominent desmoplastic stroma, not a trabecular hepatocytic pattern. - Option C: Incorrect. Combined hepatocellular–cholangiocarcinoma requires both hepatocytic and cholangiocytic components; the images mainly show hepatocytic trabeculae without convincing ductal differentiation. - Option D: Correct. Hepatocellular carcinoma arising in a fibrotic/cirrhotic background fits the trabecular growth, malignant hepatocytic cytology, and prominent inflammatory infiltrate. Overall, the lesion is a malignant hepatocytic neoplasm with a cirrhotic background, best described as hepatocellular carcinoma arising in a fibrotic/cirrhotic liver. </think> <answer> D </answer> CORRECT INCORRECT InternVL3.5-8B Question: Based on the images provided, which option best matches the overall morphologic impression? (A) Hepatocellular adenoma with background bridging fibrosis and inflammation (B) Intrahepatic cholangiocarcinoma with a prominent desmoplastic stroma (C) Combined hepatocellular–cholangiocarcinoma with mixed hepatocytic and duct- forming areas (D) Hepatocellular carcinoma arising in a fibrotic/cirrhotic background Ground truth: (D) Task: Cross-scale Diagnosis Patho-R1-7B Scheme: Semantic Reasoning PathScale-R1 Fig. 4.Cross-scale semantic reasoning VQA case study. PathScale-R1 integrates architectural, stromal, and cellular evidence across magnifications and reaches the correct diagnosis, whereas the baseline models infer unsupported features and produce incorrect predictions. <think> The image shows a well-organized glandular architecture with uniformity in the shape and arrangement of the glands. This suggests a regular, repeating pattern. The other options suggest varying degrees of uniformity or organization, which do not match the observed pattern in the image. So the answer is D. </think> <answer> D </answer> Question: Based on the histological features observable in this image, what can be (A) The glandular architecture lacks uniformity with irregular shapes and arrangements (B) The glands show slight variation but maintain basic structural integrity (C) The glandular architecture exhibits a high degree of uniformity and organization (D) The glandular architecture demonstrates a regular, repeating pattern Ground truth: (A) <think>The user wants me to analyze the histological features of the image. Image analysis: The H&E section shows a complex papillary architecture with branching fronds and gland-like structures. The glands are irregular in shape and size, with complex branching and papillary infoldings. The epithelial lining appears stratified and crowded, and the stroma is scant and fibrovascular. There is no clear, organized glandular pattern; the architecture is disorganized and complex. Option Evaluation: - Option A: Correct. The glands are irregular in shape and arrangement, with complex branching and papillary structures, indicating a lack of uniformity. - Option B: Incorrect. The glands show marked variation in size and shape rather than slight variation. - Option C: Incorrect. The architecture is not highly uniform; it is complex and disorganized. - Option D: The glands do not form a regular, repeating pattern; they are irregular and complex. Overall, the histology shows a complex papillary architecture with irregular glandular shapes and arrangements, consistent with a lack of uniformity.</think> <answer> A </answer> <think> The image shows a papillary growth pattern with uniform glandular architecture. According to the provided image, the glandular architecture in this case is highly uniform and organized, which directly supports option C. The absence of irregular shapes or arrangements (as in option A) and the lack of slight variation (option B) or a repeating pattern (option D) are inconsistent with the described uniformity. The reasoning aligns with the observation of preserved glandular organization and the absence of atypical architectural features. </think> <answer> C </answer> inferred about the uniformity of the glandular architecture? INCORRECT InternVL3.5-8B INCORRECT Patho-R1-7B CORRECT PathScale-R1 Fig. 5. Single-scale pathology VQA case study. PathScale-R1 identifies the key morphological evidence and selects the correct answer, whereas baseline models rely on coarse visual impressions and produce incorrect answers. identify spatial and morphological correspondence across di- agnostic trajectories. As shown in Fig. 6(A), PathScale-R1 achieves the highest overall grounding performance among the evaluated models and improves over its Patho-R1 back- bone without any grounding-specific training. Figure 6(B) illustrates a successful case study, in which PathScale-R1 identifies the intermediate view by matching stromal organi- zation and cellular morphology across magnifications. This suggests that semantic cross-scale optimization provides a certain degree of transfer to visual trajectory understanding. Nevertheless, grounding performance remains concentrated within a relatively low and narrow range across all general, medical, and pathology-specific VLMs, despite substantially larger differences in semantic reasoning accuracy. Cross-scale visual grounding can thus be positioned as a challenging and clinically relevant capability gap that should be further addressed in future pathology VLM development. H. Limitations and Future Work Despite the strong cross-scale reasoning performance, several directions remain for further development. First, PathScale-VQA is constructed from TCGA H&E slides and predefined 10×, 40×, and 200× trajectories; extending it to multi-center data and more flexible magnification paths would further broaden its coverage. Second, the current MCQ formulation enables controlled evaluation but does not fully represent open-ended diagnosis or interactive WSI navigation. Moreover, our optimization focuses primarily on semantic reasoning, while the visual grounding results indicate that fine-grained spatial correspondence remains an important ca- pability to improve. Future work will therefore explore joint semantic-grounding optimization and validation in broader, clinically realistic settings. V. CONCLUSION In this work, we introduced a cross-scale pathology VQA paradigm that better aligns with the multi-magnification work- flow of pathological diagnosis. We constructed high-quality PathScale-VQA from pathologist-verified diagnostic paths and designed complementary semantic reasoning and visual grounding tasks with shortcut-resistant curation. We further developed PathScale-R1, combining difficulty-driven reason- ing distillation with scale-aware RL optimization. Experiments 10IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2026 A B C D <think>The user wants me to identify the mid-magnification image that best connects the low- and high-magnification pathology views. Image analysis: - Image 1: The low-magnification image shows a heterogeneous lesion with broad fibrous stromal areas, inflammatory cells, and irregular tissue spaces. - Image 2: The high-magnification image shows atypical cells closely associated with dense collagenous stroma. The key visual feature is the transition from the fibrous, cellular architecture at low magnification to the detailed cellular–stromal interface at high magnification. Option Evaluation: - Option A: Incorrect. The field is predominantly eosinophilic and necrotic-appearing and does not preserve the fibrous, cellular architecture of the images. - Option B: Correct. It shows dense fibrous stroma with interspersed cellular areas and inflammation, providing a plausible intermediate view between the broad low-magnification architecture and the detailed high- magnification cellular–stromal appearance. - Option C: Incorrect. It contains cohesive polygonal cells with a lymphoid aggregate, representing a different tissue organization. - Option D: Incorrect. It shows prominent keratinizing epithelial nests and keratin pearls, which are not present in the target views. Overall, B best preserves the morphological continuity between the low- and high- magnification images. </think> <answer> B </answer> CORRECT HuatuoGPT-7B MedVLThinker-7B QoQ-Med-VL-7B Lingshu-7B HealthGPT-8B OctoMed-7B Quilt-LLaVA Patho-R1 CLOVER MiMo-VL-7B Qwen2.5-VL-7B Qwen3-VL-8B InternVL3.5-8B Q: Given the low- and high-mag images, which candidate mid- magnification image best represents an intermediate view connecting the two in a zoom-in sequence? Ground truth: (B) Task: PathScale-R1 LowHigh (B) (A) Cross-scale Semantic Reasoning (%) Cross-scale Visual Gr ounding (%) PathScale-R1 General VLMs Medical VLMs Pathology VLMs Scale-Bridging Better direction Fig. 6. Cross-scale visual grounding analysis. (A) Comparison of cross-scale semantic reasoning and visual grounding performance across models. PathScale-R1 achieves the strongest semantic reasoning performance and the highest grounding accuracy, while grounding performance remains limited across all evaluated models. (B) A representative case in which PathScale-R1 identifies the intermediate-magnification view by preserving morphological continuity between the low- and high-magnification images. show that PathScale-R1 substantially improves cross-scale semantic reasoning and transfers effectively to single-scale pathology VQA, demonstrating the practical value of cross- scale supervision. At the same time, our visual grounding benchmark reveals that fine-grained cross-scale correspon- dence remains challenging for current VLMs. Overall, our study highlights cross-scale reasoning as an important direc- tion for clinically aligned pathology VLMs and provides a comprehensive benchmark and training framework for advanc- ing future pathological AI systems. REFERENCES [1] R. J. Chen et al., “Towards a general-purpose foundation model for computational pathology,” Nature medicine, p. 850–862, 2024. [2] K. Chen et al., “Cost-effective instruction learning for pathology vision and language analysis,” Nature Computational Science, 2025. [3] M. S. Seyfioglu et al., “Quilt-llava: Visual instruction tuning by ex- tracting localized narratives from open-source histopathology videos,” in CVPR, 2024, p. 13 183–13 192. [4] W. Zhang et al., “Patho-r1: A multimodal reinforcement learning-based pathology expert reasoner,” arXiv preprint arXiv:2505.11404, 2025. [5] S. Jiang et al., “Pathreasoner-r1: Instilling structured reasoning into pathology vision-language model via knowledge-guided policy opti- mization,” arXiv preprint arXiv:2601.21617, 2026. [6] Z. Zhu et al., “Pathlens: A lightweight multimodal reasoner for in-depth pathology insights,” Knowledge-Based Systems, p. 116261, 2026. [7] Y. Chen et al., “Slidechat: A large vision-language assistant for whole- slide pathology image understanding,” in CVPR, 2025, p. 5134–5143. [8] Y. Liang et al., “Wsi-llava: A multimodal large language model for whole slide image,” in ICCV, 2025, p. 22 718–22 727. [9] E. Abels et al., “Computational pathology definitions, best practices, and recommendations for regulatory guidance: a white paper from the digital pathology association,” The Journal of pathology, 2019. [10] Z. Zhang et al., “Pathologist-level interpretable whole-slide cancer diagnosis with deep learning,” Nature Machine Intelligence, 2019. [11] N. Hashimoto et al., “Multi-scale domain-adversarial multiple-instance cnn for cancer subtype classification with unannotated histopathological images,” in CVPR, 2020, p. 3852–3861. [12] D. Entenberg et al., “Time-lapsed, large-volume, high-resolution intrav- ital imaging for tissue-wide analysis of single cell dynamics,” Methods, vol. 128, p. 65–77, 2017. [13] Y. Chong et al., “A stepwise approach to fine needle aspiration cytology of lymph nodes,” Journal of Pathology and Translational Medicine, vol. 57, no. 4, p. 196–207, 2023. [14] X. He et al., “Pathvqa: 30000+ questions for medical visual question answering,” arXiv preprint arXiv:2003.10286, 2020. [15] Y. Sun, H. Wu, C. Zhu et al., “Pathmmu: A massive multimodal expert- level benchmark for understanding and reasoning in pathology,” in ECCV. Springer, 2024. [16] Y. Sun, H. Wu, C. Zhu, Y. Si et al., “Pathbench: Advancing the bench- mark of large multimodal models for pathology image understanding at patch and whole slide level,” IEEE TMI, 2025. [17] P. Chen et al., “Wsi-vqa: Interpreting whole slide images by generative visual question answering,” in ECCV. Springer, 2025, p. 401–417. [18] D. Lee et al., “Breaking the visual shortcuts in multimodal knowledge- based visual question answering,” preprint arXiv:2511.22843, 2025. [19] R. Shrestha et al., “A negative case analysis of visual grounding methods for VQA,” in ACL, 2020, p. 8172–8181. [20] A. Agrawal et al., “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in CVPR, 2018, p. 4971–4980. [21] M. Asadi et al., “Mirage: The illusion of visual understanding,” arXiv preprint arXiv:2603.21687, 2026. [22] C. Phan et al., “Enhancing pathological vlms with cross-scale reason- ing,” arXiv preprint arXiv:2606.17412, 2026. [23] H. Liu et al., “Visual instruction tuning,” in NeurIPS, 2023. [24] S. Bai, K. Chen, X. Liu et al., “Qwen2. 5-vl technical report,” arXiv preprint arXiv:2502.13923, 2025. [25] “Qwen3-vl technical report,” arXiv preprint arXiv:2511.21631, 2025. [26] W. Wang et al., “Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,” arXiv preprint arXiv:2508.18265, 2025. [27] C. Li et al., “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,” NeurIPS, vol. 36, p. 28 541–28 564, 2023. [28] W. Xu et al., “Lingshu: A generalist foundation model for uni- fied multimodal medical understanding and reasoning,” arXiv preprint arXiv:2506.07044, 2025. [29] H. Zhang et al., “Huatuogpt, towards taming language models to be a doctor,” arXiv preprint arXiv:2305.15075, 2023. [30] Y. Sun et al., “Pathasst: A generative foundation ai assistant towards artificial general intelligence of pathology,” in AAAI, 2024. [31] A. Colaprico et al., “Tcgabiolinks: an r/bioconductor package for integrative analysis of tcga data,” Nucleic acids research, 2016. [32] A. Singh et al., “Openai gpt-5 system card,” arXiv preprint arXiv:2601.03267, 2025. [33] Gemini Team, “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023. [34] Qwen Team, “Qwen3-max: Just scale it,” September 2025. [35] Y. Ye et al., “Limo: Less is more for reasoning,” 2025. [Online]. Available: https://arxiv.org/abs/2502.03387 [36] C. Zhou et al., “Lima: Less is more for alignment,” NeurIPS, vol. 36, p. 55 006–55 021, 2023. [37] Qwen Team, “Qwen3.5: Towards native multimodal agents,” February 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.5 [38] Z. Shao et al., “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024. [39] Xiaomi Team, “Mimo-vl technical report,” 2025. [Online]. Available: https://arxiv.org/abs/2506.03569 [40] D. Dai et al., “Qoq-med: Building multimodal clinical foundation models with domain-aware grpo training,” NeurIPS, vol. 38, p. 37 406– 37 453, 2026. [41] X. Huang et al., “Medvlthinker: Simple baselines for multimodal med- ical reasoning,” arXiv preprint arXiv:2508.02669, 2025. [42] T. Ossowski et al., “Octomed: Data recipes for state-of-the-art multi- modal medical reasoning,” arXiv preprint arXiv:2511.23269, 2025. [43] T. Lin et al., “Healthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation,” 2025. [Online]. Available: https://arxiv.org/abs/2502.09838