Paper deep dive
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Jinjian Wu, Jiaqi Tang, Wei Wei, Yingying Yan, Jianmin Chen, Botong Geng, Lei Zhang, Qifeng Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/18/2026, 11:01:39 AM
Summary
The paper introduces IQA-T1, a tool-based visual evidence reasoning framework for Image Quality Assessment (IQA) that augments Multimodal Large Language Models (MLLMs) with explicit perceptual observations. It addresses the semantic bias of existing MLLM-based methods by autonomously invoking specialized analysis tools to generate structured visual evidence (e.g., noise residual maps, frequency spectra). The authors also present Q-Tool, a dataset of 11k multimodal reasoning chains grounded in this evidence, and demonstrate that IQA-T1 achieves state-of-the-art performance on seven IQA benchmarks.
Entities (11)
Relation Signals (9)
IQA-T1 → trainson → Q-Tool
confidence 96% · To support this paradigm, we construct Q-Tool... Extensive experiments... show that IQA-T1 achieves the best overall performance
Jinjian Wu → affiliatedwith → Northwestern Polytechnical University
confidence 95% · Jinjian Wu... School of Computer Science, Northwestern Polytechnical University
Jiaqi Tang → affiliatedwith → The Hong Kong University of Science and Technology
confidence 95% · Jiaqi Tang... The Hong Kong University of Science and Technology
IQA-T1 → uses → MLLM
confidence 95% · IQA-T1 is a framework that augments MLLM reasoning with explicit perceptual observations.
IQA-T1 → trainedwith → Supervised Fine-Tuning
confidence 94% · We address this through a two-stage training pipeline. First, supervised fine-tuning (SFT) instills foundational behaviors
IQA-T1 → trainedwith → Reinforcement Learning
confidence 94% · Second, reinforcement learning (RL) with a carefully designed reward function optimizes the tool invocation policy.
Q-Tool → contains → Fourier Magnitude Spectrum
confidence 90% · Q-Tool... containing 11k multimodal reasoning chains grounded in tool-generated evidence... and frequency spectra
Q-Tool → contains →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We propose IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations. During inference, the model autonomously invokes specialized analysis tools to generate structured visual evidence, such as noise residual maps, gradient statistics, and frequency spectra, which are progressively integrated into the reasoning process. To support this paradigm, we construct Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence. Extensive experiments on seven IQA benchmarks show that IQA-T1 achieves the best overall performance across datasets while producing interpretable and evidence-grounded quality assessments. Code and dataset are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.12375v1
- Canonical: https://arxiv.org/abs/2607.12375v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
43,826 characters extracted from source content.
Expand or collapse full text
11institutetext: School of Computer Science, Northwestern Polytechnical University, Xi’an, China 11email: wujinjian@mail.nwpu.edu.cn 11email: weiweinwpu@nwpu.edu.cn 22institutetext: The Hong Kong University of Science and Technology, Hong Kong SAR, China 22email: jtang092@connect.ust.hk 22email: cqf@ust.hk IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment Jinjian Wu * Jiaqi Tang* Wei Wei🖂 Yingying Yan Jianmin Chen Botong Geng Lei Zhang Qifeng Chen🖂 Abstract Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We propose IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations. During inference, the model autonomously invokes specialized analysis tools to generate structured visual evidence, such as noise residual maps, gradient statistics, and frequency spectra, which are progressively integrated into the reasoning process. To support this paradigm, we construct Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence. Extensive experiments on seven IQA benchmarks show that IQA-T1 achieves the best overall performance across datasets while producing interpretable and evidence-grounded quality assessments. Code and dataset are available at https://github.com/zibuyu-02/IQA-T1. †footnotetext: * Equal contribution. 🖂 Correspondence to: Wei Wei and Qifeng Chen. 1 Introduction Image Quality Assessment (IQA) aims to predict perceptual image quality in alignment with human subjective judgment. In image restoration [li2025mair, lin2025jarvisir, jiang2025survey, tang2026robustu1], enhancement [an2024hfm, yan2025hvi, ma2025bilevel, Tang_2024_CVPR], compression [he2022elic, liu2023learned, yang2023lossy], and various downstream vision applications [tang2025intelligent], image quality assessment not only plays an important role in algorithm performance but also directly affects the usability and robustness of vision systems in open-world environments. With the rapid development of multimodal large language models (MLLMs) [yang2025qwen3, touvron2023llama, li2023blip, achiam2023gpt, tang2024hawk], researchers have begun to explore leveraging their reasoning capabilities to shift IQA from score regression toward language-driven frameworks. Although this transition promises enhanced interpretability through textual rationales and improved generalization via semantic priors, we identify a fundamental limitation shared by existing MLLM-based IQA methods: their quality assessments rely on internally represented, semantically biased visual features, rather than explicit, verifiable evidence of low-level degradation. This limitation manifests across two predominant reasoning modalities, as illustrated in Fig. 1. (1) Text-only Reasoning. These methods [li2025q, wu2025visualquality] generate structured quality descriptions to interpret and predict image quality. However, the reasoning process relies on the MLLM’s internal representations and pre-trained semantic priors. (2) Region-based Visual Reasoning. To address the lack of visual grounding in text-only reasoning, recent methods [liang2026zoom, li2026q] introduce cropped or zoomed regions of the original image during multi-step reasoning, aiming to guide the model to focus on potentially degraded areas. However, these regions are essentially unstructured pixel patches that lack explicit association with specific perceptual attributes, and thus cannot form visual evidence to support the reasoning process. Figure 1: Motivation of our visual evidence reasoning framework. Existing MLLM-based IQA methods typically follow two reasoning paradigms: text-only reasoning, which is prone to semantic bias, and region-based visual reasoning using cropped image patches that lack explicit perceptual interpretation. In contrast, our framework introduces a tool-based visual evidence paradigm that generates structured visual evidence through specialized analysis tools, enabling perceptually grounded and evidence-referenced quality reasoning. To address this fundamental gap, we introduce IQA-T1, a novel framework that instantiates a visual evidence reasoning paradigm for IQA. The core insight is to augment the MLLM’s reasoning process with explicit and structured visual evidence generated by external specialized tools. During inference, the model autonomously invokes tools from a predefined library, where each tool is designed to isolate a specific perceptual attribute, to generate intermediate visual representations such as noise residual maps, gradient magnitude distributions, and Fourier magnitude spectra. These evidence images are then incorporated into the reasoning chain as additional visual observations. This process transforms quality assessment from a purely semantic inference task into a visually grounded and evidence-driven procedure. Realizing this paradigm requires equipping the model with two complementary capabilities: understanding how to invoke tools and learning when to do so strategically. We address this through a two-stage training pipeline. First, supervised fine-tuning (SFT) instills foundational behaviors: structured reasoning, standardized tool invocation syntax, and the ability to associate visual evidence with quality judgments. Second, reinforcement learning (RL) with a carefully designed reward function optimizes the tool invocation policy. The overall training reward consists of a format reward, a scoring reward, a tool usage reward, and a tool repetition penalty, which together improve scoring accuracy while suppressing redundant or improper tool calls. Since no existing IQA dataset supports visual evidence reasoning, we construct Q-Tool, containing 11k multimodal reasoning chains grounded in tool-generated visual evidence. The dataset is built through a three-stage pipeline: tool library construction, visual evidence reasoning chain generation, and reasoning chain verification. We conduct comprehensive evaluations on seven diverse IQA benchmarks. Experimental results show that IQA-T1 achieves the best overall performance, with an average PLCC / SRCC of 0.7950.795 / 0.7840.784, surpassing traditional deep learning-based approaches and recent MLLM-based methods. Further ablation studies demonstrate that incorporating structured visual evidence significantly improves the stability and accuracy of quality prediction, validating the effectiveness of the visual evidence reasoning framework in enhancing both assessment reliability and interpretability. In summary, our contributions are threefold: • We propose IQA-T1, a tool-based visual evidence reasoning framework for image quality assessment. It enables MLLMs to autonomously invoke specialized tools, dynamically constructing a structured evidence base that grounds quality assessment in explicit perceptual representations. • We introduce Q-Tool, containing over 11k high-quality multimodal reasoning chains with standardized tool invocation formats and rigorous verification. • We conduct extensive evaluations of IQA-T1 on multiple IQA benchmarks, demonstrating its SOTA performance compared to both traditional IQA approaches and recent MLLM-based baselines. 2 Related Work 2.1 Image Quality Assessment Image Quality Assessment (IQA) aims to predict perceptual image quality in alignment with human subjective judgments. Traditional methods model quality degradation and output scores under either full-reference (FR-IQA) [wang2004image] or no-reference (NR-IQA) [ma2017learning, mittal2012no, mittal2012making] settings. With the rapid advancement of multimodal large language models (MLLMs) [yang2025qwen3, touvron2023llama, li2023blip, achiam2023gpt], researchers have leveraged their cross-modal understanding to extend IQA from discriminative regression toward more generalizable, language-driven frameworks. Existing MLLM-based IQA research can be taxonomized into three evolving paradigms: (1) Score-Oriented Paradigm. Early MLLM-based methods formulate quality prediction as classification, distribution regression, or contrastive learning to improve numerical accuracy [wang2023exploring, wu2023q, zhu2024adaptive, you2025teaching, liu2024dog]. While these approaches achieve strong correlation with human judgments, their outputs are limited to numerical values, offering no interpretability regarding why an image receives a particular score. This opacity limits their utility in applications requiring explainable decisions. (2) Description-Oriented Paradigm. To enhance interpretability, description-oriented methods [chen2024grounding, wu2024q, you2024depicting, you2025enhancing] generate structured or fine-grained natural language quality analyses. By producing textual descriptions of perceived distortions, these models provide insight into their quality judgments. However, they typically rely on supervised fine-tuning with text descriptions constructed from synthetic distortions, which fail to capture the diversity and complexity of real-world degradation patterns. Moreover, their training objectives prioritize textual coherence over numerical accuracy, often resulting in unstable or imprecise quality scores. (3) Textual Reasoning Paradigm. Recent studies have introduced explicit reasoning processes, forming a "reason-then-score" joint optimization framework. Exemplified by Q-Insight [li2025q] and VisualQuality-R1 [wu2025visualquality], these methods generate step-by-step textual rationales before predicting quality scores, often incorporating reinforcement learning to align reasoning with scoring. Despite this advancement, their quality judgments remain grounded in the MLLM’s internal representations, which are inherently biased toward high-level semantics. As argued in our introduction, this semantic bias leaves these methods ill-equipped to capture low-level degradation cues, and their purely textual reasoning lacks verifiable visual anchors, rendering the generated rationales potentially disconnected from actual visual evidence. 2.2 Thinking with Images The recognition that text-only reasoning lacks visual grounding has motivated a broader exploration of paradigms where images actively participate in the reasoning process. Termed "Thinking with Images" [su2025thinking], this emerging direction transforms visual information from passive input into a dynamic cognitive workspace. Recent work can be categorized into three progressive levels [su2025thinking]: (1) tool-driven visual exploration, where models invoke external tools such as cropping and zooming to acquire fine-grained visual evidence [wang2025visuothink, zheng2025deepeyes, wang2025pixel, zhang2026cmmcot, liu2025visual]; (2) programmatic visual manipulation, where models generate executable code to perform structured edits on images [fu2025refocus, liu2025visualagentic, wang2025vrag]; and (3) intrinsic visual imagination, where models generate new visual content as intermediate reasoning steps [chen2025blip3, li2025imagine, zhao2025cot, jiang2025t2i]. These explorations mark the evolution from "thinking about images" to "thinking with images." Thinking with Images for IQA. Applying this paradigm to IQA represents a natural progression toward addressing the lack of visual evidence in text-only reasoning. Methods such as Zoom-IQA [liang2026zoom] and Q-Probe [li2026q] iteratively crop or zoom into evidence regions during inference, guiding the model to focus on locally degraded areas. While these approaches represent important steps, they remain fundamentally limited: the introduced regions are unstructured pixel patches that still rely on the model’s semantic interpretation. They do not generate explicit, structured visual evidence that directly highlights specific perceptual attributes such as noise characteristics, frequency anomalies, or structural coherence. Consequently, they fail to bridge the representational gap between semantic reasoning and low-level degradation, which motivates our proposed IQA-T1 framework based on tool-generated visual evidence. 3 Dataset Figure 2: Overview of the Q-Tool dataset construction pipeline, consisting of three key stages: (a) Perceptual tool library construction for visual evidence generation; (b) Visual evidence-grounded reasoning chain synthesis guided by human-curated priors; (c) Multi-level verification of reasoning chains. Existing MLLM-based IQA methods suffer from a fundamental limitation: their quality assessments rely on semantically biased internal representations rather than explicit, verifiable evidence of low-level degradation. Addressing this gap requires a dataset that explicitly models the relationship between perceptual degradation and structured visual evidence within a reasoning framework. However, no such dataset currently exists. To enable the training of visual evidence reasoning models, we introduce Q-Tool —the first dataset specifically designed for tool-based visual evidence reasoning in IQA. It comprises 11k high-quality multimodal reasoning chains, each integrating explicit tool-generated visual evidence with human-aligned textual rationales. The construction of Q-Tool is guided by three requirements: (1) the dataset must include explicit visual evidence that captures low-level perceptual attributes; (2) the reasoning chains must systematically integrate this evidence with textual analysis; and (3) the overall structure must support supervised fine-tuning of MLLMs to learn evidence-grounded reasoning behaviors. To meet these requirements, we design a systematic three-stage construction pipeline, illustrated in Fig. 2. The following sections describe each stage in detail. 3.1 Tool Library Construction The foundation of visual evidence reasoning is the ability to generate explicit representations of low-level perceptual attributes. However, as argued in Sec. 1, MLLM internal representations are inherently biased toward high-level semantics and fail to capture degradation cues such as noise, blur, or frequency anomalies. To overcome this limitation, we construct a library of specialized perceptual tools, each designed to isolate and visualize a specific quality-relevant attribute. Design rationale. We identify seven perceptual attributes critical for IQA, including structure, sharpness, noise, artifacts, luminance, color, and naturalness. Based on these attributes, we design 15 corresponding types of visual evidence. Detailed descriptions are provided in the supplementary material. Each tool TkT_k maps an input image I to a structured visual representation ek=ψ(Tk(I))e_k=ψ(T_k(I)) that highlights degradation patterns invisible to standard MLLM representations. For example, the NoiseResidualMap isolates sensor or compression noise by subtracting a denoised version from the original; the FourierMagnitudeSpectrum reveals periodic artifacts through frequency-domain analysis; and GradientOrientationCoherenceMap exposes structural inconsistencies caused by blur or over-smoothing. Implementation considerations. To ensure stable training and inference, all tools share three properties. First, they adopt unified invocation interfaces, allowing the MLLM to call any tool using a consistent syntax. Second, outputs are moderately downsampled to balance spatial structure preservation with visual token efficiency. Third, all tools are deterministic, guaranteeing identical outputs across data construction, supervised fine-tuning, and reinforcement learning stages—a critical requirement for consistent policy learning. This tool library provides the perceptual foundation for subsequent reasoning chain construction, enabling the explicit representation of degradation cues that semantic MLLM representations inherently lack. 3.2 Visual Evidence Reasoning Chain Construction With the tool library established, we next construct multimodal reasoning chains that systematically integrate visual evidence into quality assessment. A reasoning chain consists of (i) explicit tool invocations at appropriate analytical stages, (i) interpretation of the generated visual evidence, and (i) a final quality judgment logically derived from the accumulated evidence. Challenge of naive generation. Directly prompting large language models to generate such chains autonomously leads to semantic hallucinations, logical discontinuities, and improper tool usage. The resulting chains lack the structured, verifiable relationship between evidence and judgment required for effective supervision. Human-curated reasoning templates. To address this, we first develop human-curated reasoning templates that encode expert knowledge about perceptual analysis. These templates define: the logical decomposition of quality assessment into stages (e.g., "examine structure," "analyze noise," "assess color fidelity"); the mapping from perceptual attributes to appropriate tools; and the causal relationships between observed evidence and final quality judgment. These templates serve as semantic constraints, ensuring that generated chains follow human-aligned perceptual logic. Constrained generation with GPT-4o. Under these constraints, we employ GPT-4o to generate complete multimodal reasoning chains. For each image, the model receives: the image itself, the corresponding reasoning template, and access to the tool library. At each designated stage, the model inserts explicit tool invocations, and the resulting visual evidence is integrated into subsequent analysis. This process transforms purely textual reasoning into an evidence-grounded procedure where each inference step relies on observable visual signals. By combining human perceptual priors with controlled model generation, we produce reasoning chains that are both logically structured and grounded in verifiable evidence. 3.3 Reasoning Chain Verification The reliability of Q-Tool as a training resource depends on the quality of its constituent reasoning chains. We therefore conduct systematic multi-level verification to ensure consistent alignment among tool usage, visual evidence, and reasoning logic. Verification protocol. Each generated chain is examined along three dimensions. (1) Tool appropriateness: Are tool invocations aligned with the intended reasoning stage and analytical purpose? We verify that tools are neither redundant nor improperly used. (2) Evidence grounding: Is the generated visual evidence correctly referenced, interpreted, and integrated into the textual reasoning? We ensure that each evidence image serves a genuine analytical function rather than appearing as a decorative element. (3) Logical consistency: Is the final quality judgment sufficiently supported by the preceding evidence accumulation? We reject chains where conclusions lack adequate evidentiary support or exhibit reasoning discontinuities. Quality refinement. Chains failing any verification criterion are removed. A subset of borderline cases undergoes manual refinement to correct minor inconsistencies while preserving the overall structure. This rigorous verification process yields a final dataset of 11k high-quality reasoning chains, each exhibiting: (i) appropriate tool utilization, (i) coherent integration of visual evidence, and (i) logically sound quality judgments. The resulting Q-Tool dataset provides the necessary supervision for learning evidence-grounded reasoning behaviors in the subsequent two-stage training pipeline (Sec. 4.3 and Sec. 4.3), serving as the foundation for the visual evidence reasoning paradigm introduced in this work. 4 Methodology Figure 3: Overview of the proposed IQA-T1 framework. The model operates in two stages: (1) Supervised fine-tuning (SFT) on the Q-Tool dataset to learn evidence-grounded reasoning and tool invocation syntax; (2) Reinforcement learning (RL) with a tailored reward function to optimize the adaptive tool invocation policy. During inference, the model autonomously invokes tools to acquire visual evidence rE_r, which is integrated into the reasoning chain to ground quality assessment. In this section, we first formalize the problem of MLLM-based IQA and highlight the semantic bias inherent in standard visual representations (Sec. 4.1). We then introduce the proposed visual evidence reasoning framework (Sec. 4.2) and describe the two-stage training pipeline comprising supervised fine-tuning (Sec. 4.3) and reinforcement learning (Sec. 4.3) to equip the model with both foundational reasoning capabilities and an adaptive tool invocation policy. 4.1 Problem Formulation MLLM-based IQA as Sequence Generation. We formalize image quality assessment using a multimodal large language model (MLLM) as a conditional sequence generation problem. Let I∈ℝH×W×3I ^H× W× 3 denote the input image and P the textual prompt (e.g., “Assess the quality of this image”). The target output is a sequence Y=y1,y2,…,yTY=\y_1,y_2,…,y_T\ that includes both a reasoning chain and a final quality score s∈ℝs (typically normalized to [1,5][1,5]). A standard MLLM comprises a visual encoder ℰϕE_φ, a modality projector ℳM, and a large language model (LLM) θG_θ. The image is first encoded and projected into a sequence of visual tokens aligned with the textual embedding space: v=ℳ(ℰϕ(I))∈ℝL×d,z_v=M (E_φ(I) ) ^L× d, where L is the number of visual tokens and d the embedding dimension. The inference process then models the posterior distribution of the output sequence conditioned on the multimodal inputs: p(Y∣I,P)=∏t=1Tpθ(yt∣v,p,y<t),p(Y I,P)= _t=1^Tp_θ(y_t _v,h_p,y_<t), (1) with ph_p denoting the prompt embeddings and y<ty_<t the previously generated tokens. Semantic Bias in Visual Representations. A growing body of work [chen2026mitigating, li2025investigate] has shown that the visual representation vz_v produced by MLLMs is heavily biased toward high-level semantic content, while offering limited sensitivity to low-level perceptual degradations critical for IQA, such as noise, blur, or compression artifacts. Formally, consider an image I and its perceptually degraded counterpart IdI_d (e.g., with added noise). Although human observers would assign significantly different quality scores s(I)s(I) and s(Id)s(I_d), the corresponding visual representations often remain nearly identical: ‖v(I)−v(Id)‖2≈0,|s(I)−s(Id)|≫ϵ,\|z_v(I)-z_v(I_d)\|_2≈ 0, |s(I)-s(I_d)| ε, (2) where ϵε is a small tolerance. This observation reveals that vz_v alone is insufficient to capture the degradation cues necessary for accurate quality assessment, motivating the introduction of complementary perceptual evidence. 4.2 Visual Evidence Reasoning To compensate for the information loss in vz_v, we propose a framework that augments the MLLM’s reasoning process with explicit, structured visual evidence generated by external tools. Let =Tkk=1KT=\T_k\_k=1^K denote a library of specialized perceptual analysis tools, each mapping an image to a degradation-oriented feature space: Tk:ℝH×W×3→k,T_k:R^H× W× 3 _k, where kD_k is the output domain (e.g., a residual map, histogram, or spectrum). The set of all possible visual evidence for image I is defined as =ek∣ek=ψ(Tk(I)),k=1,…,K,E=\e_k e_k=ψ(T_k(I)),\;k=1,…,K\, (3) with ψ(⋅)ψ(·) projecting tool outputs into the LLM’s context space (e.g., downsampling and tokenization). During inference, the model does not access the entire set E upfront. Instead, it dynamically constructs a subset r⊆E_r by invoking relevant tools at appropriate reasoning steps. The generation process thus conditions on both the original visual tokens and the progressively acquired evidence: Y^=argmaxY∑t=1Tlogpθ(yt∣v,p,r,y<t). Y= *arg\,max_Y _t=1^T p_θ(y_t _v,h_p,E_r,y_<t). (4) The evidence rE_r serves as a dynamically acquired perceptual complement to vz_v, enabling the model to ground each reasoning step in verifiable, low-level visual cues. This formulation transforms IQA from a purely semantic inference task into an evidence-driven reasoning process. 4.3 Two-Stage Training Pipeline Realizing the visual evidence reasoning paradigm requires equipping the MLLM with two complementary capabilities: (i) understanding how to invoke tools and interpret the resulting evidence, and (i) learning when to invoke tools strategically to maximize assessment accuracy while avoiding redundancy. We address these via a two-stage pipeline: supervised fine-tuning (SFT) on the Q-Tool dataset establishes foundational reasoning behaviors, followed by reinforcement learning (RL) that optimizes the adaptive tool invocation policy. Stage I: Evidence-Grounded Reasoning Learning In this stage, we train the model on the Q-Tool dataset to learn structured, evidence-grounded reasoning chains. Each training sample consists of the original image I, a set of tool-generated visual evidence V (corresponding to the ground-truth invocations), and a textual reasoning chain T that integrates the evidence and culminates in a quality score sgts_gt. Crucially, the model is not required to generate the visual evidence itself, as the evidence is deterministically produced by the tools and provided as input. During SFT, we apply loss masking to the tokens representing the evidence images, as they are not prediction targets. The objective is to maximize the likelihood of the reasoning tokens conditioned on the complete context: ℒSFT=−∑t=1LlogPθ(yt∣y<t,I,V,T),L_SFT=- _t=1^L P_θ(y_t y_<t,I,V,T), (5) where yty_t denotes the t-th ground-truth token of the reasoning sequence (including the final score) and L is its length. Optimizing Eq. (5) instills three essential behaviors: (i) the ability to follow structured reasoning templates, (i) correct syntax for tool invocation, and (i) the capacity to associate visual evidence with quality judgments. This stage initializes a policy πSFT _SFT that serves as the foundation for subsequent RL fine-tuning. Stage I: Adaptive Tool Invocation Policy Optimization While SFT establishes correct reasoning patterns, it does not explicitly optimize the strategy of when to invoke which tool. To learn an adaptive invocation policy, we employ reinforcement learning with a reward function designed to balance prediction accuracy, evidence efficiency, and reasoning stability. Reinforcement Learning Setup. We adopt Group Relative Policy Optimization (GRPO) [guo2025deepseek], a variant of PPO that estimates advantages from within-group comparisons, eliminating the need for a separate value network [tang2026lpo]. Given a prompt x, the current policy πθ _θ samples a group of n reasoning chains y1,…,yn\y_1,…,y_n\. The objective is: ℒGRPO(θ)=−x∼,yi∼πθ(⋅|x)[A^ilogπθ(yi|x)]+βDKL(πθ∥πref),L_GRPO(θ)=-E_x ,\,y_i _θ(·|x) [ A_i _θ(y_i|x) ]+β\,D_KL\! ( _θ\,\|\, _ref ), (6) where A^i=(Ri−μR)/σR A_i=(R_i- _R)/ _R is the normalized advantage within the group, πref _ref is the reference policy from Stage I, and β controls the KL penalty to prevent catastrophic forgetting. Reward Design. The reward function R(x,y)R(x,y) is a weighted combination of four components: (1) Format reward RfmtR_fmt: enforces that the output adheres to the expected structured format (e.g., contains tool invocation tags and a final score). It is 11 if the format is correct, 0 otherwise. (2) Scoring reward RscoreR_score: measures accuracy of the predicted quality score spreds_pred relative to the ground truth sgts_gt: Rscore=exp(−α|spred−sgt|),R_score= (-α\,|s_pred-s_gt| ), (7) with α>0α>0 controlling the sensitivity. (3) Tool usage reward RtoolR_tool: encourages efficient evidence acquisition by rewarding an appropriate number of tool invocations. Let N=|r|N=|E_r| be the number of distinct evidence items collected. Then Rtool(N)=0,N=0,1,1≤N≤M,1−γ(N−M)2,N>M,R_tool(N)= cases0,&N=0,\\[4.0pt] 1,&1≤ N≤ M,\\[4.0pt] 1-γ(N-M)^2,&N>M, cases (8) where M is a predefined maximum (set to 44 in our experiments) and γ controls the penalty for excessive invocations.This reward is intentionally designed as a weak constraint rather than a strict tool-level supervision signal, since real-world IQA images often involve mixed degradations and may admit multiple valid tool combinations. Therefore, RtoolR_tool encourages the model to acquire sufficient visual evidence while preserving flexibility in adaptive tool selection. (4) Repetition penalty RrepR_rep: prevents degenerate behavior where the model repeatedly invokes the same tool without gaining new information: Rrep=−1,∃k∈1,…,K:nk>1,0,otherwise,R_rep= cases-1,&∃\,k∈\1,…,K\:n_k>1,\\[2.84526pt] 0,&otherwise, cases (9) where nkn_k is the number of invocations of tool TkT_k. The overall reward is R=λfmtRfmt+λscoreRscore+λtoolRtool+λrepRrep,R= _fmtR_fmt+ _scoreR_score+ _toolR_tool+ _repR_rep, (10) with weights λ balancing the contributions (set empirically as λfmt=1 _fmt=1, λscore=5 _score=5, λtool=1 _tool=1, λrep=2 _rep=2). Through this reward formulation, the RL stage refines the tool invocation policy to maximize scoring accuracy while maintaining efficient and non-redundant use of perceptual evidence, ultimately yielding a model whose reasoning is both accurate and interpretable. 5 Experiments Table 1: PLCC/SRCC comparison on score regression tasks between our method and other competitive IQA methods. All methods except handcrafted ones are trained on the KonIQ dataset. AVG. denotes the average PLCC/SRCC across all evaluation datasets. The best result is bolded and the second best is underlined. Methods Wild Images Synthetic Distortion AI-Generated AVG. KonIQ [hosu2020koniq] SPAQ [fang2020perceptual] LiveW [ghadiyaram2015live] KADID [lin2019kadid] PIPAL [jinjin2020pipal] CSIQ [larson2010most] AGIQA [li2023agiqa] Handcrafted Methods NIQE [mittal2012making] 0.533 / 0.530 0.679 / 0.664 0.493 / 0.449 0.468 / 0.405 0.195 / 0.161 0.718 / 0.628 0.560 / 0.533 0.521 / 0.481 BRISQUE [mittal2012no] 0.225 / 0.226 0.490 / 0.406 0.361 / 0.313 0.429 / 0.356 0.267 / 0.232 0.740 / 0.556 0.541 / 0.497 0.436 / 0.369 Non-MLLM Methods NIMA [talebi2018nima] 0.896 / 0.859 0.838 / 0.856 0.814 / 0.771 0.532 / 0.535 0.390 / 0.399 0.695 / 0.649 0.715 / 0.654 0.697 / 0.675 HyperIQA [su2020blindly] 0.917 / 0.906 0.791 / 0.788 0.772 / 0.749 0.506 / 0.468 0.410 / 0.403 0.752 / 0.717 0.702 / 0.640 0.693 / 0.667 DBCNN [zhang2018blind] 0.884 / 0.875 0.812 / 0.806 0.773 / 0.755 0.497 / 0.484 0.384 / 0.381 0.586 / 0.572 0.730 / 0.641 0.667 / 0.645 MUSIQ [ke2021musiq] 0.924 / 0.929 0.868 / 0.863 0.789 / 0.830 0.575 / 0.556 0.431 / 0.431 0.771 / 0.710 0.722 / 0.630 0.726 / 0.707 ManIQA [yang2022maniqa] 0.849 / 0.834 0.768 / 0.758 0.849 / 0.832 0.499 / 0.465 0.457 / 0.452 0.623 / 0.627 0.723 / 0.636 0.681 / 0.658 MLLM-based Methods (w/o reasoning) CLIP-IQA+ [wang2023exploring] 0.909 / 0.895 0.866 / 0.864 0.832 / 0.805 0.653 / 0.654 0.427 / 0.419 0.772 / 0.719 0.736 / 0.685 0.742 / 0.720 C2Score [zhu2024adaptive] 0.923 / 0.910 0.867 / 0.860 0.786 / 0.772 0.500 / 0.453 0.354 / 0.342 0.735 / 0.705 0.777 / 0.671 0.706 / 0.673 Q-Align [wu2023q] 0.941 / 0.940 0.886 / 0.887 0.853 / 0.860 0.674 / 0.684 0.403 / 0.419 0.671 / 0.737 0.772 / 0.735 0.743 / 0.752 DeQA [you2025teaching] 0.953 / 0.941 0.895 / 0.896 0.892 / 0.879 0.694 / 0.687 0.472 / 0.478 0.787 / 0.744 0.809 / 0.729 0.786 / 0.765 MLLM-based Methods (w/ reasoning) Q-Insight [li2025q] 0.918 / 0.895 0.903 / 0.903 0.870 / 0.839 0.702 / 0.702 0.458 / 0.435 0.685 / 0.640 0.816 / 0.766 0.765 / 0.740 VisualQuality-R1 [wu2025visualquality] 0.910 / 0.896 0.889 / 0.892 0.856 / 0.827 0.703 / 0.712 0.451 / 0.441 0.768 / 0.707 0.817 / 0.760 0.771 / 0.748 Zoom-IQA [liang2026zoom] 0.938 / 0.922 0.902 / 0.900 0.887 / 0.870 0.701 / 0.700 0.468 / 0.465 0.797 / 0.754 0.816 / 0.765 0.787 / 0.768 IQA-T1 0.942 / 0.925 0.905 / 0.908 0.888 / 0.860 0.686 / 0.708 0.480 / 0.490 0.850 / 0.840 0.816 / 0.754 0.795 / 0.784 5.1 Experimental Setup Datasets and Metrics. To evaluate the robustness of IQA-T1 under diverse degradations, we follow prior works by training the model on KonIQ [hosu2020koniq] and evaluating its cross-distribution generalization on six IQA benchmarks. KonIQ [hosu2020koniq], SPAQ [fang2020perceptual], and LIVE-Wild [ghadiyaram2015live] represent real-world distortions; KADID [lin2019kadid] and CSIQ [larson2010most] contain synthetic distortions; and PIPAL [jinjin2020pipal] focuses on algorithm-induced degradations. We further include AGIQA [li2023agiqa] to evaluate AI-generated content. Following common IQA practice, performance is measured using PLCC and SRCC. Implementation Details. We build our framework upon Qwen3-VL-4B and adopt a two-stage training pipeline. In the first stage, we perform SFT on our Q-Tool datasets using cross-entropy loss. Training is conducted with a batch size of 1, 16 gradient accumulation steps, a learning rate of 1×10−51× 10^-5, and a warm-up ratio of 0.05 for two epochs. In the second stage, we apply GRPO to further align the model with perceptual quality preferences. The GRPO training initializes from the SFT checkpoint and uses a global batch size of 12 with a micro-batch size of 1 per GPU. The learning rate is set to 1×10−61× 10^-6, and each prompt generates N=8N=8 rollout samples. Training is performed on six NVIDIA RTX 4090 GPUs, and the KL-penalty coefficient is fixed at 1×10−21× 10^-2. Baseline Methods. We evaluate IQA-T1 on the image quality score regression task against four categories of baselines. (1) Traditional handcrafted IQA metrics, including NIQE [mittal2012making] and BRISQUE [mittal2012no]. (2) Deep learning-based IQA models such as NIMA [talebi2018nima], HyperIQA [su2020blindly], DBCNN [zhang2018blind], MUSIQ [ke2021musiq], and ManIQA [yang2022maniqa]. (3) MLLM-based score prediction methods, including CLIP-IQA+ [wang2023exploring], C2Score [zhu2024adaptive], Q-Align [wu2023q], and DeQA [you2025teaching]. (4) MLLM-based frameworks with reasoning, such as Q-Insight [li2025q], VisualQuality-R1 [wu2025visualquality], and Zoom-IQA [liang2026zoom]. Figure 4: Qualitative comparison of IQA-T1 with competing methods. We highlight: incorrect descriptions, and the visual evidence reasoning unique to our model. 5.2 Main Results Table 1 presents a comprehensive comparison of PLCC and SRCC across all methods and datasets. Our proposed IQA-T1 achieves the highest average PLCC (0.795) and SRCC (0.784) across the seven datasets. Comparison with MLLM-based methods without reasoning. Methods such as Q-Align [wu2023q] and DeQA [you2025teaching] focus exclusively on score regression and achieve strong performance, particularly on in-distribution datasets like KonIQ [hosu2020koniq]. However, they lack the ability to explain their judgments. IQA-T1 not only remains highly competitive in score accuracy but also generates interpretable, evidence-grounded reasoning chains, offering a unique combination of precision and transparency. Comparison with MLLM-based reasoning methods. Among methods that produce both scores and rationales, IQA-T1 exhibits strong cross-distribution generalization (Qualitative examples are shown in Fig. 4). It demonstrates robust performance on real-world images (e.g., SPAQ [fang2020perceptual]), establishes new highest results on challenging datasets for typical synthetic and algorithm-induced distortions (CSIQ [larson2010most] and PIPAL [jinjin2020pipal], respectively), and remains highly competitive on AI-generated content (AGIQA [li2023agiqa]). These results provide strong empirical evidence for the central thesis of this work: by augmenting MLLM reasoning with explicit, tool-generated visual evidence, IQA-T1 effectively reduces the semantic bias inherent in standard MLLM representations. 5.3 Ablation Studies and Analysis Figure 5: Qualitative comparison of IQA-T1 with competing methods. Table 2: Ablation studies of individual components measured by PLCC and SRCC, where all models are trained using the KonIQ dataset. Tool SFT toolR_ tool repR_ rep KonIQ [hosu2020koniq] SPAQ [fang2020perceptual] LiveW [ghadiyaram2015live] KADID [lin2019kadid] PIPAL [jinjin2020pipal] CSIQ [larson2010most] AGIQA [li2023agiqa] AVG. × ✓ × × 0.843 / 0.817 0.848 / 0.836 0.789 / 0.742 0.621 / 0.603 0.412 / 0.418 0.675 / 0.640 0.747 / 0.701 0.705 / 0.680 ✓ ✓ × × 0.915 / 0.901 0.891 / 0.890 0.826 / 0.774 0.660 / 0.705 0.464 / 0.455 0.768 / 0.730 0.770 / 0.712 0.756 / 0.738 ✓ ✓ ✓ × 0.933 / 0.917 0.903 / 0.904 0.881 / 0.858 0.684 / 0.704 0.479 / 0.481 0.841 / 0.835 0.811 / 0.739 0.790 / 0.777 ✓ ✓ × ✓ 0.935 / 0.919 0.902 / 0.902 0.874 / 0.840 0.689 / 0.708 0.481 / 0.485 0.838 / 0.832 0.804 / 0.742 0.789 / 0.775 ✓ ✓ ✓ ✓ 0.942 / 0.925 0.905 / 0.908 0.888 / 0.860 0.686 / 0.708 0.480 / 0.490 0.850 / 0.840 0.816 / 0.754 0.795 / 0.784 Table 3: Ablation study of different tool invocation strategies on the KonIQ dataset. Method PLCC SRCC Avg. Tools Avg. Time(s) Peak Mem.(GB) Qwen3-VL-4B 0.785 0.744 0.00 4.62 8.52 IQA-T1 (All Tools) 0.865 0.841 15.00 6.06 9.13 IQA-T1 (Fixed-K) 0.895 0.869 3.00 5.32 8.70 IQA-T1 (Random-K) 0.871 0.854 3.00 5.42 8.74 IQA-T1 (Dynamic) 0.942 0.925 2.34 5.09 8.67 Table 4: Generalization of IQA-T1 across different MLLM backbones. Results are reported as PLCC/SRCC. Method KonIQ [hosu2020koniq] SPAQ [fang2020perceptual] LiveW [ghadiyaram2015live] KADID [lin2019kadid] PIPAL [jinjin2020pipal] CSIQ [larson2010most] AGIQA [li2023agiqa] IQA-T1 (Qwen3-VL-2B) 0.929 / 0.910 0.892 / 0.898 0.851 / 0.800 0.654 / 0.700 0.471 / 0.470 0.812 / 0.773 0.792 / 0.723 IQA-T1 (InternVL3.5-4B) 0.898 / 0.897 0.881 / 0.875 0.842 / 0.821 0.701 / 0.719 0.443 / 0.437 0.749 / 0.733 0.822 / 0.774 IQA-T1 (Qwen3-VL-4B) 0.942 / 0.925 0.905 / 0.908 0.888 / 0.860 0.686 / 0.708 0.480 / 0.490 0.850 / 0.840 0.816 / 0.754 We provide a set of complementary analyses to better understand the effectiveness and generality of IQA-T1. Table 2 isolates the contribution of individual components, Table 3 compares different tool invocation strategies and their inference costs, and Table 4 evaluates the robustness of the proposed framework across different MLLM backbones. Effectiveness of Visual Evidence. We first examine the effect of tool-generated visual evidence by comparing two SFT variants: one where tool calls return only placeholders without actual evidence (Table 2 Row 1), and another where visual evidence is available during reasoning (Table 2 Row 2). Since both variants share identical training objectives and reasoning templates, the difference isolates the impact of visual evidence. The results show consistent improvements across all datasets, indicating that structured visual evidence effectively compensates for the low-level information loss in vz_v and mitigates the semantic bias discussed in Sec. 4.1. Effectiveness of Dynamic Tool Invocation. We compare dynamic tool invocation with all, fixed, and random tool selection strategies in Table 3. Compared with the base Qwen3-VL-4B, tool-generated evidence brings clear gains with moderate inference overhead. However, more evidence is not necessarily better, as irrelevant tools may introduce redundant visual tokens and distracting cues. In contrast, dynamic invocation adaptively selects relevant evidence for each image, achieving the best performance with fewer tool calls. The qualitative case study in Fig. 5 further shows that tools aligned with dominant degradations provide stronger support for quality judgment, indicating that IQA-T1 benefits from evidence reasoning rather than simple score fitting. Effectiveness of Reward Design. We next analyze the contribution of the RL-stage reward components. Starting from the full model, we ablate either the tool usage reward RtoolR_tool or the repetition penalty RrepR_rep in Table 2. Removing either component leads to degraded performance, indicating that both rewards are necessary for stable tool-based reasoning. Specifically, RtoolR_tool encourages the model to acquire sufficient perceptual evidence, while RrepR_rep suppresses redundant or repetitive tool calls. Their combination enables the model to balance evidence sufficiency and evidence diversity, leading to more reliable quality prediction. Generality across Backbones. To examine whether the proposed framework depends on a specific MLLM architecture, we instantiate IQA-T1 with different backbones, including Qwen3-VL-2B, InternVL3.5-4B, and Qwen3-VL-4B. As shown in Table 4, IQA-T1 maintains strong performance across different model scales and architectures. Although the absolute performance varies with backbone capacity and pretraining characteristics, the overall trend remains consistent across real-world, synthetic, algorithm-induced, and AI-generated image quality benchmarks. These results suggest that tool-based visual evidence reasoning is a general mechanism for enhancing MLLM-based IQA, rather than a backbone-specific design. 6 Conclusion In this work, we introduced IQA-T1, a tool-based visual evidence reasoning framework that grounds image quality assessment in explicit perceptual evidence rather than semantically biased internal representations. By enabling MLLMs to autonomously invoke specialized analysis tools, our approach transforms IQA from purely semantic inference into an evidence-driven reasoning process. Supported by the proposed Q-Tool dataset and a two-stage training strategy combining supervised learning and reinforcement learning, IQA-T1 achieves the best average performance across multiple IQA benchmarks while providing interpretable reasoning grounded in visual evidence. We hope this work encourages future research on integrating structured visual evidence into multimodal reasoning systems, advancing more reliable and interpretable perception in multimodal large language models. Acknowledgements This work was supported in part by the National Natural Science Foundation of China (No. 62472359, No. 62372379), in part by Xi’an’s Key Industrial Chain Core Technology Breakthrough Project: AI Core Technology Breakthrough under Grand 24ZDCYJSGG0003, and in part by the Hong Kong University of Science and Technology under Grant No. WEB26EG02. References