Paper deep dive
HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
Yucan Guo, Xiaohan Wang, Miao Su, Saiping Guan, Zhongni Hou, Jiajun Chai, Wei Lin, Guojun Yin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/25/2026, 7:31:06 AM
Summary
The paper introduces HiDiffTIR, a hierarchical difficulty-aware policy optimization framework for multi-turn Tool-Integrated Reasoning (TIR) in Large Language Models (LLMs). It addresses the limitation of existing Reinforcement Learning (RL) methods that assign uniform advantages to all correct tool calls. HiDiffTIR performs fine-grained credit assignment at both trajectory and turn levels by estimating difficulty using group-level statistics from standard RL rollouts, without additional supervision. Experiments on benchmarks like FTRL, ToolHop, and BFCL show that HiDiffTIR improves tool-use accuracy and reasoning performance over strong RL baselines like GRPO and MatchTIR.
Entities (10)
Relation Signals (8)
HiDiffTIR ā solves ā Tool-Integrated Reasoning
confidence 95% Ā· HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR.
HiDiffTIR ā employs ā credit assignment
confidence 93% Ā· HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels
HiDiffTIR ā uses ā Reinforcement Learning
confidence 92% Ā· Reinforcement Learning (RL) has become the dominant paradigm... HiDiffTIR... relies solely on group-level statistics derived from standard RL rollouts.
HiDiffTIR ā outperforms ā GRPO
confidence 90% Ā· HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines... GRPO
HiDiffTIR ā outperforms ā MatchTIR
confidence 88% Ā· HiDiffTIR consistently improves... over strong RL baselines... MatchTIR
HiDiffTIR ā evaluatedon ā FTRL
confidence 85% Ā· Extensive experiments on three tool-using benchmarks... FTRL
HiDiffTIR ā evaluatedon ā ToolHop
confidence 85% Ā· Extensive experiments on three tool-using benchmarks... ToolHop
HiDiffTIR ā evaluatedon ā BFCL
confidence 85% Ā· Extensive experiments on three tool-using benchmarks... BFCL
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, existing approaches typically assign uniform trajectory-level advantages and treat all correct tool calls equally, ignoring the varying difficulty and learning value across trajectories and reasoning steps. This can lead to imprecise learning signals that do not adequately distinguish between trivial and challenging tool-use patterns. To address this limitation, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR. HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels, enabling the policy to focus on more informative trajectories and harder reasoning steps. Notably, this fine-grained optimization is achieved without additional supervision, relying solely on group-level statistics derived from standard RL rollouts. Extensive experiments on three tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents.
Tags
Links
- Source: https://arxiv.org/abs/2608.21863v1
- Canonical: https://arxiv.org/abs/2608.21863v1
Trouble viewing inline? Open PDF directly ā
Full Text
79,147 characters extracted from source content.
Expand or collapse full text
HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning Yucan Guo Thanks: Co-first authors. Affiliation: State Key Laboratory of AI Safety Affiliation: Institute of Computing Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Xiaohan Wang Affiliation: Meituanguoyucan23z, guansaiping, jinxiaolong@ict.ac.cn Miao Su Affiliation: State Key Laboratory of AI Safety Affiliation: Institute of Computing Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Saiping Guan Thanks: Corresponding authors. Zhongni Hou Affiliation: Meituanguoyucan23z, guansaiping, jinxiaolong@ict.ac.cn Jiajun Chai Affiliation: Meituanguoyucan23z, guansaiping, jinxiaolong@ict.ac.cn Wei Lin Affiliation: Meituanguoyucan23z, guansaiping, jinxiaolong@ict.ac.cn Guojun Yin Affiliation: Meituanguoyucan23z, guansaiping, jinxiaolong@ict.ac.cn Xiaolong Jin Affiliation: State Key Laboratory of AI Safety Affiliation: Institute of Computing Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Jiafeng Guo Affiliation: State Key Laboratory of AI Safety Affiliation: Institute of Computing Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Xueqi Cheng Affiliation: State Key Laboratory of AI Safety Affiliation: Institute of Computing Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Abstract Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning (RL) has become the dominant paradigm for enabling this capability. However, existing approaches typically assign uniform trajectory-level advantages and treat all correct tool calls equally, ignoring the varying difficulty and learning value across trajectories and reasoning steps. This can lead to imprecise learning signals that do not adequately distinguish between trivial and challenging tool-use patterns. To address this limitation, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for multi-turn TIR. HiDiffTIR performs difficulty-aware credit assignment at both trajectory and turn levels, enabling the policy to focus on more informative trajectories and harder reasoning steps. Notably, this fine-grained optimization is achieved without additional supervision, relying solely on group-level statistics derived from standard RL rollouts. Extensive experiments on three tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy over strong RL baselines, highlighting the necessity of difficulty-aware credit assignment for effective policy optimization in tool-integrated LLM agents. 11 1 The code is publicly available at https://github.com/YucanGuo/HiDiffTIR. 1 Introduction Large Language Models (LLMs) have demonstrated strong reasoning capabilities across a wide range of tasks 23; 3; 41, leading to the emergence of LLM agents that can plan, act, and interact with external environments 10; 37; 21. A central capability of such agents is effective tool use, where they invoke external tools to solve diverse tasks 22; 39. Tool-Integrated Reasoning (TIR) formalizes this process by enabling LLM agents to iteratively interact with external tools 7; 27, such as APIs 26; 28, search engines 16; 12, and code interpreters 7; 18. This paradigm extends LLM agents beyond their static parametric knowledge and their known limitations in domains such as precise numerical computation 5, using specific tools to solve complex multi-turn tasks that are otherwise challenging for purely inner knowledge-based reasoning 20; 30. Reinforcement learning (RL) has emerged as the prevailing paradigm for training LLMs to perform TIR 27; 40. By employing RL algorithms like Group Relative Policy Optimization (GRPO) 32, models learn to interleave reasoning, tool invocation, and answer generation in an end-to-end manner. Most existing TIR methods adopt trajectory-level optimization, which has evolved through two primary stages. Early methods rely on outcome-based rewards 18; 4, assessing only the final correctness of a reasoning chain, and thus suffer from reward sparsity. To provide denser feedback, recent works have shifted toward turn-level verification 27; 43, where the correctness of each individual tool-call is evaluated and subsequently aggregated to form the trajectory-level reward. Despite this progress, these methods typically assign uniform advantages across all steps within a trajectory, implicitly treating each tool call as equally informative for learning. This coarse-grained credit assignment can obscure the contribution of individual reasoning steps, especially in multi-turn settings where errors accumulate and only a subset of decisions are critical to success. Figure 1: Comparison of credit assignment manners in TIR. Existing methods treat all correct tool calls uniformly, while HiDiffTIR performs difficulty-aware credit assignment. To address this limitation, more recent works have explored finer-grained credit assignment by incorporating turn-level signals and combining them with trajectory-level objectives 17; 29. These methods typically perform trajectory-turn fusion, where trajectory-level rewards are used in their original form, and turn-level signals are either directly applied or heuristically adjusted and propagated across steps. However, they treat all correct tool calls equivalently, without explicitly modeling the varying difficulty of different reasoning steps. As a result, the learning signal remains insensitive to the relative importance and challenge of different tool-use patterns. In this paper, we revisit credit assignment in TIR from a difficulty-aware perspective. As shown in Figure 1, we argue that not all trajectories or tool-calling steps contribute equally to policy improvement. Some trajectories are more informative due to their relative difficulty, and some reasoning steps are inherently harder and thus more valuable for learning. Motivated by this, we propose HiDiffTIR, a Hierarchical Difficulty-aware policy optimization framework for TIR that performs fine-grained credit assignment without requiring additional supervision. HiDiffTIR operates at two levels. At the trajectory level, it differentiates rollouts based on their relative difficulty, enabling the policy to prioritize more informative training signals. At the turn level, it further refines credit assignment by redistributing advantages across reasoning steps according to their difficulty, allowing the model to focus on harder tool-calling decisions. Importantly, the judgment of tool-using difficulty relies solely on group-level statistics derived from standard RL rollouts and does not require external supervision such as LLM-based judges or rubrics. Our contributions are summarized as follows: ⢠We propose HiDiffTIR, a hierarchical policy optimization framework that performs difficulty-aware credit assignment at both the trajectory and turn levels. ⢠We introduce a supervision-free method for fine-grained policy optimization that relies solely on statistics from standard RL rollouts. ⢠Extensive experiments demonstrate that HiDiffTIR improves tool-use accuracy and multi-turn reasoning performance across multiple benchmarks. 2 Related Work 2.1 Tool-Integrated Reasoning TIR has become a central paradigm for equipping LLM agents with external tools to solve complex tasks. Early work primarily relies on Supervised Fine-Tuning (SFT), where models are trained on solution annotations from strong LLMs that interleave reasoning steps and tool invocations 31; 35; 26; 28. While effective in establishing basic tool-use capabilities, recent research has increasingly shifted toward RL to further improve policy learning through interaction 40; 18; 34. Within RL-based TIR, methods have evolved from trajectory-level optimization with outcome-based rewards 18; 4 to incorporating process-based verification signals, such as tool name, parameter name, and argument values, to provide denser supervision 27; 43. More recent approaches further integrate trajectory-level and turn-level signals to provide individual advantage for each turn. For instance, DeepAgent 17 combines outcome rewards with turn-level tool call rewards, while MatchTIR 29 further propagates turn-level rewards across steps using discounted accumulation. Despite these advances, existing methods generally treat all correct tool calls as equally informative, without explicitly modeling the varying difficulty of different trajectories or reasoning steps. Figure 2: Illustration of the HiDiffTIR framework. HiDiffTIR implements a hierarchical difficulty-aware credit assignment mechanism: (a) trajectory-level, where raw turn rewards are reweighted based on turn difficulties to compute a difficulty-aware trajectory advantage, and (b) turn-level, where the trajectory advantage is redistributed to yield fine-grained reweighted turn advantages for policy optimization. 2.2 Reinforcement Learning for LLM Reasoning RL has been widely adopted to improve the reasoning capabilities of LLMs beyond SFT. A prominent line of work is Reinforcement Learning from Human Feedback (RLHF), where models are optimized using preference signals collected from human annotators 1; 24. Recent work has demonstrated the effectiveness of Reinforcement Learning from Verifiable Rewards (RLVR), where models are optimized using automatically computed correctness signals 32; 3; 41. To improve reward quality, a line of studies also concentrates on enhancing credit assignment in LLM reasoning through process-oriented supervision 36; 19. Such approaches typically rely on additional supervision signals, including process reward models 45; 38, LLM-based judges guided by evaluation rubrics 8; 11, to assess intermediate reasoning steps and provide more fine-grained feedback. In parallel, there is growing interest in difficulty-aware RL 44; 6; 9, where training emphasizes more challenging samples to improve model performance, with sample difficulty estimated based on outcome-level signals. However, these methods are largely designed for general reasoning tasks, making them insufficient for jointly modeling process-level feedback and difficulty in TIR. 3 HiDiffTIR 3.1 Problem Formulation Given a user query q and a set of available tools =t1,ā¦,tnT=\t_1,ā¦,t_n\, the goal of a TIR agent ĻĪø _Īø is to solve the query through a multi-turn reasoning and tool use interaction process Ļ=(z1,a1,f1),ā¦,(zTā1,aTā1,fTā1),(zT,aānās)Ļ=\(z_1,a_1,f_1),ā¦,(z_T-1,a_T-1,f_T-1),(z_T,ans)\, where T denotes the total number of turns. At each intermediate turn i<Ti<T, the agent generates a reasoning step ziz_i and a set of tool actions aia_i, where each action specifies a selected tool from T along with its arguments. The environment executes these tool calls and returns feedback fif_i. At the final turn, the agent produces a reasoning step zTz_T followed by the final answer aānāsans without invoking any tool. The objective of a TIR agent is to learn a policy that generates trajectories leading to correct final answers by effectively interleaving reasoning and tool use. 3.2 Overall Framework HiDiffTIR is a hierarchical policy optimization framework for TIR that introduces difficulty-aware credit assignment at two levels. The core idea is to distinguish the relative learning importance of different trajectories and turns, rather than treating them uniformly. As shown in Figure 2, at the trajectory level, HiDiffTIR distinguishes trajectories according to the relative difficulty of all constituent tool calls (§3.3). At the turn level, it further refines credit assignment across reasoning steps, considering both tool-selection difficulty and tool-using difficulty (§3.4). These two components together enable fine-grained optimization of the policy model (§3.5). 3.3 Trajectory-Level Difficulty-Aware Credit Assignment We now introduce trajectory-level difficulty-aware credit assignment, which extends turn-level rewards based on tool-use correctness. Specifically, we first compute the similarity between predicted and ground-truth tool calls to derive turn-level raw rewards. Next, we estimate the relative difficulty of each turn and reweight the raw rewards accordingly. Finally, we normalize these rewards to construct a difficulty-aware trajectory-level advantage. Raw Reward Computation. Following previous works 27; 29, we compute the raw reward for each turn in a trajectory using tool-level correctness. Given a trajectory Ļ(i)Ļ^(i) with T turns, let Ci=ciā1,ā¦,ciāmC_i=\c_i1,ā¦,c_im\ denote the predicted tool calls and Ciā=ciā1ā,ā¦,ciānāC^*_i=\c^*_i1,ā¦,c^*_in\ the ground-truth calls. A similarity matrix MiāāmĆnM_i ^mĆ n is defined as Miā[u,v]=simā(ciāu,ciāvā),M_i[u,v]=sim(c_iu,c^*_iv), (1) where uā1,āÆ,muā\1,Ā·s,m\, vā1,āÆ,nvā\1,Ā·s,n\, and the similarity function simā(ā ,ā )ā[0,1]sim(Ā·,Ā·)ā[0,1] measures agreement on tool name, parameter names, and parameter values. Detailed computation of the similarity function is provided in Appendix A.3. A maximum-weight bipartite matching Ļiā _i^* is then obtained via the Hungarian algorithm 13; 14: Ļiā=argā”maxā”ā(u,v)āĻāĪ ā”Miā[u,v], _i^*= _Ļā _(u,v)āĻM_i[u,v], (2) where Ī denotes all one-to-one matchings between CiC_i and CiāC^*_i. Let Ci,jāCiC_i,j C_i denote the subset of tool calls issued at turn j. The raw reward for turn j is defined as ri,j=1|Ci,j|āciāuāCi,j[(u,v)āĻiā]ā Mi[u,v],r_i,j= 1|C_i,j| _c_iuā C_i,jI[(u,v)ā _i^*]Ā· M_i[u,v], (3) where unmatched tool calls contribute zero. These turn-level rewards serve as the initial supervision signal for subsequent trajectory-level optimization. Difficulty-Aware Reward Reweighting. To model the relative difficulty of different tool-use behaviors, group-level statistics are estimated over trajectories sampled for the same query q. Let ā”(q)=Ļ(1),āÆ,Ļ(G)G(q)=\Ļ^(1),Ā·s,Ļ^(G)\ denote the group of trajectories generated for q. For each ground-truth tool call cāc^*, we compute its average reward across the group: rĀÆā(cā)=1|ā”(q)|āāĻ(k)āā”(q)r(k)ā(cā), r(c^*)= 1|G(q)| _Ļ^(k) (q)r^(k)(c^*), (4) where r(k)ā(cā)r^(k)(c^*) denotes the reward assigned to the tool call in trajectory Ļ(k)Ļ^(k) that is matched to cāc^* via bipartite matching. For each turn j in trajectory Ļ(i)Ļ^(i), let Ci,jāC^*_i,j denote the set of ground-truth tool calls associated with that turn. We define the turn-level average reward as rĀÆi,j=1|Ci,jā|āācāāCi,jārĀÆā(cā). r_i,j= 1|C^*_i,j| _c^*ā C^*_i,j r(c^*). (5) The difficulty of the turn is then defined as di,j=1ārĀÆi,j,d_i,j=1- r_i,j, (6) and used to reweight the raw reward: r~i,j=di,jā ri,j. r_i,j=d_i,jĀ· r_i,j. (7) This weighting emphasizes turns associated with harder tool-use decisions while down-weighting consistently easy ones. Trajectory-Level Score Construction. To ensure stable optimization, the reweighted rewards are normalized within each group. A min-max normalization is first applied: r^i,j=r~i,jāminĻāā”(q)ā”r~Ļ,jmaxĻāā”(q)ā”r~Ļ,jāminĻāā”(q)ā”r~Ļ,j, r_i,j= r_i,j- _Ļ (q) r_Ļ,j _Ļ (q) r_Ļ,j- _Ļ (q) r_Ļ,j, (8) followed by rescaling to preserve the total reward mass: r~i,jnorm=r^i,jā āĻāā”(q)r~Ļ,jāĻāā”(q)r^Ļ,j. r_i,j^norm= r_i,jĀ· _Ļ (q) r_Ļ,j _Ļ (q) r_Ļ,j. (9) The trajectory-level score is then obtained by aggregating normalized turn rewards: Rā”(Ļ(i))=āj=1Tir~i,jnorm,R(Ļ^(i))= _j=1^T_i r_i,j^norm, (10) and converted into a group-relative advantage: Aā”(Ļ(i))=Rā”(Ļ(i))āμqĻq,A(Ļ^(i))= R(Ļ^(i))- _q _q, (11) where μq _q and Ļq _q denote the mean and standard deviation of trajectory scores within ā”(q)G(q). This formulation yields a difficulty-aware trajectory-level objective that emphasizes relatively harder yet successful trajectories while maintaining stable comparisons within each group. 3.4 Turn-Level Difficulty-Aware Credit Assignment While trajectory-level optimization captures the overall difficulty of a rollout, it does not distinguish which reasoning steps are more challenging within the trajectory. To further refine credit assignment, we first estimate the relative difficulty of each turn based on group-level statistics, and then redistribute the trajectory-level advantage across individual turns based on these turn-level difficulties. Group-Level Tool Selection Success Rate. For each ground-truth tool call cāc^*, we compute its tool selection success rate over the group ā”(q)G(q): sā”(cā)=1|ā”(q)|+Ī»āāĻ(k)āā”(q)ā”(r(k)ā(cā)>0)+Ī»,s(c^*)= 1|G(q)|+Ī» _Ļ^(k) (q)I (r^(k)(c^*)>0 )+Ī», (12) where ā”(ā )I(Ā·) is an indicator function, and Ī» is a smoothing constant. Since a positive reward indicates that the correct tool is selected, sā”(cā)s(c^*) effectively measures how frequently each tool call is correctly invoked across trajectories. After that, we aggregate the tool selection success rates into a turn-level statistic for each turn j in trajectory Ļ(i)Ļ^(i): si,j=mincāāCi,jāā”sā”(cā),s_i,j= _c^*ā C^*_i,js(c^*), (13) which emphasizes the most challenging tool call within the turn. Difficulty-Aware Advantage Reweighting. Based on the estimated tool selection success rates, we construct turn-level weights that emphasize more difficult decisions. For trajectories with positive advantage, we assign higher weights to turns with lower success rates while incorporating the overall tool-using quality: w~i,j=ri,jsi,j. w_i,j= r_i,j s_i,j. (14) This formulation suppresses incorrect turns while prioritizing harder correct decisions. For trajectories with negative advantage, we do not apply difficulty-based reweighting and instead use uniform weights, ensuring stable penalization of unsuccessful rollouts. The weights are then normalized within each trajectory: w^i,j=w~i,j1Tiāāk=1Tiw~i,k. w_i,j= w_i,j 1T_i _k=1^T_i w_i,k. (15) Finally, the trajectory-level advantage is redistributed across turns as Ai,j=Aā”(Ļ(i))ā (α+(1āα)āw^i,j),A_i,j=A(Ļ^(i))Ā· (α+(1-α) w_i,j ), (16) where αā[0,1]αā[0,1] controls the balance between trajectory- and turn-level difficulty-aware credit assignment, preserving the global optimization signal while enabling fine-grained emphasis on harder reasoning turns. 3.5 Policy Optimization The final advantage assigned to each token is constructed by combining trajectory-level and turn-level signals in a hierarchical manner. The turn-level advantages Ai,jA_i,j are assigned to all tokens within the corresponding turn, resulting in token-level advantages A~i,t A_i,t. Using the integrated advantage A~i,t A_i,t, we optimize the policy under the GRPO algorithm. Given a batch of queries qā¼q and sampled trajectories Ļii=1Gā¼ĻĪøold(ā ā£q)\ _i\_i=1^G _ _old(Ā· q), the objective is defined as (Īø)=q,Ļi1Gāi=1G1|Ļi|āt=1|Ļi|[min(wi,tA~i,t,clip(wi,t,1āϵ,1+ϵ)A~i,t)āβKL(ĻĪøā„Ļref)], splitJ(Īø)=E_q,\ _i\ 1G _i=1^G 1| _i| _t=1^| _i| [ (w_i,t A_i,t,\\ clip(w_i,t,1-ε,1+ε) A_i,t )- _KL( _Īø\,\|\, _ref) ], split (17) where wi,t=ĻĪøā(yi,tā£yi,<t,q)ĻĪøoldā(yi,tā£yi,<t,q)w_i,t= _Īø(y_i,t y_i,<t,q) _ _old(y_i,t y_i,<t,q) denotes the token-level importance sampling ratio, and β controls the strength of KL regularization. This objective maintains the standard GRPO optimization structure while incorporating difficulty-aware hierarchical credit assignment, enabling more effective learning for multi-turn TIR. 4 Experiments In this section, we evaluate HiDiffTIR through comparative experiments (§4.2), ablation study (§4.3), and tool-use accuracy analysis (§4.4), with additional results detailed in Appendix B. Methods In-Domain (ID) Out-of-Domain (OOD) FTRL BFCL ToolHop Avg. Solve-P Solve-R Solve-F1 Avg. Base MF MP LC AC Qwen3-4B Vanilla 30.78 29.65 25.85 28.76 41.50 31.00 26.50 27.50 31.63 31.63 GRPO 31.13 32.83 30.67 31.54 45.00 37.50 26.50 29.50 37.25 35.15 ToolRL 28.26 28.32 23.78 26.79 32.50 31.00 22.50 20.00 30.28 27.26 FTRL 34.10 34.07 31.54 33.24 43.00 35.50 31.00 28.00 38.63 35.23 MatchTIR (OT) 31.79 37.52 32.60 33.97 50.00 40.50 26.50 35.00 41.95 38.79 MatchTIR (KM) 32.39 39.70 34.21 35.43 50.50 47.00 28.50 36.50 42.55 41.01 HiDiffTIR 35.41 44.74 38.29 39.48 51.00 39.50 31.50 37.00 49.95 41.79 Qwen3-8B Vanilla 28.08 36.55 29.74 31.46 47.50 46.00 37.50 34.50 42.21 41.54 GRPO 31.59 39.75 32.54 34.63 52.50 45.50 34.50 36.50 40.64 41.93 ToolRL 25.57 35.31 26.72 29.20 41.00 39.50 31.50 25.00 32.93 33.99 FTRL 32.32 38.87 32.85 34.68 51.50 45.00 35.50 34.00 36.72 40.54 MatchTIR (OT) 33.61 42.56 33.61 36.59 55.50 52.00 38.50 36.00 45.80 45.56 MatchTIR (KM) 36.33 44.18 37.33 39.28 60.00 49.00 39.00 40.50 46.16 46.93 HiDiffTIR 36.90 44.82 39.04 40.25 61.50 49.00 40.00 42.00 51.26 48.75 Table 1: Performance comparison between HiDiffTIR and the baselines on in-domain and out-of-domain benchmarks across Qwen3-4B and Qwen3-8B backbones. The best results within each backbone group are indicated in bold, while the underlined values represent the second-best results. 4.1 Experimental Setting Training Dataset. We conduct training on the FTRL dataset 43, which comprises 2,215 training data generated via an automated environment construction pipeline. By executing all tools locally as code, it circumvents the unreliability of online APIs, providing stable and verifiable tool-use training. Evaluation Datasets. We evaluate HiDiffTIR on three widely used TIR benchmarks: (1) FTRL 43, the held-out test dataset generated by the FTRL pipeline; (2) ToolHop 42, a query-driven benchmark specifically designed to evaluate multi-hop tool reasoning; and (3) the Berkeley Function Call Leaderboard (BFCL) 25, a comprehensive and executable function call benchmark that evaluates function-calling ability across diverse APIs and parameter configurations. Baselines. We compare HiDiffTIR against several strong TIR baselines: (1) Vanilla, which directly performs TIR using the backbone Qwen3 41 model without RL training; (2) GRPO 32, standard GRPO framework, which computes advantages solely based on the outcome reward; (3) ToolRL 27, an RL-based TIR training framework that introduces process-based tool-call verification rewards; (4) FTRL 43, a feedback-driven training framework using a verifiable reward mechanism based on precision and completeness; and (5) MatchTIR 29, a recent fine-grained credit assignment framework that derives turn-level rewards with two credit assignment manners: MatchTIR (KM), which formulates the assignment as a bipartite matching problem using the Kuhn-Munkres algorithm 13, and MatchTIR (OT), which smooths the assignment through optimal transport 2. Implementation Details. We conduct experiments using Qwen3-4B and Qwen3-8B 41 as the backbone models. All RL training is implemented based on the verl framework 33, with vLLM 15 accelerating the rollout generation. For GRPO settings, we sample a group size of 16 responses per prompt at a temperature of 1.0. The policy model is trained for 5 epochs using a global batch size of 256, restricting the multi-turn reasoning process to a maximum of 10 turns per trajectory. We set the smoothing constant Ī» to 1.0 and the fusion coefficient α to 0.7. All experiments are executed on a single node equipped with 8 NVIDIA H20 GPUs. Additional implementation details are provided in Appendix A. 4.2 Main Results Methods FTRL BFCL ToolHop Avg. Solve-P Solve-R Solve-F1 Avg. Base MF MP LC AC HiDiffTIR 35.41 44.74 38.29 39.48 51.00 39.50 31.50 37.00 49.95 41.79 w/o trajectory-level 33.65 40.69 35.77 36.70 49.50 40.00 30.00 32.50 51.36 40.67 w/o turn-level 33.43 42.45 36.13 37.34 51.50 38.50 29.00 34.50 51.96 41.09 Table 2: Ablation study on credit assignment using Qwen3-4B. The main experimental results are presented in Table 1. On the basis of these results, we make the following observations. HiDiffTIR consistently achieves the best overall performance across different evaluation benchmarks. HiDiffTIR outperforms all baselines on the average score for both Qwen3-4B and Qwen3-8B. Compared with existing RL-based TIR methods, our method demonstrates greater improvements in both in-domain and out-of-domain settings. In particular, our method consistently surpasses recent strong baselines such as FTRL and MatchTIR, showing that incorporating hierarchical difficulty-aware optimization provides complementary benefits beyond existing process-based and turn-level reward modeling strategies. The effectiveness of HiDiffTIR generalizes across different backbone sizes. Although larger backbones generally achieve stronger overall performance, HiDiffTIR consistently improves upon the corresponding baselines under both the 4B and 8B settings. Notably, the performance gains on the in-domain FTRL benchmark are more pronounced for Qwen3-4B, suggesting that smaller models benefit more from fine-grained difficulty-aware optimization. This observation indicates that improved credit assignment can partially compensate for weaker intrinsic reasoning and tool-use capabilities in smaller models. Meanwhile, the consistent improvements on Qwen3-8B demonstrate the robustness and scalability of our framework. HiDiffTIR brings larger improvements on datasets with a larger number of tools. As shown in Table 5, datasets such as FTRL and ToolHop contain substantially larger numbers of tools compared with BFCL. Correspondingly, HiDiffTIR achieves more evident improvements on these benchmarks, which require reasoning over thousands of candidate tools. In contrast, on datasets with relatively fewer tools, such as BFCL, the performance improvements are comparatively modest, while stronger backbone models already obtain noticeable benefits from scaling alone. These findings highlight that the advantages of difficulty-aware optimization become increasingly important as the complexity of the tool environment grows. Current RL-based TIR methods still struggle with missing-information scenarios. Although most methods achieve clear improvements on the BFCL benchmark, the gains on the Missing Functions (MF) and Missing Parameters (MP) subsets remain modest, as they require the model to recognize when the currently available tools or user-provided information are insufficient. Existing RL-based TIR training mainly emphasizes successful tool execution and task completion, which may bias the policy toward aggressively invoking tools rather than learning insufficiency detection. Specifically, because our difficulty-aware reweighting amplifies successful complex tool sequences, it inadvertently reinforces this aggressive calling bias in sparse tool settings. Consequently, while HiDiffTIR significantly outperforms most baselines on the MF subset, it slightly trails MatchTIR. This observation suggests an important future direction for TIR, i.e., integrating self-awareness of capability boundaries into RL optimization. 4.3 Ablation Study We further conduct ablation studies on Qwen3-4B to evaluate the contribution of each component in HiDiffTIR. Effectiveness of Hierarchical Components. As shown in Table 2, removing either trajectory-level or turn-level difficulty-aware credit assignment leads to consistent performance degradation compared with the full model, demonstrating that both components contribute to the final performance. In particular, removing trajectory-level optimization results in a larger overall drop on FTRL, indicating that distinguishing the relative difficulty of different trajectories provides an important optimization signal for RL training. Furthermore, performance variations on Out-Of-Domain (OOD) datasets like BFCL and ToolHop are less pronounced when omitting specific components. Since our difficulty signals are derived dynamically during training, they naturally target bottlenecks within the training distribution. Thus, these components are highly effective for seen tools but have a more constrained impact on completely unseen tools in OOD scenarios, although the full HiDiffTIR model still achieves the best overall average performance. (a) Average Response Length (b) Generation Time Figure 3: Training dynamics regarding the average response length and generation time of different variants. Table 3: Comparison of reasoning efficiency at the final training step. Methods Avg. Resp. Len. Gen. Time (s) HiDiffTIR 1,893 1,058 w/o trajectory-level 1,870 1,007 w/o turn-level 2,085 1,213 Efficiency of Hierarchical Components. The training dynamics of reasoning efficiency are visualized in Figure 3, with convergence values summarized in Table 3. We observe that HiDiffTIR achieves higher reasoning compactness compared to the w/o turn-level variant, which relies solely on trajectory-level credit assignment. Specifically, at the final training step, HiDiffTIR reduces the average response length from 2,085 to 1,893 tokens, resulting in a 12.8% decrease in generation latency. This suggests that fine-grained turn-level reweighting effectively identifies critical reasoning steps and suppresses redundant tool calls. Although the w/o trajectory-level variant exhibits slightly lower latency, its performance is substantially inferior to the full model, confirming that our hierarchical design achieves an optimal trade-off between reasoning accuracy and computational efficiency. 4.4 Tool-Use Accuracy Analysis Figure 4: Training dynamics regarding tool-use accuracy of different models. Methods FTRL Solve-P Solve-R Solve-F1 Avg. HiDiffTIR (α=0.3α=0.3) 32.72 37.70 33.67 34.70 HiDiffTIR (α=0.5α=0.5) 33.96 40.33 35.79 36.69 HiDiffTIR (α=0.7α=0.7) 35.41 44.74 38.29 39.48 Table 4: Performance of models on FTRL dataset under different fusion coefficients. In this section, we analyze how difficulty-aware credit assignment influences the tool-use accuracy. Impact on Tool Selection Precision. As illustrated in Figure 4, all variants of HiDiffTIR consistently outperform the competitive baseline MatchTIR (KM) in tool-use accuracy throughout the training process. This convergence to higher accuracy confirms that integrating difficulty-aware credit assignment provides a clearer gradient signal, effectively guiding the policy to master correct tool invocation patterns. Specifically, variants with smaller fusion coefficients (α=0.3α=0.3 and α=0.5α=0.5) exhibit higher peak tool accuracy during training compared to α=0.7α=0.7. This trend occurs because a lower α assigns a larger proportional weight to the turn-level difficulty-aware credit assignment, forcing the policy to aggressively optimize high-difficulty tool calls within training rollouts. Sensitivity Analysis of Fusion Coefficient. Despite the higher training accuracy at lower α values, Table 4 reveals that α=0.7α=0.7 achieves the highest average score on the FTRL test set. This counterintuitive phenomenon highlights the cooperative dynamics between the two credit assignment levels. A smaller α over-emphasizes turn-level difficulty-aware credit assignment, rigidly rewarding only ground-truth tool invocations and suppressing intermediate turns with necessary trial-and-error exploration. In complex TIR, such exploration is not meaningless but is indispensable for the agent to gather feedback and learn from failures. Conversely, setting α=0.7α=0.7 provides a well-suited balance, leveraging trajectory-level credit assignment to safeguard the value of the exploratory sequence while utilizing turn-level credit assignment to resolve specific tool-use bottlenecks. 5 Conclusions In this paper, we propose HiDiffTIR, a hierarchical policy optimization framework for TIR that performs difficulty-aware credit assignment at both the trajectory and turn levels. By modeling the relative difficulty of trajectories and redistributing advantages across reasoning steps, our method provides fine-grained learning signals that better capture the heterogeneous importance of tool-calling decisions, without requiring additional supervision. Extensive experiments on multiple tool-using benchmarks demonstrate that HiDiffTIR consistently improves multi-turn TIR performance and tool invocation accuracy, suggesting that incorporating difficulty-aware signals is a promising direction for improving the training of tool-using LLM agents. Limitations Despite the strong empirical results, our proposed HiDiffTIR has several limitations. Primarily, owing to limited computational resources, our empirical evaluations were restricted to moderately sized open-source models, specifically Qwen3-4B and Qwen3-8B. Although the core algorithmic design of our hierarchical difficulty-aware policy optimization framework is inherently model-agnostic, its behavioral dynamics and scalability when integrated into LLMs with larger parameters have not yet been fully validated. Additionally, because our difficulty-aware signals are derived dynamically from online group-level rollout statistics, the framework implicitly assumes that the base model possesses a foundational level of instruction-following capability. In exceptionally intricate tasks where the initial policy completely fails to produce any valid tool invocations, the empirical difficulty estimation might suffer from extreme reward sparsity, potentially necessitating a supervised warm-up phase to kickstart effective reinforcement learning exploration. Finally, the difficulty estimation mechanism in our framework serves as a heuristic. It captures empirical learning difficulty from the perspective of operational bottlenecks for the policy during RL rollouts, rather than providing an absolute measure of static semantic complexity. References Christiano et al. (2017) P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei Deep reinforcement learning from human preferences. Advances in neural information processing systems 30. External Links: Link Cited by: §2.2. Cuturi (2013) M. Cuturi Sinkhorn distances: lightspeed computation of optimal transport. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2, NIPSā13, Red Hook, NY, USA, p. 2292ā2300. Cited by: 5th item, §4.1. DeepSeek-AI et al. (2025) DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948. External Links: Link Cited by: §1, §2.2. Dong et al. (2025) G. Dong, Y. Chen, X. Li, J. Jin, H. Qian, Y. Zhu, H. Mao, G. Zhou, Z. Dou, and J. Wen Tool-star: empowering llm-brained multi-tool reasoner via reinforcement learning. arXiv preprint arXiv:2505.16410. Cited by: §1, §2.1. Frieder et al. (2023) S. Frieder, L. Pinchetti, C. Chevalier, R. Griffiths, T. Salvatori, T. Lukasiewicz, P. Petersen, and J. Berner Mathematical capabilities of chatgpt. Advances in neural information processing systems 36, p. 27699ā27744. Cited by: §1. Gao et al. (2026) H. Gao, zhenyu zhang, L. Pang, F. Guo, douhongjian, G. Lv, S. Liu, T. Gao, H. Shen, and X. Cheng DIVA-GRPO: enhancing multimodal reasoning through difficulty-adaptive variant advantage. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2.2. Gou et al. (2024) Z. Gou, Z. Shao, Y. Gong, yelong shen, Y. Yang, M. Huang, N. Duan, and W. Chen ToRA: a tool-integrated reasoning agent for mathematical problem solving. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1. Gunjal et al. (2025) A. Gunjal, A. Wang, E. Lau, V. Nath, Y. He, B. Liu, and S. M. Hendryx Rubrics as rewards: reinforcement learning beyond verifiable domains. In NeurIPS 2025 Workshop on Efficient Reasoning, External Links: Link Cited by: §2.2. Heckel et al. (2026) R. Heckel, M. Soltanolkotabi, and C. Thramboulidis Asymmetric prompt weighting for reinforcement learning with verifiable rewards. arXiv preprint arXiv:2602.11128. Cited by: §2.2. Huang et al. (2024) X. Huang, W. Liu, X. Chen, X. Wang, H. Wang, D. Lian, Y. Wang, R. Tang, and E. Chen Understanding the planning of llm agents: a survey. arXiv preprint arXiv:2402.02716. Cited by: §1. Huang et al. (2025) Z. Huang, Y. Zhuang, G. Lu, Z. Qin, H. Xu, T. Zhao, R. Peng, J. Hu, Z. Shen, X. Hu, et al. Reinforcement learning with rubric anchors. arXiv preprint arXiv:2508.12790. Cited by: §2.2. Jin et al. (2025) B. Jin, H. Zeng, Z. Yue, J. Yoon, S. O. Arik, D. Wang, H. Zamani, and J. Han Search-r1: training LLMs to reason and leverage search engines with reinforcement learning. In Second Conference on Language Modeling, External Links: Link Cited by: §1. Kuhn (1955) H. W. Kuhn The hungarian method for the assignment problem. Naval Research Logistics Quarterly 2 (1-2), p. 83ā97. External Links: Document, Link, https://onlinelibrary.wiley.com/doi/pdf/10.1002/nav.3800020109 Cited by: 5th item, §3.3, §4.1. Kuhn (1956) H. W. Kuhn Variants of the hungarian method for assignment problems. Naval Research Logistics Quarterly 3 (4), p. 253ā258. External Links: Document, Link, https://onlinelibrary.wiley.com/doi/pdf/10.1002/nav.3800030404 Cited by: §3.3. Kwon et al. (2023) W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, External Links: Link Cited by: §A.5, §4.1. Li et al. (2025a) X. Li, G. Dong, J. Jin, Y. Zhang, Y. Zhou, Y. Zhu, P. Zhang, and Z. Dou Search-o1: agentic search-enhanced large reasoning models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 5420ā5438. Cited by: §1. Li et al. (2026) X. Li, W. Jiao, J. Jin, G. Dong, J. Jin, Y. Wang, H. Wang, Y. Zhu, J. Wen, Y. Lu, and Z. Dou DeepAgent: a general reasoning agent with scalable toolsets. In Proceedings of the ACM Web Conference 2026, W ā26, New York, NY, USA, p. 2219ā2230. External Links: ISBN 9798400723070, Link, Document Cited by: §1, §2.1. Li et al. (2025b) X. Li, H. Zou, and P. Liu Torl: scaling tool-integrated rl. arXiv preprint arXiv:2503.23383. Cited by: §1, §1, §2.1. Lightman et al. (2023) H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe Letās verify step by step. In The twelfth international conference on learning representations, Cited by: §2.2. Lin and Xu (2025) H. Lin and Z. Xu Understanding tool-integrated reasoning. arXiv preprint arXiv:2508.19201. Cited by: §1. Luo et al. (2025) J. Luo, W. Zhang, Y. Yuan, Y. Zhao, J. Yang, Y. Gu, B. Wu, B. Chen, Z. Qiao, Q. Long, R. Tu, X. Luo, W. Ju, Z. Xiao, Y. Wang, M. Xiao, C. Liu, J. Yuan, S. Zhang, Y. Jin, F. Zhang, X. Wu, H. Zhao, D. Tao, P. S. Yu, and M. Zhang Large language model agent: a survey on methodology, applications and challenges. arXiv preprint arXiv:2503.21460. Cited by: §1. Masterman et al. (2024) T. Masterman, S. Besen, M. Sawtell, and A. Chao The landscape of emerging ai agent architectures for reasoning, planning, and tool calling: a survey. arXiv preprint arXiv:2404.11584. Cited by: §1. OpenAI et al. (2023) OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Å. Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Å. Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. MĆ©ly, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. OāKeefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph Gpt-4 technical report. arXiv preprint arXiv:2303.08774. External Links: Link Cited by: §1. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. Advances in neural information processing systems 35, p. 27730ā27744. External Links: Link Cited by: §2.2. Patil et al. (2025) S. G. Patil, H. Mao, F. Yan, C. C. Ji, V. Suresh, I. Stoica, and J. E. Gonzalez The berkeley function calling leaderboard (BFCL): from tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: 3rd item, §4.1. Patil et al. (2024) S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez Gorilla: large language model connected with massive apis. Advances in Neural Information Processing Systems 37, p. 126544ā126565. Cited by: §1, §2.1. Qian et al. (2025) C. Qian, E. C. Acikgoz, Q. He, H. Wang, X. Chen, D. Hakkani-Tür, G. Tur, and H. Ji Toolrl: reward is all tool learning needs. arXiv preprint arXiv:2504.13958. Cited by: 3rd item, §A.3, §1, §1, §2.1, §3.3, §4.1. Qin et al. (2024) Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, dahai li, Z. Liu, and M. Sun ToolLLM: facilitating large language models to master 16000+ real-world APIs. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1, §2.1. Qu et al. (2026) C. Qu, S. Dai, H. Cai, J. Xu, S. Wang, and D. Yin MatchTIR: fine-grained supervision for tool-integrated reasoning via bipartite matching. arXiv preprint arXiv:2601.10712. Cited by: 5th item, §A.3, §1, §2.1, §3.3, §4.1. Qu et al. (2025) C. Qu, S. Dai, X. Wei, H. Cai, S. Wang, D. Yin, J. Xu, and J. Wen Tool learning with large language models: a survey. Frontiers of Computer Science 19 (8), p. 198343. Cited by: §1. Schick et al. (2023) T. Schick, J. Dwivedi-Yu, R. DessƬ, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom Toolformer: language models can teach themselves to use tools. Advances in neural information processing systems 36, p. 68539ā68551. Cited by: §2.1. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo Deepseekmath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. External Links: Link Cited by: 2nd item, §1, §2.2, §4.1. Sheng et al. (2025) G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu Hybridflow: a flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, p. 1279ā1297. External Links: Link Cited by: §4.1. Singh et al. (2025) J. Singh, Y. Pandya, P. Vajreshwari, R. Magazine, and A. Nambi Agentic reasoning and tool integration for LLMs via reinforcement learning. In First Workshop on Foundations of Reasoning in Language Models, External Links: Link Cited by: §2.1. Tang et al. (2023) Q. Tang, Z. Deng, H. Lin, X. Han, Q. Liang, B. Cao, and L. Sun Toolalpaca: generalized tool learning for language models with 3000 simulated cases. arXiv preprint arXiv:2306.05301. Cited by: §2.1. Uesato et al. (2022) J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275. Cited by: §2.2. Wang et al. (2024) L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, W. X. Zhao, Z. Wei, and J. Wen A survey on large language model based autonomous agents. Front. Comput. Sci. 18 (6). External Links: ISSN 2095-2228, Link, Document Cited by: §1. Wang et al. (2025) W. Wang, Z. Gao, L. Chen, Z. Chen, J. Zhu, X. Zhao, Y. Liu, Y. Cao, S. Ye, X. Zhu, L. Lu, H. Duan, Y. Qiao, J. Dai, and W. Wang Visualprm: an effective process reward model for multimodal reasoning. arXiv preprint arXiv:2503.10291. Cited by: §2.2. Xu et al. (2025) W. Xu, C. Huang, S. Gao, and S. Shang LLM-based agents for tool learning: a survey. Data Science and Engineering, p. 1ā31. Cited by: §1. Xue et al. (2025) Z. Xue, L. Zheng, Q. Liu, Y. Li, X. Zheng, Z. Ma, and B. An Simpletir: end-to-end reinforcement learning for multi-turn tool-integrated reasoning. arXiv preprint arXiv:2509.02479. Cited by: §1, §2.1. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. arXiv preprint arXiv:2505.09388. External Links: Link Cited by: 1st item, §1, §2.2, §4.1, §4.1. Ye et al. (2025a) J. Ye, Z. Du, X. Yao, W. Lin, Y. Xu, Z. Chen, Z. Wang, S. Zhu, Z. Xi, S. Yuan, T. Gui, Q. Zhang, X. Huang, and J. Chen ToolHop: a query-driven benchmark for evaluating large language models in multi-hop tool use. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 2995ā3021. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: 2nd item, §4.1. Ye et al. (2025b) J. Ye, C. Jiang, Z. Du, Y. Xu, X. Yao, Z. Xi, X. Fan, Q. Zhang, X. Huang, and J. Chen Feedback-driven tool-use improvements in large language models via automated build environments. arXiv preprint arXiv:2508.08791. Cited by: 1st item, 4th item, §1, §2.1, §4.1, §4.1, §4.1. Zhang and Zuo (2025) J. Zhang and C. Zuo Grpo-lead: a difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 5642ā5665. Cited by: §2.2. Zhang et al. (2025) Z. Zhang, C. Zheng, Y. Wu, B. Zhang, R. Lin, B. Yu, D. Liu, J. Zhou, and J. Lin The lessons of developing process reward models in mathematical reasoning. In Findings of the Association for Computational Linguistics: ACL 2025, p. 10495ā10516. Cited by: §2.2. Zhao et al. (2023) Y. Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li PyTorch fsdp: experiences on scaling fully sharded data parallel. Proc. VLDB Endow. 16 (12), p. 3848ā3860. External Links: ISSN 2150-8097, Link, Document Cited by: §A.5. Appendix A Implementation Details A.1 Datasets and Evaluation Metrics Datasets. The statistics of the training and evaluation datasets are shown in Table 5. Table 5: Dataset statistics Dataset # of queries # of tools FTRL (train) 2,215 4,433 FTRL (test) 200 1,109 ToolHop 995 3,912 BFCL (multi-turn) 800 84 Evaluation Metrics. We evaluate model performance using metrics of each benchmark. ⢠FTRL 43. Let NgāeānN_gen denote the total number of tool calls generated by the model, NgātN_gt denote the total number of required tool calls in the ground-truth trajectory, and NmāaātācāhN_match denote the number of successfully matched tool calls. The metrics of FTRL are defined as follows: ā Solve-P measures the precision of tool invocations. Solve-P=NmāaātācāhNgāeānif āNgāeān>01.0if āNgāeān=0Solve-P= cases N_matchN_gen&if N_gen>0\\ 1.0&if N_gen=0 cases (18) ā Solve-R evaluates the completeness of the task execution. Solve-R=NmāaātācāhNgātSolve-R= N_matchN_gt (19) ā Solve-F1 is the harmonic mean of Solve-P and Solve-R. Solve-F1=2ā NmāaātācāhNgāeān+NgātSolve-F1= 2Ā· N_matchN_gen+N_gt (20) ⢠ToolHop 42. The metric of ToolHop is Answer Correctness (AC). Assume that for a test instance, the standard answer is a and the modelās final response is o. The performance is evaluated as: AC=1if āaāo0otherwise.AC= cases1&if aā o\\ 0&otherwise cases. (21) ⢠BFCL 25. In BFCL V3, model performance is evaluated across several categories, including Base Multi-Turn (Base), Missing Functions (MF), Missing Parameters (MP), and Long-Context Multi-Turn (LC). A.2 Baselines We compare HiDiffTIR with following baselines: ⢠Vanilla. We deploy the backbone Qwen3 41 model using standard generation configurations without any further RL training. This serves as the lower bound, demonstrating the inherent instruction-following and tool-use capabilities of the model. ⢠GRPO 32. GRPO serves as the standard trajectory-level RL baseline. During training, for a given prompt, it generates a group of G trajectories, evaluates the final outcome of each trajectory to assign a reward. The advantages are computed by normalizing these rewards within the group and are subsequently broadcast uniformly to all tokens in the trajectory. ⢠ToolRL 27. This framework introduces process-based supervision to TIR. Rather than relying solely on outcome rewards, ToolRL incorporates intermediate verification rewards such as the correctness of tool names, parameter names, and parameter values. These intermediate rewards are aggregated into a single trajectory reward, and the resulting advantage is then shared across the entire trajectory during policy optimization. ⢠FTRL 43. The method accompanying the FTRL dataset. It employs a verifiable reward mechanism that evaluates the precision of tool invocations and the completeness of task execution. These factors are combined to calculate a total reward, which is used to derive a single advantage value broadcast to all tokens in the sequence. ⢠MatchTIR 29. The state-of-the-art fine-grained credit assignment TIR framework that derives dense, turn-level advantages by aligning generated trajectories with ground-truth tool-use sequences. We evaluate two implementations of its alignment module: (1) MatchTIR (KM) formulates the alignment as a hard bipartite matching problem. It uses the Kuhn-Munkres (KM) algorithm 13 to find the optimal one-to-one mapping between generated and ground-truth tool calls based on semantic similarity, distributing rewards strictly to matched pairs. (2) MatchTIR (OT) employs Optimal Transport (OT) 2 with Sinkhorn solver to create a soft alignment matrix. This allows for a probabilistic distribution of rewards across multiple generated turns that share semantic overlap with the ground truth. Table 6: Training prompt template for the policy model. Training Prompt Template for the Policy Model system # Tools You may call one or more functions to assist with the user query. You are provided with function signatures within <tools></tools> XML tags: <tools> Tool Descriptions </tools> For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags: <tool_call> ānameā: <function-name>, āargumentsā: <args-json-object> </tool_call> user Please call given tools to answer the question. Please note that all your information must be obtained by calling tools and not by answering the question directly. If the call fails, you need to try to correct it and continue until you arrive at an answer. Only output the final answer (in words, numbers or phrase) inside the <answer></answer> tag, without any explanations or extra information. Question: question assistant A.3 Tool Call Similarity Computation For reward computation, a matching matrix MiāāmĆnM_i ^mĆ n is constructed for each trajectory Ļ(i)Ļ^(i), where each entry Miā[u,v]M_i[u,v] measures the alignment between a predicted tool call ciāuāCic_iuā C_i and a ground-truth tool call ciāvāāCiāc^*_ivā C^*_i. The similarity score Miā[u,v]M_i[u,v] consists of three components below, following previous works 27; 29. Tool Name Matching. Let ciāunamec_iu^name and ciāvnameāc_iv^name* denote the tool names of the predicted and ground-truth calls. The tool name score is sname=(ciāuname=ciāvnameā)ā0,1,s_name=I(c_iu^name=c_iv^name*)ā\0,1\, (22) where ā”[ā ]I[Ā·] is an indicator function. Parameter Name Matching. If the tool names match, the parameter name similarity is computed as the Jaccard similarity over the sets of parameter names PciāuP_c_iu and PciāvāP_c^*_iv: sparam=|Pciāuā©Pciāvā||PciāuāŖPciāvā|ā[0,1].s_param= |P_c_iuā© P_c^*_iv||P_c_iuāŖ P_c^*_iv|ā[0,1]. (23) Parameter Value Matching. Finally, the correctness of parameter values is assessed for each ground-truth parameter kāPciāvākā P_c^*_iv: svalue=ākāPciāvāā”(ciāuā[k]=ciāvāā[k])ā[0,|Pciāvā|].s_value= _kā P_c^*_ivI(c_iu[k]=c^*_iv[k])ā[0,|P_c^*_iv|]. (24) The three components are combined and normalized to produce the final similarity score: Miā[u,v]=snameā sname+sparam+svalue2+|Pciāvā|ā[0,1].M_i[u,v]=s_nameĀ· s_name+s_param+s_value2+|P_c^*_iv|ā[0,1]. (25) A.4 Training Prompt Template The training prompt template for the policy model is shown in Table 6. A.5 Training Details To facilitate reproducibility, we provide the complete set of hyperparameters and infrastructure configurations used for training HiDiffTIR. Infrastructure and Hardware. All models are trained on a single compute node equipped with 8 NVIDIA H20 GPUs. We leverage Fully Sharded Data Parallel (FSDP) 46 to distribute the policy model weights, avoiding parameter and optimizer offloading to maximize throughput. For rollout generation, we utilize vLLM 15 with a GPU memory utilization threshold of 0.7. Data and Generation Settings. We format the prompts using the standard Qwen3 system style and explicitly enable the generation of reasoning trajectories. To handle long trajectories, we set the maximum model length to 32,768 tokens and the maximum response length to 8,192 tokens. During the rollout phase, we sample 16 responses per prompt (G=16G=16) with a generation temperature of 1.0 and a maximum turn of 10. Optimization and Loss Formulation. The policy is optimized over 5 total epochs. We use a global training batch size of 256 prompts, broken down into mini-batch sizes of 32. The actor learning rate is fixed at 10ā610^-6 without a critic warmup phase. The entropy coefficient is set to 0.001, and the KL penalty coefficient is set to 0.001. Appendix B Additional Experiments B.1 Group Size Analysis Table 7: Performance comparison of HiDiffTIR on Qwen3-4B with different group sizes G. Methods FTRL BFCL ToolHop Avg. Solve-P Solve-R Solve-F1 Avg. Base MF MP LC AC HiDiffTIR (G=4) 30.24 36.93 31.93 33.03 49.50 41.00 26.00 34.50 49.95 40.19 HiDiffTIR (G=8) 32.90 40.65 34.74 36.10 47.50 37.50 30.00 32.50 51.36 39.77 HiDiffTIR (G=16) 35.41 44.74 38.29 39.48 51.00 39.50 31.50 37.00 49.95 41.79 The group size G is a critical hyperparameter in GRPO-based frameworks, as it determines the sample size used to estimate the baseline and normalize rewards within each group. In HiDiffTIR, G further influences the robustness of the empirical difficulty estimation used for both trajectory-level and turn-level credit assignment. To investigate this impact, we conduct experiments on Qwen3-4B with α=0.7α=0.7 across three different group sizes: 4, 8, and 16. The experimental results are summarized in Table 7. We observe distinct performance trends between the In-Domain (ID) and Out-Of-Domain (OOD) metrics. For the ID FTRL dataset, performance consistently improves as the group size increases. The ID average score rises steadily from 33.03 (G=4G=4) to 39.48 (G=16G=16). This trend indicates that our hierarchical difficulty-aware mechanisms benefit from larger group sizes, which provide more statistically robust estimations of group-level statistics within each training iteration. Accurately quantifying the relative difficulty of trajectories and turns allows the model to better identify and optimize informative reasoning steps in a difficulty-aware manner, thus improving ID task success rates. Conversely, for OOD scenarios, the relationship is not strictly monotonic. While the optimal OOD average score of 41.79 is achieved at G=16G=16, G=8G=8 actually shows a slight regression (39.77) compared to G=4G=4 (40.19). This volatility stems from the intrinsic mechanism of our difficulty-aware framework. By design, HiDiffTIR utilizes group rollouts to intensively optimize the policy toward mastering seen tools, particularly focusing on resolving high-difficulty tool-calling bottlenecks. While this targeted credit assignment significantly enhances the modelās proficiency and precision on seen tool-use patterns, it can lead to behavioral variances when the model encounters entirely unseen tools in OOD environments, as their underlying API structures and difficulty characteristics differ from the training set. Nevertheless, G=16G=16 still yields a well-suited balance, achieving the highest performance on both in-domain tasks and overall OOD generalization. B.2 Case Study To intuitively demonstrate the superiority of HiDiffTIR, we present a qualitative case study involving a complex 5-hop reasoning query. As shown in Table 8, the user query requires identifying a specific local handicraft through a chain of five implicit entities. We compare our approach against the untrained Qwen3-4B backbone and MatchTIR (KM), a strong baseline built on the same backbone, with results reported in Tables 9-11, respectively. As illustrated in the interaction trajectories, HiDiffTIR exhibits the following distinct advantages. Mitigation of Parameter Hallucination and Error Loops. Both baseline models exhibit a tendency to hallucinate tool parameters during the initial reasoning phase. As shown in Table 10 and Table 11, when calling the initial famous_peak_identifier tool, both Qwen3-4B and MatchTIR incorrectly append unnecessary geographic constraints, hallucinating āAlpsā or āHimalayasā instead of passing the required empty arguments. While MatchTIR manages to correct itself after wasting two turns, the untrained Qwen3-4B model lacks this reflective capability, entering a fatal trial-and-error loop that continuously guesses random regions until it exhausts the maximum turn limit. In contrast, HiDiffTIR completely suppresses these hallucinated priors, immediately generating the precise arguments required to navigate the tool logic. Calibration of Overconfidence and Adaptation to Environmental Feedback. A further striking observation from the trajectories is the backbone modelās overconfidence in its parametric knowledge. When facing consecutive empty responses from the tools, the untrained Qwen3-4B exhibits severe resistance to self-correction. Instead of rectifying its over-specified parameters, the model stubbornly suspects that the tool itself is flawed (e.g., explicitly assuming in Turn 6 and Turn 9 that āthe toolās dataā or āthe toolās databaseā is incomplete). This overconfidence traps the agent, preventing it from utilizing negative environmental feedback. While the RL-trained MatchTIR largely mitigates this extreme overconfidence and demonstrates the capacity to adjust based on environmental feedback, it still exhibits a lingering over-reliance on its internal parametric knowledge, frequently defaulting to hallucinated assumptions during exploration. Moving beyond this behavior, HiDiffTIR effectively neutralizes this stubbornness, teaching the LLM agent to respect and adapt to strict tool constraints rather than blindly persisting with its hallucinated assumptions. Enhanced Reasoning Efficiency and Minimal Redundancy. Beyond hallucination and overconfidence, baseline models also struggle with inefficient and redundant exploration. As observed in the trajectories, whenever the untrained Qwen3-4B faces an obstacle, it continuously appends arbitrary parameters (e.g., include_protected in Turn 8 and environmental_certification in Turn 9) in futile attempts to force a result. Even the RL-trained MatchTIR, despite eventually reaching the correct answer, exhibits similar inefficiency by repeatedly over-specifying non-essential arguments (e.g., guessing radius in Turn 4 and extra boolean constraints in Turn 7). In contrast, HiDiffTIR learns a highly compact reasoning policy with minimal redundancy. In this specific case, it executes the entire task flawlessly in exactly five optimal tool-calling turns, achieving absolute zero redundancy. This highlights the core advantage of our hierarchical difficulty-aware policy optimization framework. By accurately amplifying the learning signals for high-difficulty bottleneck steps, the framework implicitly penalizes superfluous exploratory actions, yielding a much more efficient and precise reasoning trajectory compared to uniform or heuristic-based credit assignments. Table 8: The user query and ground truth answer of the case study. User Query: What is the local handicraft of the town where the woodworker who carves sculptures using the wood from the forest near the mountain with a famous peak? Ground Truth Tool Calls: 1. ānameā: āfamous_peak_identifierā, āargumentsā: 2. ānameā: ānearby_forest_locatorā, āargumentsā: ālandmarkā: āMount Everestā 3. ānameā: āwood_sourcing_locatorā, āargumentsā: āforest_nameā: āLangtang National Park Forestā, āproduct_typeā: āsculpturesā 4. ānameā: āindividual_location_finderā, āargumentsā: āindividual_nameā: āRamesh Thapaā 5. ānameā: ālocal_handicraft_identifierā, āargumentsā: ālocationā: āPokharaā Ground Truth Answer: Wooden masks Table 9: Interaction trajectory of HiDiffTIR. The model conducts precise TIR by successfully chaining five sequential tool calls without any hallucinations or redundant trial-and-error. HiDiffTIR [Turn 1] <think> Okay, letās tackle this question step by step ⦠So, the plan is to first find the mountain with a famous peak. Then, determine the forest near that mountain. Next, find the woodworker who uses that forestās wood to carve sculptures. Then, locate the town where this woodworker is based, and finally, identify the local handicraft of that town ⦠Starting with the mountain with a famous peak. The tool āfamous_peak_identifierā can help here ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: </tool_call> <tool_response> Mount Everest - The highest mountain in the world, known for its climbing fame. </tool_response> [Turn 2] <think> Okay, so the first step was to find a mountain with a famous peak, and the response was Mount Everest. Now, the next part of the question is to find a forest near this mountain. The tool to use here is nearby_forest_locator ⦠</think> <tool_call> ānameā: ānearby_forest_locatorā, āargumentsā: ālandmarkā: āMount Everestā </tool_call> <tool_response> āforestā: āLangtang National Park Forestā </tool_response> [Turn 3] <think> Okay, so the forest near Mount Everest is Langtang National Park Forest. Now the next step is to find a woodworker who carves sculptures using wood from this forest. The relevant tool here is wood_sourcing_locator ⦠</think> <tool_call> ānameā: āwood_sourcing_locatorā, āargumentsā: āforest_nameā: āLangtang National Park Forestā, āproduct_typeā: āsculpturesā </tool_call> <tool_response> Ramesh Thapa is a woodworker who sources wood from Langtang National Park Forest. </tool_response> [Turn 4] <think> Okay, now I know the woodworker is Ramesh Thapa. The next step is to find out where heās based. The tool to use here is individual_location_finder ⦠</think> <tool_call> ānameā: āindividual_location_finderā, āargumentsā: āindividual_nameā: āRamesh Thapaā </tool_call> <tool_response> ālocationā: āPokharaā </tool_response> [Turn 5] <think> Okay, now I know Ramesh Thapa is based in Pokhara. The next step is to find the local handicraft of Pokhara. The tool to use here is local_handicraft_identifier ⦠</think> <tool_call> ānameā: ālocal_handicraft_identifierā, āargumentsā: ālocationā: āPokharaā </tool_call> <tool_response> Wooden masks </tool_response> [Turn 6] <think> ⦠All the tool calls led to the correct answers, so the final answer should be wooden masks. </think> The local handicraft of the town where the woodworker who carves sculptures using wood from the forest near the mountain with a famous peak (Mount Everest) is wooden masks. Table 10: Interaction trajectory of Qwen3-4B. The model hallucinates tool parameters, enters a trial-and-error loop, and ultimately fails to complete the reasoning chain within the turn limit. Qwen3-4B [Turn 1] [Issue: Parameter Hallucination & Redundant Over-specification] <think> Okay, letās tackle this question step by step ⦠So, the plan is to first find the mountain with a famous peak ⦠Letās assume the mountain is in a region like the Alps ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āregionā: āAlpsā, āfame_typeā: āclimbingā </tool_call> <tool_response> No famous peaks found matching the criteria. </tool_response> [Turn 2] [Issue: Parameter Hallucination & Redundant Over-specification] <think> Okay, the first tool call didnāt find any famous peaks in the Alps with a climbing reputation. Maybe I need to adjust the parameters ⦠So, Iāl call the famous_peak_identifier again with region set to āHimalayasā and fame_type as āscenicā. </think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āregionā: āHimalayasā, āfame_typeā: āscenicā </tool_call> <tool_response> No famous peaks found matching the criteria. </tool_response> [Turns 3-5] (Repeated failed attempts to find the peak) [Issue: Trapped in Error Loop] [Turn 6] [Issue: Parameter Hallucination & Redundant Over-specification & Overconfidence] <think> Okay, after several attempts to find a famous peak using different regions and fame types, it seems the tool isnāt returning any results. Maybe the issue is with the parameters or the toolās data ⦠Letās try using a specific mountain name, like Mount Fuji, and check if itās in the Himalayas ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āregionā: āJapanā, āfame_typeā: āculturalā, āhistorical_eventā: āfirst_ascentā </tool_call> <tool_response> Mount Fuji - A cultural icon of Japan. </tool_response> [Turn 7] [Issue: Redundant Over-specification] <think> ⦠The next step is to find the forest near Mount Fuji. Using the ānearby_forest_locatorā tool with the landmark as Mount Fuji and a reasonable radius, say 50 km ⦠</think> <tool_call> ānameā: ānearby_forest_locatorā, āargumentsā: ālandmarkā: āMount Fujiā, āregionā: āJapanā, āradiusā: 50 </tool_call> <tool_response> āforestā: āNo specific forest found near the given landmark.ā </tool_response> [Turn 8] [Issue: Parameter Hallucination & Redundant Over-specification] <think> ⦠Alternatively, perhaps the forest is part of a protected area or has a different name ⦠Let me try adjusting the radius or region ⦠So, next step: call ānearby_forest_locatorā again with ⦠a larger radius, say 100 km, to cover more area. </think> <tool_call> ānameā: ānearby_forest_locatorā, āargumentsā: ālandmarkā: āMount Fujiā, āregionā: āJapanā, āradiusā: 100, āinclude_protectedā: true </tool_call> <tool_response> āforestā: āNo specific forest found near the given landmark.ā </tool_response> [Turn 9] [Issue: Parameter Hallucination & Redundant Over-specification & Overconfidence] <think> ⦠Maybe the issue is that the landmark name isnāt precise enough, or the toolās database doesnāt have entries for forests near Mount Fuji ⦠Let me try a different approach ⦠Let me use the āwood_sourcing_locatorā tool with the forest name as āMount Fuji Forestā and product type as āsculpturesā ⦠</think> <tool_call> ānameā: āwood_sourcing_locatorā, āargumentsā: āforest_nameā: āMount Fuji Forestā, āproduct_typeā: āsculpturesā, āregionā: āJapanā, āenvironmental_certificationā: āFSCā </tool_call> <tool_response> No specific woodworker information available for the given forest. </tool_response> [Turn 10] [Issue: Parameter Hallucination & Redundant Over-specification] <think> ⦠Now, since the tools arenāt yielding specific results, maybe the approach needs adjustment ⦠Letās try using the ālocal_handicraft_identifierā tool with Japan as the location, assuming the handicraft is related to woodcarving ⦠</think> <tool_call> ānameā: ālocal_handicraft_identifierā, āargumentsā: ālocationā: āJapanā, ācraft_typeā: āwoodworkā, ācultural_significanceā: true </tool_call> <tool_response> Information on local handicrafts is not available for the specified location. </tool_response> [Generation Terminated: Maximum turn limit (10) reached without providing a final answer.] Table 11: Interaction trajectory of MatchTIR (KM). Although the model successfully reasons out the correct target answer, it exhibits redundancy in the early reasoning stages due to over-specification of geographic constraints. MatchTIR (KM) [Turn 1] [Issue: Parameter Hallucination & Redundant Over-specification] <think> Okay, letās tackle this question step by step ⦠So, the first thing I need to do is find out which mountain has a famous peak. The tool āfamous_peak_identifierā can help with that ⦠Letās assume the mountain is a well-known one, maybe in a region like the Alps or Himalayas ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āregionā: āAlpsā, āfame_typeā: āclimbingā </tool_call> <tool_response> No famous peaks found matching the criteria. </tool_response> [Turn 2] [Issue: Parameter Hallucination & Redundant Over-specification] <think> Okay, the first tool call didnāt return any results. Maybe the mountain isnāt in the Alps. ⦠Letās adjust the region to āHimalayasā and see if that works ⦠Wait, the toolās parameters allow for optional region, so maybe I should specify a region where famous peaks are common ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āregionā: āHimalayasā, āfame_typeā: āscenicā </tool_call> <tool_response> No famous peaks found matching the criteria. </tool_response> [Turn 3] [Issue: Redundant Over-specification] <think> Okay, maybe the issue is with the region. Letās try a different approach. Instead of specifying a region, just use the fame_type as āclimbingā and see if the tool can find a mountain with a famous peak ⦠</think> <tool_call> ānameā: āfamous_peak_identifierā, āargumentsā: āfame_typeā: āclimbingā </tool_call> <tool_response> Mount Everest - The highest mountain in the world, known for its climbing fame. </tool_response> [Turn 4] [Issue: Redundant Over-specification] <think> ⦠The tool returned Mount Everest. Now, the next part of the question is to find the forest near this mountain ⦠I need to use the ānearby_forest_locatorā tool ⦠I should set the landmark to āMount Everestā and maybe specify a radius to search around it. The default radius is 50 km, which should be sufficient ⦠</think> <tool_call> ānameā: ānearby_forest_locatorā, āargumentsā: ālandmarkā: āMount Everestā, āradiusā: 50 </tool_call> <tool_response> āforestā: āLangtang National Park Forestā </tool_response> [Turn 5] <think> Okay, so the forest near Mount Everest is Langtang National Park Forest. Now the next step is to find the woodworker who carves sculptures using wood from this forest ⦠So Iāl call the wood_sourcing_locator with forest_name set to Langtang National Park Forest and product_type as sculptures. </think> <tool_call> ānameā: āwood_sourcing_locatorā, āargumentsā: āforest_nameā: āLangtang National Park Forestā, āproduct_typeā: āsculpturesā </tool_call> <tool_response> Ramesh Thapa is a woodworker who sources wood from Langtang National Park Forest. </tool_response> [Turn 6] <think> ⦠Ramesh Thapa is the woodworker who sources wood from Langtang National Park Forest. The next step is to find out where he is based ⦠The tool that can help with that is the āindividual_location_finderā ⦠The required parameter is individual_name, so Iāl set that to āRamesh Thapaā and maybe include the country as Nepal to narrow it down ⦠</think> <tool_call> ānameā: āindividual_location_finderā, āargumentsā: āindividual_nameā: āRamesh Thapaā, ācountryā: āNepalā </tool_call> <tool_response> ālocationā: āPokharaā </tool_response> [Turn 7] [Issue: Redundant Over-specification] <think> Okay, so now I know that Ramesh Thapa is based in Pokhara. The next step is to find the local handicraft of Pokhara⦠Letās call the function with location: Pokhara. </think> <tool_call> ānameā: ālocal_handicraft_identifierā, āargumentsā: ālocationā: āPokharaā, ācultural_significanceā: true, āmaterials_usedā: true </tool_call> <tool_response> Wooden masks </tool_response> [Turn 8] <think> Okay, so the final answer is that the local handicraft in Pokhara is wooden masks ⦠All steps seem to connect correctly. The answer should be wooden masks. </think> The local handicraft of the town where the woodworker who carves sculptures using wood from the forest near the mountain with a famous peak is wooden masks.