Paper deep dive
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
Linghua Zhang, Jun Wang, Jingtong Wu, Zhisong Zhang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 5:42:18 AM
Summary
RetailBench is a high-fidelity benchmark designed to evaluate the long-horizon autonomous decision-making capabilities of LLM agents in complex, dynamic retail environments. The paper introduces the 'Evolving Strategy & Execution' framework, which decouples strategic reasoning from operational execution to improve stability. Experiments across eight state-of-the-art LLMs reveal that while the proposed framework outperforms existing baselines, current models struggle to scale decision-making as task complexity increases, often failing to maintain coherent strategies over extended horizons.
Entities (4)
Relation Signals (3)
RetailBench → evaluates → LLM Agents
confidence 95% · RetailBench, a high-fidelity benchmark designed to evaluate long-horizon autonomous decision-making
Evolving Strategy & Execution → improves → Operational Stability
confidence 90% · our framework improves operational stability and efficiency compared to other baselines
LLM Agents → operatesin → RetailBench
confidence 90% · We evaluate eight state-of-the-art LLMs in this environment.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Model (LLM)-based agents have achieved notable success on short-horizon and highly structured tasks. However, their ability to maintain coherent decision-making over long horizons in realistic and dynamic environments remains an open challenge. We introduce RetailBench, a high-fidelity benchmark designed to evaluate long-horizon autonomous decision-making in realistic commercial scenarios, where agents must operate under stochastic demand and evolving external conditions. We further propose the Evolving Strategy & Execution framework, which separates high-level strategic reasoning from low-level action execution. This design enables adaptive and interpretable strategy evolution over time. It is particularly important for long-horizon tasks, where non-stationary environments and error accumulation require strategies to be revised at a different temporal scale than action execution. Experiments on eight state-of-the-art LLMs across progressively challenging environments show that our framework improves operational stability and efficiency compared to other baselines. However, performance degrades substantially as task complexity increases, revealing fundamental limitations in current LLMs for long-horizon, multi-factor decision-making.
Tags
Links
- Source: https://arxiv.org/abs/2603.16453v1
- Canonical: https://arxiv.org/abs/2603.16453v1
Trouble viewing inline? Open PDF directly →
Full Text
79,270 characters extracted from source content.
Expand or collapse full text
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments Linghua Zhang 1 Jun Wang 2 Jingtong Wu 2 Zhisong Zhang † 1 Ant Group 2 City University of Hong Kong zlh20011228@gmail.com, nanrong.wj@ant-intl.com jingtong.wujt@ant-intl.com, zhisong.zhang@cityu.edu.hk Abstract Large Language Model (LLM)-based agents have achieved notable success on short-horizon and highly structured tasks, yet their ability to maintain coherent decision-making over long horizons in realistic and dynamic envi- ronments remains an open challenge. We in- troduce RetailBench, a high-fidelity benchmark designed to evaluate long-horizon autonomous decision-making in realistic commercial scenar- ios, where agents must operate under stochastic demand and evolving external conditions. We further propose the Evolving Strategy & Execution framework, which separates high- level strategic reasoning from low-level ac- tion execution, enabling adaptive and inter- pretable strategy evolution over time. This de- sign is crucial for long-horizon tasks, where non-stationary environments and error accumu- lation require strategies to be revised at a differ- ent temporal scale than action execution. Exper- iments on eight state-of-the-art LLMs across progressively challenging environments show that our framework improves operational stabil- ity and efficiency compared to other baselines. However, performance degrades substantially as task complexity increases, revealing funda- mental limitations in current LLMs for long- horizon, multi-factor decision-making. 1 Introduction Recent large language models (LLMs), particularly when augmented with reasoning and tool-use ca- pabilities, have demonstrated strong performance on a variety of cognitively demanding tasks, in- cluding code editing, mathematical problem solv- ing, and complex information retrieval (Jimenez et al., 2024; Phan et al., 2025; Gao et al., 2024). However, accumulating empirical evidence indi- cates that such capabilities fail to generalize into robust, domain-agnostic autonomy, especially in realistic settings requiring long-horizon planning, persistent objective alignment, and stable behav- ioral consistency—core prerequisites for partici- pation in real-world economic systems (Amodei, 2024; Kwa et al., 2025; METR, 2025). Corre- spondingly, existing agent benchmarks—spanning web interaction (Mialon et al., 2023; Wei et al., 2025; Zhou et al., 2024; Deng et al., 2023; Jimenez et al., 2024; Team, 2025b) primarily focus on short- horizon or highly structured tasks, limiting their ability to evaluate sustained interaction with com- plex, evolving environments. Recent studies on long-horizon autonomy consistently demonstrate that even state-of-the-art agents struggle to main- tain coherent strategies over extended time spans (Nof1.ai, 2025; Andon Labs, 2025; Backlund and Petersson, 2025). To systematically study this challenge, we intro- duce RetailBench, a new benchmark grounded in real-world commercial data and informed by estab- lished economic modeling principles.RetailBench centers on a supermarket operation scenario that demands long-horizon decision-making, sustained interaction with a dynamic environment, and the integration of heterogeneous historical informa- tion. The benchmark evaluates whether LLM- based agents can autonomously sustain realistic business operations under complex, multi-factor conditions. We evaluate eight state-of-the-art LLMs in this environment. Our results show that current models struggle to maintain stable decision quality as the decision space expands and often fail to incorporate all relevant information. Moreover, hallucinations and economically irrational behaviors frequently emerge during long-horizon execution, leading to environment collapse and preventing sustained au- tonomous operation. Our contributions are summarized as follows: •We introduce RetailBench, a high-fidelity bench- mark for evaluating long-horizon autonomous 1 arXiv:2603.16453v1 [cs.AI] 17 Mar 2026 RetailBench Environment Intra-day Loop: Multiple steps within day t RL Agent Product & Inventory (Sale / Shelf life / Historial Sales) Financial States Info Query State Triggered Day Transition to day t + 1 External Events Demand Signals Supply Chain Action End-of-Day Pipeline 1. Customer Traffic Sampling 2. Sales Volumn Update 3. Reviews & Returns Genration 4. Inventory Update 5. Financial & Exogeous Update Next Day State Cannot Pay for 5 consecutive days Episode Terminate Actions Price Adjustment Infro Queries Replenishment Order Memory Read / Write end_today Figure 1: Overview of the hierarchical supermarket environment, illustrating intra-day agent–environment interac- tions and end-of-day state transition dynamics. decision-making in realistic retail environments. • We propose the Evolving Strategy & Execution agent framework, which improves operational stability compared to a Reflection-based base- line. •Through extensive experiments, we identify sys- tematic failure modes of current LLM-based agents in long-horizon, multi-factor decision- making settings. 2 Environment Construction 2.1 Problem Formulation and Overview We model supermarket operations as a Markov De- cision Process (MDP), in which an autonomous agent manages a single retail store over a finite horizon of days. At each dayt, the agent makes a sequence of operational decisions that jointly deter- mine the store’s daily outcomes. Formally, the MDP is defined as(S,A,T ,R,γ), whereSdenotes the state space,Athe action space, Tthe stochastic transition dynamics,Rthe reward function, and γ ∈ (0, 1] the discount factor. At the beginning of each dayt, the agent ob- serves the initial stateS t,0 and executes a sequence of intra-day actions indexed byk. Each action a t,k ∈Ainduces a transition to the next intra-day stateS t,k+1 . When the agent issues the end-today action, the environment transitions to the initial state of the next day,S t+1,0 , according to the tran- sition dynamicsT. Figure 1 illustrates the overall interaction process. The environment supports long-horizon opera- tion over more than one thousand simulated days, with episodes terminating if the store fails to pay rent for five consecutive days. 2.2 State Space The intra-day stateS t,k summarizes the complete operational context of the store at stepkof dayt and is composed of multiple interdependent compo- nents: S t,k = S prod t,k ,S inv t,k ,S sup t,k ,S dem t,k ,S ext t,k ,S fin t,k . • S prod t,k andS inv t,k encode product-level attributes and on-hand inventory status, including prices, shelf life, and historical sales records. Product demand is grounded in real-world retail data de- rived from the Dominick’s dataset (Kilts Cen- ter for Marketing). • S sup t,k represents the supply chain state, including supplier prices, quality levels, and delivery lead times, constructed to reflect empirically observed price–quality relationships (Grewal et al., 2014). • S dem t,k captures demand-side signals such as recent customer traffic and aggregated review statistics, 2 which influence consumer purchasing behavior (Fedewa et al., 2021). • S ext t,k represents external contextual informa- tion, including active news events with market-, category-, or product-level impact scope, syn- thesized to resemble real-world financial news (ashraq, 2025). • S fin t,k denotes the financial state of the store, in- cluding available cash and estimated net worth, accounting for inventory depreciation over prod- uct shelf life. Together, these components define a richly struc- tured yet partially stochastic state space. Further details of the environment design are deferred to Appendix A. 2.3 Action Space At each time stept, the agent selects an action a t ∈A, defined as a tuple of decision components: a t = a price t ,a repl t ,a info t ,a mem t ,a end t . Each compo- nent corresponds to a distinct class of operational decisions: • a price t denotes pricing decisions for individual SKUs; • a repl t denotes inventory replenishment decisions, including supplier selection and order quantities; • a info t denotes information acquisition actions, such as querying sales history, customer reviews, supplier conditions, or current news; • a mem t denotes memory operations that allow the agent to write and retrieve persistent notes across days; • a end t denotes a day-termination action that con- cludes the current decision phase. Within a single day, the agent may execute mul- tiple information queries and operational adjust- ments before issuing thea end t action to advance the environment to the next day. 2.4 Day Transition Dynamics After the agent issues the day-termination action, the environment transitions to the next day accord- ing toS t,k+1 ∼T (·| S t,k ,a t ), where the transition functionTfactorizes into a sequence of stochas- tic updates that jointly model daily supermarket operations. Specifically, each day-level transition consists of the following steps: 1. Customer traffic sampling. The total cus- tomer traffic for the day is sampled according to stochastic demand dynamics. 2.Sales realization. Given the sampled customer traffic and exogenous factors (e.g., promotions, seasonal effects, and news signals), the realized sales volume of each product is determined. 3.Reviews and returns generation. Customer reviews and product return events are generated conditional on sales outcomes, product quality, and external influences. 4.Inventory update. Sold units are deducted from on-hand inventory, and replenishment or- ders scheduled to arrive on the current day are added to stock. 5.Financial and exogenous state update. The agent’s financial state is updated based on real- ized revenues and costs, after which new exoge- nous information (e.g., news or market signals) for the next day is generated. 3 Evolving Strategy & Execution Framework 3.1 Framework Detail Prior agent frameworks, including ReAct, Reflec- tion, and Plan-and-Act (Erdogan et al., 2025a; Yao et al., 2023b), either do not maintain a persistent global strategy, leading to inconsistent behaviors, or revise strategies during action execution, which can induce oscillation and gradual goal drift in long-horizon environments. To address these limitations, we propose an explicit two-stage interaction framework, termed Evolving Strategy & Execution, that separates strategic deliberation from operational execution. The framework maintains a persistent global strat- egy and restricts strategy updates to a day-level granularity to prevent excessive short-term fluctua- tions. In the Evolving Strategy stage, the agent may invoke observation and analysis tools to examine environmental feedback and historical outcomes, but cannot execute actions that directly modify the environment. Based on this information, it deter- mines whether to revise the inherited strategy by 3 ModelAvg. Days↑Avg. Daily Sales↑Avg. Daily Income↑Expiry Ratio↓Return Ratio↓Max Days↑ Framework: Evolving Strategy & Execution GLM-4.652.40174.34124.670.07730.129358 Kimi-K2 (Thinking)54.25260.68168.720.02390.117958 GPT-5.281.00457.21358.270.06600.114181 Average (3 models)62.55297.41217.220.05570.120465.67 Framework: Reflection (Day-Level) GLM-4.655.00160.70125.670.01940.117662 Kimi-K2 (Thinking)58.33216.51184.010.09640.125571 GPT-5.264.00283.88260.880.17740.088764 Average (3 models)59.11220.36190.190.09770.110665.67 Framework: Reflection (Step-Level) GLM-4.651.6792.3577.180.01480.107257 Kimi-K2 (Thinking)51.67181.19111.290.03530.136253 GPT-5.256.00398.71324.010.15360.104856 Average (3 models)53.11224.08170.830.06790.116155.33 Framework: Plan-and-Act GLM-4.648.33231.35113.010.01980.109655 Kimi-K2 (Thinking)48.67170.26105.360.00000.105651 GPT-5.264.00323.88193.020.01520.109664 Average (3 models)53.67241.83137.130.01170.108356.67 Heuristic Policy (Upper Bound, Easy) Hand-crafted Policy180.00674.18729.460.02660.0070180 Table 1: Performance comparison of three representative large language models under four agent frameworks in the EASY environment. A hand-crafted heuristic policy is included as an approximate upper bound. adding, refining, or removing components. This stage ensures that long-term intent and planning logic remain explicit rather than implicitly entan- gled with immediate action choices. In the Execution stage, the strategy is fixed and treated as immutable. The agent generates concrete actions strictly consistent with the current strategy, without further modification. This controlled sepa- ration enables clearer attribution between strategy and behavioral outcomes. By alternating between these stages, the frame- work enforces a principled decomposition of decision-making, reduces uncontrolled strategy drift, and promotes behavioral stability and inter- pretability over extended horizons. An illustration is provided in Appendix D.1. 3.2 Hierarchical Policy Representation To support structured, interpretable, and tempo- rally extended decision-making under the proposed framework, we represent the agent policy using a hierarchical abstraction that separates strategic in- tent from executable actions. Each policy consists of three conceptual layers: 1.Macro Strategy, which captures high-level managerial objectives that persist across mul- tiple decision steps; 2. Execution Strategy, which encodes structured operational guidance in a machine-readable in- termediate representation; 3. Daily Actions, which specify concrete exe- cutable operations issued to the environment. Detailed policy configurations and example poli- cies are provided in Appendix A.2.1. 4 Experiment Settings We conduct experiments under three environment configurations with increasing levels of difficulty. These configurations vary in market complexity, budget constraints, and the presence of exogenous dynamics. 4.1 Environment Configurations We employ a heuristic policy with full access to the environment’s internal state as a calibration base- line for each environment variant. Environment pa- rameters are tuned such that the heuristic policy re- mains stable across different difficulty levels while still experiencing meaningful operational pressure 4 in Appendix A.4. To evaluate models’ information- processing and external perception capabilities, we design three environment configurations: •Easy: A controlled environment without dynamic news events or adaptive supplier price–quality relationships. The market contains five product categories. The agent is initialized with a budget of 10,000 and incurs a fixed daily rent of 250. • Middle: A moderately complex environment that expands the product space to all twenty cate- gories while still excluding dynamic news events and supplier adaptations. The initial budget is increased to 50,000, with a daily rent of 1,000. •Hard: The most challenging and realistic environ- ment, incorporating dynamically generated news events and time-varying supplier price–quality relationships. The market includes all twenty product categories. The agent starts with a bud- get of 50,000, pays a daily rent of 1,000, and receives twenty news items per day. Detailed specifications of all environment configu- rations are provided in Appendix A.3. 4.2 Evaluation Metrics Metrics.We evaluate store-level operational per- formance using the following metrics (↑indicates higher is better;↓ indicates lower is better): •Days (↑): the number of operating days before episode termination; • MaxDays (↑): the maximum number of operating days achieved across three rollouts; • Avg. Daily Sales (↑): the average number of items sold per day; • Avg. Daily Income (↑): the average money earned per day; •Expiry Ratio (↓): the fraction of products that expire before being sold; • Return Ratio (↓): the fraction of sold products that are returned by customers. All reported metrics are averaged over three inde- pendent rollouts, each subject to a fixed maximum execution horizon. 4.3 Experimental Setup Agent Frameworks.Preliminary experiments in- dicate that simple ReAct-style interaction frame- works are unstable in long-horizon settings, fre- quently exhibiting premature episode termination during mid-horizon execution. This observation motivates the evaluation of agent frameworks that explicitly support sustained strategic control. We compare our proposed Evolving Strategy & Execution framework with step-level Reflection, day-level Reflection, and the Plan-and-Act base- line (Shinn et al., 2023; Erdogan et al., 2025a). For the Reflection framework, we implement two variants with different reflection frequencies: (1) step-level reflection, where the agent reflects after each action, and (2) day-level reflection, where re- flection is performed once per day after completing daily actions. Full prompt specifications for all frameworks are provided in Appendix C. Models and Evaluation Protocol. We evalu- ate a diverse set of contemporary large language models, including Qwen-235B (Thinking) (Team, 2025a), Kimi K2 (Thinking) (Team et al., 2025b), GLM-4.6 (Team et al., 2025a), DeepSeek-V3.2- Exp (DeepSeek-AI et al., 2025), Gemini-3-Flash- Preview (Google DeepMind, 2024), Grok-4.1 Fast (xAI, 2025), GPT-5-Mini (OpenAI, 2025a), and GPT-5.2 (OpenAI, 2025b). To manage the substan- tial token costs of long-horizon rollouts, lower-cost variants are used for closed-source models when available. Each model is evaluated across three environ- ment configurations, with three independent roll- outs per configuration; results are averaged across runs. To isolate the effect of the agent framework, we conduct controlled comparisons between our framework and baseline methods in the Easy en- vironment under identical conditions. We further implement a hand-crafted policy based on privi- leged internal knowledge unavailable to the agents, which serves as an approximate upper bound on performance. 5 Results 5.1 Performance Comparison between Agent Frameworks Due to computational budget constraints, we con- duct a comprehensive four-framework comparison on three representative models: GLM-4.6, Kimi- 5 Grok-4.1 FastGemini-3 (Fast)DeepSeek-V3.2 (Exp.) Model 0 20 40 60 80 100 Value Per Category 101.6 67.4 36.0 19.9 25.2 18.2 79.9 58.9 22.0 22.5 25.7 32.5 45.8 36.7 13.2 12.8 8.3 8.8 Sales and Income Per Category Comparison Easy (Sales) Easy (Income) Middle (Sales) Middle (Income) Hard (Sales) Hard (Income) Figure 2: Category-level sales and profit per category across Easy, Middle, and Hard environments. Results are shown for three representative models. K2 (Thinking), and GPT-5.2. As reported in Ta- ble 1, Evolving Strategy & Execution consistently outperforms alternative agent frameworks on core metrics, achieving higher sales and profit while substantially reducing product expiration rates. Based on these results, we select the two strongest-performing frameworks for broader eval- uation across all eight models. When averaged over all models, our proposed framework demonstrates clear improvements over Reflection (Day-Level). Despite these gains, a substantial gap remains relative to the hand-crafted heuristic policy, which serves as an approximate upper bound. This gap highlights the current limitations of LLM-based agents in long-horizon decision-making. Notably, Table 1 also suggests that larger closed-source mod- els, such as GPT-5.2, tend to achieve more stable and sustained store operation compared to smaller or open-source counterparts, indicating that model capacity remains an important factor alongside framework design. Table 8 presents the complete results across all eight evaluated models. Across models, we ob- serve recurring and systematic failure modes in the current environment, suggesting that the perfor- mance bottlenecks are structural rather than model- specific. 5.2 Performance across Environments with Varying Difficulty As environment difficulty increases, all models ex- hibit consistent performance degradation, includ- ing shorter operational durations, higher expiry and return ratios, and reduced long-horizon stability. Although the expansion of product categories in the Middle and Hard settings enables higher aggre- gate sales and profit for some models, Sales per Category and Profit per Category decline substan- ModelContext EasyMiddleHard SKU↑Cat↑SKU↑Cat↑SKU↑Cat↑ DeepSeek-V3.2 (Exp.)128K4.753.316.835.346.645.45 Gemini-3 (Fast)1M8.724.3413.6012.4321.5713.56 GLM-4.6200K4.822.983.232.467.08 5.22 OpenAI-5.1 Mini400K3.661.944.104.106.674.79 Grok-4.1 Fast200K5.783.684.873.085.403.20 Kimi-K2 (Thinking)256K5.063.095.174.776.914.75 Qwen-235B (Thinking)256K4.593.673.172.365.114.51 Heuristic Strategy–9.034.8735.3418.6134.8218.52 Table 2: SKU and Category counts sold by different models across difficulties. Context denotes the maxi- mum supported context length of each model. Bolded indicates the best model;underlinedindicates the sec- ond best (Human excluded). tially (Figure 2), indicating persistent challenges in effective resource allocation within increasingly high-dimensional decision spaces. The relatively small performance gap between the Middle and Hard environments can be partly attributed to the delayed impact of intensified news dynamics, whose effects typically unfold over longer time horizons. Appendix B.3 reports the full results across all environments under the proposed framework. Overall, while models demonstrate limited adaptability as task complexity increases, their performance remains substantially below the heuristic upper bound, highlighting persistent limi- tations in robust long-horizon decision-making un- der complex and dynamic conditions. 6 Analysis Based on both quantitative evaluation results and manual inspection of trajectories, we identify sev- eral key factors that explain why current models fail to operate reliably in the environment. We analyze these failure modes from four com- plementary perspectives. 6.1 Non-scalable Decision-Making Capability As shown in Appendix B.3, models achieve rea- sonable performance in the Easy setting but exhibit consistent degradation as the environment scales in complexity. To better understand this phenomenon, we analyze both the average number of stock keep- ing units (SKUs) and product categories sold per day, as well as the number of SKUs and categories explicitly considered during the Evolving Strategy phase (Tables 2 and Appendix B.6). Across most models, decision-making capabil- ity does not scale proportionally with the size of the environment. Instead, performance remains relatively flat despite a substantial increase in the number of available options. Notably, Gemini-3 6 ModelSupplierInventoryReturnRatingPriceReviewSalesHistory DeepSeek-V3.2 (Exp.)41.4100.026.775.46.60.689.20.0 Gemini-3 (Fast)69.6100.041.082.367.76.387.62.9 GLM-4.693.998.379.877.884.15.196.30.0 OpenAI-5.1 Mini58.899.05.178.63.910.684.70.3 Grok-4.1 Fast89.999.884.094.088.636.096.40.3 Kimi-K2 (Thinking)57.290.425.951.436.211.064.00.5 Qwen-235B (Thinking)76.878.331.558.219.823.674.60.0 Average69.795.142.074.043.813.384.71.0 Table 3: Percentage of days on which each model queries specific information sources when making de- cisions for SKUs included in the final daily strategy. Higher values indicate greater reliance on the corre- sponding data source. (Fast), which supports the largest maximum con- text window, performs better than other models in this respect, suggesting that larger context capac- ity helps retain salient information even when the effective interaction context is constrained. Nevertheless, even the strongest models fail to cover the full decision space. This indicates that current systems are unable to expand their effective decision-making scope as the environment grows, leading to systematic performance degradation in larger and more complex settings. 6.2 Incomplete Decision-Making Due to Limited Information Coverage We analyze the information sources that models attend to when making operational decisions by examining the data queried for the set of SKUs included in each day’s final strategy. This analy- sis reveals a clear concentration of attention on a limited subset of signals. Specifically, most models primarily rely on sup- plier prices, inventory levels, SKU ratings, and historical sales records when deciding how to oper- ate selected SKUs in Table 3. In contrast, several other critical signals—such as recent customer re- views, return rates, and current selling prices—are consistently underutilized or entirely ignored. Further correlation analysis shows a strong posi- tive relationship between the frequency of SKU re- views queries and average daily sales performance in Appendix B.5. This observation aligns with the underlying environment dynamics and indi- cates that incomplete information coverage is a key factor limiting decision quality. Overall, these results suggest that models often fail to perform sufficiently comprehensive information gathering, leading to systematically suboptimal operational decisions. Macro StrategyExecution Strategy ModelStd_diff (↓)MAC (↓)TV (↓)Std_diff (↓)MAC (↓)TV (↓) DeepSeek-V3.2(Exp.)0.1300.0905.000.2890.1447.93 Gemini-3 (Fast)0.1360.0853.820.2930.22510.18 GLM-4.60.1310.0964.520.2640.2099.87 OpenAI-5.1 Mini0.0690.0512.500.2290.1668.00 Grok-4.1 Fast0.0790.0452.710.2040.1519.39 Kimi K2(Thinking)0.1790.1336.950.3350.27114.08 Qwen-235B (Thinking)0.2400.1896.510.3910.30610.57 Table 4: Temporal instability metrics of macro- and execution-level strategy similarity in the Easy environ- ment. All metrics are lower-is-better (↓). Larger values indicate greater temporal instability. Column-wise max- imum values are highlighted in bold. 6.3 Temporal Instability in Execution-Level Decision-Making Beyond limitations in decision scalability and infor- mation coverage, we identify temporal instability in execution-level decision-making as a key contribu- tor to long-horizon failure. Even under relatively stable environmental conditions, agents frequently revise their strategies across consecutive days, lead- ing to inconsistent execution trajectories. Measuring Strategy Similarity. We quantify temporal instability by measuring the similarity be- tween strategies on adjacent days at both macro and execution levels.Macro strategy similar- ity is assessed using an LLM-based prompt that evaluates semantic consistency between consec- utive high-level plans. Execution strategy sim- ilarity is computed via set-based Jaccard sim- ilarity over key fields, includingfocus_skus, sku_supplier_mapping,news_to_monitor, and sku_to_monitor, with the final execution similar- ity obtained by averaging across fields. Instability Metrics. To characterize temporal fluctuations, we compute three complementary met- rics: the standard deviation of first-order differ- ences (Std_diff ) to capture short-term volatility, the mean absolute change (MAC) to estimate typical day-to-day variation, and total variation (TV ) to reflect cumulative long-term instability. Across all three environment configurations, we observe consistent temporal instability in both macro- and execution-level strategies. For clarity, we present detailed results from the Easy environ- ment in Table 4 and Appendix B.4 . Despite its reduced complexity, the Easy setting already ex- hibits pronounced temporal fluctuations. In partic- ular, macro strategies remain relatively stable over time, whereas execution strategies show substan- tially larger variability across most models. This effect is especially pronounced for Qwen-235B 7 (Thinking), which also performs poorly across all environments. Overall, these results indicate that long-horizon failures arise not only from suboptimal strategy formulation, but more fundamentally from the in- ability to maintain temporally consistent execution policies over extended horizons. 6.4 Hallucinations and Invalid Actions Finally, manual inspection reveals recurrent failure patterns that directly break planning correctness and action validity. We distinguish two closely related but practically different issues: Hallucinations in reasoning Models occasion- ally generate reasoning traces that reference non- existent SKUs or fabricate numerical quantities, and subsequently incorporate these hallucinated elements into multi-step plans (examples are pro- vided in Appendix B.1). As a result, decisions become misaligned with the true environment state, even when the overall planning structure appears internally coherent. Invalid or irrational actions.Models sometimes output actions that violate basic constraints or are inconsistent with historical demand, such as nega- tive order quantities or implausible pricing in Ap- pendix B.2. Even though state-modifying oper- ations are restricted to only a small set of tools, these invalid actions occur with non-negligible fre- quency and can quickly destabilize the system in long-horizon operation. In realistic operational settings, both hallucina- tions and invalid actions would be unacceptable and could directly trigger system collapse. 7 Related Work Long-horizon planning benchmarks for LLMs. In parallel, a growing body of benchmarks targets long-horizon planning abilities of large language models across structured and interactive environ- ments. PlanBench focuses on classical plan gen- eration, while WebShop, Mind2Web, and Science- World (Deng et al., 2023; Wang et al., 2022; Yao et al., 2023a) evaluate multi-step interaction, er- ror recovery, and adaptive behavior in dynamic settings. More recent benchmarks—such as Her- oBench, OdysseyBench, UltraHorizon (Luo et al., 2025; Anokhin et al., 2025; Wang et al., 2025), and explicitly stress extended, interdependent decision sequences that require persistent memory, hierar- chical reasoning, and long-term strategy mainte- nance. Collectively, these benchmarks suggest that while short-horizon tasks are increasingly tractable, robust long-horizon planning and execution remain a central challenge for LLM-based agents. Long-horizon agent frameworks. Recent ad- vances in agent design address these challenges by introducing structured frameworks that de- couple high-level planning from low-level execu- tion. Approaches such as Plan-and-Act (Erdogan et al., 2025b) and EAGLET (Si et al., 2025) adopt planner–executor architectures to support hierar- chical decision-making, dynamic replanning, and improved execution stability over long horizons. Extensions to multi-agent settings ELHPlan (Ling et al., 2025) and plan-aware context management frameworks PAACE (Yuksel, 2025) further em- phasize task decomposition, explicit strategy rep- resentation, and memory-aware context control. Together, these works indicate that scalable long- horizon autonomy increasingly relies on structured planning modules and controlled execution mecha- nisms rather than monolithic end-to-end policies. 8 Conclusion This paper introduces RetailBench, a benchmark for evaluating long-horizon autonomous decision- making in realistic retail environments. Retail- Bench models supermarket operations as a stochas- tic, multi-factor, and temporally extended process, requiring agents to reason jointly about pricing, inventory, information acquisition, and financial sustainability. We further propose the Evolving Strategy & Execution framework, which decouples high-level strategy evolution from low-level execu- tion to better support long-horizon autonomy. Experiments across eight state-of-the-art large language models show that our framework im- proves operational stability and economic perfor- mance compared to a Reflection-based baseline. However, performance degrades sharply as envi- ronment complexity increases, revealing persistent limitations in scalability, information use, execu- tion stability, and action validity. These results suggest that while structured agent frameworks mit- igate some challenges, current LLM-based agents remain far from robust strategy-aware autonomy in complex dynamic environments. RetailBench thus provides a principled testbed for advancing re- search on long-horizon decision-making and agen- tic reasoning. 8 9 Limitations Despite its realism and scale, this work has sev- eral limitations. First, RetailBench focuses on a single-store supermarket setting; while expressive, it does not capture multi-store coordination, com- petitive markets, or strategic interactions among multiple autonomous agents. Second, although the environment incorporates stochastic demand, news dynamics, and supply-chain delays, it remains a simulation grounded in historical data and simpli- fied economic assumptions, which may not fully reflect the complexities of real-world retail sys- tems. Third, our evaluation is limited to prompting- based LLM agents without parameter updates or long-term learning across episodes; stronger perfor- mance may be achievable through reinforcement learning, fine-tuning, or hybrid neuro-symbolic ap- proaches. Finally, while we identify key failure modes such as hallucinations and economically ir- rational actions, we do not propose explicit algorith- mic mechanisms to enforce economic constraints or factual grounding during execution. Address- ing these limitations—through richer environments, multi-agent extensions, learning-based adaptation, and constraint-aware action control—remains an important direction for future research. References Dario Amodei. 2024.Machines of loving grace.https://w.darioamodei.com/essay/ machines-of-loving-grace. Accessed: 2025-12- 01. Andon Labs. 2025.Vending-bench 2: A bench- mark for long-horizon business simulation.https: //andonlabs.com/evals/vending-bench-2. Ac- cessed: 2025-12-10. Petr Anokhin, Roman Khalikov, Stefan Rebrikov, Viktor Volkov, Artyom Sorokin, and Vincent Bissonnette. 2025. Herobench: A benchmark for long-horizon planning and structured reasoning in virtual worlds. Preprint, arXiv:2508.12782. ashraq. 2025.financial-news-articles.https: //huggingface.co/datasets/ashraq/ financial-news-articles.Accessed:2025- 12-01. Axel Backlund and Lukas Petersson. 2025. Vending- bench: A benchmark for long-term coherence of au- tonomous agents. Preprint, arXiv:2502.15840. DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chen- hao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, and 245 others. 2025. Deepseek-v3.2: Pushing the frontier of open large language models. Preprint, arXiv:2512.02556. Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2web: Towards a generalist agent for the web. Preprint, arXiv:2306.06070. Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025a. Plan-and-act: Improving planning of agents for long-horizon tasks. Preprint, arXiv:2503.09572. Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025b. Plan-and-act: Improving planning of agents for long-horizon tasks. Preprint, arXiv:2503.09572. Dave Fedewa, Chris Holder, Wynn Teichner, and Ben Wiseman. 2021. Five-star growth: Using online rat- ings to design better products. McKinsey & Company. Accessed: 2025-11-30. Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024. Omni-math: A univer- sal olympiad level mathematic benchmark for large language models. Preprint, arXiv:2410.07985. Google DeepMind. 2024. Gemini 3 flash (preview). https://deepmind.google/technologies/ gemini/.Large multimodal language model, preview version. Dhruv Grewal, Jens Nordfält, Anne Roggeveen, Rainer Olbrich, and Hans Christian Jansen. 2014. Price- quality relationship in pricing strategies for private labels. Journal of Product and Brand Management, 23(6):429–438. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. Swe-bench: Can language mod- els resolve real-world github issues?Preprint, arXiv:2310.06770. University of Chicago Booth School of Business Kilts Center for Marketing. Dominick’s dataset. https://w.chicagobooth.edu/research/ kilts/research-data/dominicks.Accessed: 2025-11-01. Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, 9 Seraphina Nix, Tao Lin, Neev Parikh, and 6 others. 2025. Measuring ai ability to complete long tasks. Preprint, arXiv:2503.14499. Shaobin Ling, Yun Wang, Chenyou Fan, Tin Lun Lam, and Junjie Hu. 2025. Elhplan: Efficient long-horizon task planning for multi-agent collaboration. Preprint, arXiv:2509.24230. Haotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang, Zeyu Qin, Wenjie Lu, Guozheng Ma, Haiying He, Yingsha Xie, Qiyang Zhou, Zixuan Hu, Hongze Mi, Yibo Wang, Naiqiang Tan, Hong Chen, Yi R. Fung, Chun Yuan, and Li Shen. 2025. Ultrahori- zon: Benchmarking agent capabilities in ultra long- horizon scenarios. Preprint, arXiv:2509.21766. Daniel McFadden. 1974. Conditional logit analysis of qualitative choice behavior. In Paul Zarembka, editor, Fontiers in Econometrics, pages 105–142. Academic press, New York. METR. 2025. Measuring ai ability to complete long tasks. METR blog. Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. Gaia: a benchmark for general ai assistants. Preprint, arXiv:2311.12983. Nof1.ai. 2025. Alpha arena — exploring the limits of large language models as quant traders.https: //nof1.ai/blog/TechPost1. Accessed: 2025-12- 10. OpenAI. 2025a.Gpt-5 mini.https://platform. openai.com/docs/models/gpt-5-mini.Large language model — cost-efficient GPT-5 variant. OpenAI. 2025b. Introducing gpt-5.2. Accessed: 2026- 02-19. Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, and 1093 others. 2025. Hu- manity’s last exam. Preprint, arXiv:2501.14249. Noah Shinn, Federico Cassano, Edward Berman, Ash- win Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal rein- forcement learning. Preprint, arXiv:2303.11366. Shuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen, Fanchao Qi, Minjia Zhang, Baobao Chang, and Maosong Sun. 2025. A goal without a plan is just a wish: Efficient and effective global plan- ner training for long-horizon agent tasks. Preprint, arXiv:2510.05608. GLM Team, Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, Kedong Wang, Lucen Zhong, Mingdao Liu, Rui Lu, Shulin Cao, Xiaohan Zhang, Xuancheng Huang, Yao Wei, and 152 others. 2025a. Glm-4.5: Agentic, reason- ing, and coding (arc) foundation models. Preprint, arXiv:2508.06471. Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jia- hao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, Jialei Cui, Hao Ding, Mengnan Dong, Angang Du, Chen- zhuang Du, Dikang Du, Yulun Du, Yu Fan, and 150 others. 2025b. Kimi k2: Open agentic intelligence. Preprint, arXiv:2507.20534. Qwen Team. 2025a. Qwen3 technical report. Preprint, arXiv:2505.09388. The Terminal-Bench Team. 2025b. Terminal-bench: A benchmark for ai agents in terminal environments. Mark D. Uncles. 1987. Discrete choice analysis: Theory and application to travel demand. Journal of the Operational Research Society, 38(4):370–371. Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu. 2022. Scienceworld: Is your agent smarter than a 5th grader?Preprint, arXiv:2203.07540. Weixuan Wang, Dongge Han, Daniel Madrigal Diaz, Jin Xu, Victor Rühle, and Saravan Rajmohan. 2025. Odysseybench: Evaluating llm agents on long-horizon complex office application workflows. Preprint, arXiv:2508.09124. Jason Wei, Zhiqing Sun, Spencer Papay, Scott McK- inney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. 2025. Browsecomp: A simple yet chal- lenging benchmark for browsing agents. Preprint, arXiv:2504.12516. xAI. 2025. Grok 4.1 fast and agent tools api.https:// x.ai/news/grok-4-1-fast. Accessed [Your Ac- cess Date]. Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2023a. Webshop: Towards scalable real-world web interaction with grounded language agents. Preprint, arXiv:2207.01206. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023b. React: Synergizing reasoning and acting in language models. Preprint, arXiv:2210.03629. Kamer Ali Yuksel. 2025.Paace: A plan-aware automated agent context engineering framework. Preprint, arXiv:2512.16970. Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Gra- ham Neubig. 2024. Webarena: A realistic web envi- ronment for building autonomous agents. Preprint, arXiv:2307.13854. 10 A Environment Configuration Details A.1 State Space Decomposition We construct the environment using real-world re- tail data from the Dominick’s dataset (Kilts Cen- ter for Marketing). From the 20 available product categories, we select 96 SKUs with the most com- plete and informative sales records. The richness of these observations enables reliable modeling of the relationship between pricing decisions and realized demand. Key Symbols. LetJdenote the set of selected SKUs, with cardinality|J| = J(whereJ = 96in our experiments). Each SKUj ∈ Jbelongs to a product categorycat(j) ∈ 1,..., 20. For each SKUj, letK(j)denote its associated supplier set, with|K(j)| = 5. A.1.1 Product State S prod t For each SKUj, we maintain a set of static at- tributes together with time-varying operational sig- nals: S prod t,j = id j , desc j , L j , p jt , h sales j,t , r j,t . (1) Here,id j anddesc j denote the unique identifier and textual description of SKUj, respectively.L j denotes the shelf life (in days), andp jt is the retail price on dayt. The termh sales j,t encodes historical sales statistics, whiler j,t summarizes aggregated customer review signals, such as average rating and review volume. Data grounding. The SKU set and cost priors are constructed from the Dominick’s dataset (Kilts Center for Marketing). Historical sales records from the same source are used to calibrate the de- mand model. A.1.2 Inventory State S inv t Inventory is represented at the SKU level with age tracking. LetI jt denote the on-hand units of SKU jat the beginning of dayt. To support shelf-life constraints and depreciation, we track the arrival time and remaining shelf life of each unit. Capacity constraint and pending queue. Let Capdenote the total inventory capacity of the store. If incoming replenishment orders would violate this constraint, excess units are placed into a first- in-first-out (FIFO) pending queueQ t and become available only after sufficient inventory space is released: X j∈J L j X a=0 I (a) jt ≤ Cap.(2) Stockout signal.When realized demand exceeds the available sellable inventory of SKUjon day t, a stockout indicator is triggered. This signal is recorded at the end of daytand exposed to the agent as explicit operational feedback. Expiration and destruction.At each day transi- tion, inventory units whose residence time exceeds their shelf life are removed from the system. A.1.3 Supply Chain State S sup t Each SKUjis associated with a set of suppliers k ∈K(j). The supplier state includes procurement pricec jk , quality levelq jk , and lead-time distribu- tion: S sup t,j = (c jk ,q jk ,L jk ) : k ∈K(j) .(3) Lead timeℓ jk ∼ L jk is sampled when an order is placed. Procurement prices are discretized into tiers using Dominick’s cost signals (Kilts Center for Marketing), following empirically observed posi- tive correlations between price tier and quality level (Grewal et al., 2014). Enforced diversity. To avoid degenerate sup- plier configurations, we enforce that each SKU has (i) one supplier with maximal quality, (i) one supplier with minimal price and minimal qual- ity, and remaining suppliers spanning intermediate price–quality levels. A.1.4 Demand Signals and Demand Generation Demand-related components capture both observ- able market signals and the stochastic process gov- erning realized demand. We define S dem t = N t , r j,t j∈J ,(4) where N t denotes daily customer traffic. Customer traffic. Daily customer traffic is de- rived from the Dominick’s dataset. Consumer choice and demand generation. Given customer trafficN t , pricesp jt j∈J , review signalsr j,t , and external newsE t , we generate re- alized demand using a discrete-choice model. Fol- lowing standard practice in empirical economics, 11 we adopt a Multinomial Logit (MNL) framework (McFadden, 1974; Uncles, 1987). For SKU j, the raw utility is defined as U raw jt = α j + β j p jt + ε jt ,(5) whereα j captures intrinsic preference,β j models price sensitivity, andε jt represents idiosyncratic shocks. We further incorporate review effects, news sig- nals, and within-category substitution: ̃ U jt = U raw jt + ∆ j (r j,t ,E t ) + X i̸=j cat(i)=cat(j) γ ji exp( ̃ U it ). (6) With the outside option utility normalized to zero, the purchase probability for SKUjon dayt is p ∗ jt = exp( ̃ U jt ) 1 + P J i=1 exp( ̃ U it ) .(7) Realized demand.Given trafficN t and purchase probabilitiesp ∗ jt , potential demand is sampled as y jt ∼ Binomial(N t , p ∗ jt ),(8) and realized sales are capped by available on-hand inventory. A.1.5 External Information S ext t (News Module) We maintain a set of active news eventsE t . Each event e∈E t is represented as e = (type, scope, target, side, sign, η, text, ttl) (9) wherescope∈ macro, category, product, neutral ,side ∈ demand, supply, both,sign ∈ +1,−1, ηdenotes impact magnitude, andttlis the time-to-live in days. News texts are synthesized using an LLM to resemble financial news corpora (ashraq, 2025). Impact application. News impacts are incorpo- rated into (i) demand utilities and/or (i) supplier prices depending onsideandscope. Neutral news has η = 0 by design. A.1.6 Financial State S fin t LetF t denote available funds at the start of dayt. Net worth is computed as NW t = F t + X j∈J L j X a=0 I (a) jt · v j (a),(10) wherev j (a)denotes the per-unit value at agea. We adopt linear shelf-life depreciation: v j (a) = c j · max 0, 1− a L j ,(11) withc j denoting a reference procurement cost (e.g., the mean or realized supplier cost). A.2 Policy Detail A.2.1 Policy Presentation To support structured and interpretable decision making, we represent the agent policy using a hier- archical abstraction that separates strategic intent from executable actions. Specifically, each policy is composed of three layers: 1.Macro Strategy, which captures high-level managerial principles that persist across days; 2.Execution Strategy, which encodes struc- tured operational guidance in a machine- readable form; 3.Daily Actions, which enumerate concrete ex- ecutable operations submitted to the environ- ment. This design enables the agent to reason at dif- ferent temporal and semantic granularities, while maintaining a clear separation between planning and execution. Macro Strategy.The macro strategy consists of a set of natural-language statements that describe high-level objectives (e.g., prioritizing inventory turnover or focusing on high-margin products). These statements are non-executable and serve as persistent guidance for downstream decision mak- ing. Execution Strategy. The execution strategy consists of six key components.focus_skus specifies the SKUs that require immediate attention.sku_supplier_mappingdenotes the corresponding suppliers for each SKU. news_to_monitoridentifies relevant news signals 12 to track.skus_to_reorderindicates SKUs that require replenishment.price_adjustmentspec- ifies the SKUs whose prices should be adjusted along with the corresponding adjustment magni- tudes.sku_to_monitordenotes SKUs under ob- servation. Theotherfield captures additional exe- cution directives. Daily Actions.Daily actions consist of two oper- ation types:place_order, which places purchase orders for selected SKUs, andmodify_sku_price, which adjusts the prices of specified SKUs. Strategy–Execution Protocol.Policy usage fol- lows a two-phase protocol. In the strategy phase, the agent analyzes the environment and constructs its macro and execution strategies. In the subse- quent execution phase, the finalized strategy is con- sumed to generate executable daily actions. The strategy is immutable during execution, enforcing a clear separation between deliberation and action. A.2.2 Policy Example See in Table 5. A.3 Environment Setting A.3.1 Easy Environment Configuration TheEasyconfiguration is designed for simpler sce- narios with limited resources and a reduced cate- gory set. • Time Range: – Data begin time: 06/06/91 – Data end time: 12/31/95 – Store begin time: 09/07/91 – Store ID: 15 • Financial Parameters: – Initial funds: 10,000 – Daily rent: 250 – Inventory capacity: 10,000 • Feature Enablement: – Review enabled: True – Review ratio: 0.02 – News enabled: False • Selected Categories (5 categories): – Bathroom_Tissues – Canned_Soup – Cigarettes – Front_end_candies – Soft_Drinks •Category Effects: All categories have a uni- form effect of -0.2 • Random Seed: 42 (for reproducibility) A.3.2 Middle Environment Configuration TheMiddleconfiguration uses static data but with full resource capacity and all categories enabled. • Time Range: Same as Easy • Financial Parameters: – Initial funds: 50,000 – Daily rent: 1,000 – Inventory capacity: 40,000 • Feature Enablement: – Review enabled: True – Review ratio: 0.02 – News enabled: False • Selected Categories (20 categories): –Bathroom_Tissues, Beer, Bottled_Juices, Canned_Soup, Canned_Tuna – Cereals, Cheeses, Cigarettes, Cookies, Crackers –Dish_Detergent,Fabric_Softeners, Front_end_candies, Frozen_Entrees –Frozen_Juices, Oatmeal, Paper_Towels, Snack_Crackers – Soft_Drinks, Toothpastes •Category Effects: All categories have a uni- form effect of -0.2 • Random Seed: 42 A.3.3 Hard Environment Configuration TheHardconfiguration uses dynamic data with news events enabled, providing the most complex simulation environment. • Time Range: Same as other configurations • Financial Parameters: – Initial funds: 50,000 – Daily rent: 1,000 – Inventory capacity: 40,000 13 • Feature Enablement: – Review enabled: True – Review ratio: 0.02 – News enabled: True • News Configuration: – News impact base scale: 0.4 – News daily count: 20 – News random seed: 42 – News sample ratios: * Neutral: 0.9 * Single category: 0.02 * Macro all: 0.03 * SKU level: 0.05 – News impact mode weights: * Neutral: 0.0 * Macro all: 1.0 * Single category: 1.0 * SKU level: 1.2 •Selected Categories: Same 20 categories as Middle •Category Effects: All categories have a uni- form effect of -0.2 • Random Seed: 42 A.4 Environment Simulation See in Figure 3 B Rollout Details B.1Hallucinations in Reasoning and Planning This appendix provides representative examples of hallucinations observed during long-horizon roll- outs, focusing on reasoning-time errors that prop- agate into multi-step planning despite internally coherent logic. B.1.1 Non-existent SKUs During execution, models occasionally reference or plan around SKUs that do not exist in the envi- ronment. Across all rollouts, we identify 14 dis- tinct non-existent SKUs that repeatedly appear in final daily strategies, indicating persistent hallu- cinations rather than isolated parsing or format- ting errors. Representative examples include SKU identifiers such as10700013100,166051312, and 440004627. 020406080100120140160180 Days 10000 20000 30000 40000 50000 60000 70000 80000 90000 Net worth Net worth over time easy hard middle 020406080100120140160180 Days 0 10000 20000 30000 40000 50000 60000 70000 Money balance Money balance over time easy hard middle Figure 3: Net worth and available funds trajectories of the heuristic policy under different environment config- urations, illustrating the calibrated difficulty levels. The Middle and Hard settings involve a larger number of product categories, enabling higher potential net worth and cash accumulation over the course of an episode. Example. In the following strategy output, the model includes a non-existent SKU (440004627) in its set of focus SKUs. This hallucinated iden- tifier is subsequently treated as a valid product in downstream reasoning and planning: "day": 14, "current_date": "1991-09-22", "strategy": "macro_strategy": ***, "execute_strategy": "focus_skus": [ "3000001460", "3700063037", "440004627", "5100001251" ], "news_to_monitor" ***, , "today_action": ***, 14 B.1.2 Hallucinated Dates Models also hallucinate temporal information dur- ing tool-based reasoning when required to provide calendar dates that are not explicitly specified by the environment. Instead of querying the current simulation date, the model resolves this underspec- ification by fabricating a plausible calendar map- ping in order to proceed with execution. Example. The following excerpt illustrates the model’s reasoning process when attempting to is- sue a tool call that requires valid date strings: Tool requires dates in Y-M-D or M/D/Y. Only day index is known: Day 14. No calendar start date is provided. Assume Day 1 = 2023-01-01. Infer Day 14 = 2023-01-14. start_date = 2023-01-13 end_date = 2023-01-14 Based on this fabricated assumption, the model issues a syntactically valid tool call using the in- ferred dates. Although the reasoning chain is inter- nally consistent, the assumed calendar mapping is not grounded in the true environment state, result- ing in a semantically incorrect interaction. These examples highlight a recurring failure mode in which models resolve underspecified vari- ables through plausible fabrication rather than ex- plicit information acquisition, leading to planning trajectories that appear coherent yet diverge from the actual environment dynamics. B.2 Invalid or Irrational Actions We identify invalid or economically irrational ac- tions by applying a set of heuristic rules to model- issued tool calls. Specifically, we flag actions that violate basic economic or operational constraints, including: (i) setting product prices to zero, nega- tive values, or unrealistically high levels (e.g., ex- ceeding 50); and (i) placing orders with quantities far beyond plausible operational scales for a single SKU. Across all evaluated rollouts, we detect more than 300 such anomalous tool calls. Among them, 197 correspond to invalid or irrational price modifi- cations, while 125 involve excessively large order- ing decisions. These behaviors are observed across multiple models and environment configurations. Excessive Ordering. The following example is taken from a rollout of the Kimi K2 Thinking model in the Hard environment. The model places an implausibly large order for a single SKU, far ex- ceeding realistic replenishment volumes: tool: place_order model: Kimi K2 Thinking sku_id: 5100000011 quantity: 18000 Although syntactically valid, such actions intro- duce extreme inventory shocks that are misaligned with realistic retail operations. InvalidPriceModifications. Irrational pricing behavior is also observed across models.In the following example,the model updates the price of a SKU by sev- eral orders of magnitude (log excerpted from still_hard/kimi_thinking/tool_calls.jsonl): tool: modify_sku_price model: Kimi K2 Thinking old_price: 0.25 new_price: 999.0 In other cases, models attempt to assign zero or negative prices, which are explicitly rejected by the environment. While such invalid actions are prevented at execution time, extreme but valid val- ues can still propagate downstream and destabilize subsequent decision-making. Overall, these results indicate that current LLM- based agents lack reliable internal mechanisms for enforcing basic economic plausibility at the action level, even when tool interfaces enforce syntactic validity and detailed execution logs are available for inspection. B.3 Rollout Results See in Table ?? B.4 Strategy Score Analysis See in Figure 4, 5 B.5 Tool use Analysis See in Figure 6 B.6 Category Analysis See in Table 7 B.7 All Model and Framework Results s See in Table 8 15 020406080 Days 0.0 0.2 0.4 0.6 0.8 1.0 Macro Strategy Similarity Easy Environment - Macro Strategy Score deepseekv3_2 gemini3_fast glm4_6 gpt5_1mini grok4_fast kimi_thinking qwen_235b Figure 4: Macro strategy similarity over time in the easy environment. Higher values indicate greater consistency in high-level decisions across days. 020406080 Days 0.0 0.2 0.4 0.6 0.8 1.0 Execute Strategy Similarity Easy Environment - Execution Strategy Score deepseekv3_2 gemini3_fast glm4_6 gpt5_1mini grok4_fast kimi_thinking qwen_235b Figure 5: Execution strategy similarity over time in the easy environment. Execution-level behaviors exhibit substantially higher temporal variability than macro strategies. C Prompt C.1 Evolving Strategy & Execution Framework Evolving Strategy Phase System Prompts You are a retail strategy analyst. Your task is to analyze current business data and determine whether the current strategy needs adjustment. ,→ ,→ ,→ Environment Characteristics - The store operates with a large number of SKUs, where products within the same category interact and may substitute or cannibalize each other's demand. ,→ ,→ ,→ ,→ 0.00.20.40.6 Avg Calls per Focus SKU-Day (SKU Reviews) 200 400 600 800 1000 Avg Daily Sales r = 0.528 p = 0.000 Runs: 65 (using: 59) deepseekv3_2 gemini3_fast glm4_6 gpt5_1mini grok4_fast k2_thinking kimi_thinking qwen_235b Trend SKU Reviews vs Avg Daily Sales (All Runs) Figure 6: Correlation between the frequency of SKU re- view queries during rollouts and average daily sales across all runs. Each point represents a single roll- out. We observe a positive association between review- related tool usage and sales performance. The Pearson correlation coefficientrand the correspondingp-value are reported in the figure. - Historical sales data provides essential signals for future decision-making. ,→ ,→ - Customer reviews dynamically influence product demand and sales velocity, with recent reviews having stronger effects. ,→ ,→ ,→ news_characteristic - Supply chains involve delivery lead times, requiring forward-looking inventory planning. ,→ ,→ - Orders require delivery time: When you place an order (place_order), the items will not arrive immediately. The delivery time varies and can take up to 7 days (within 7 days). You should plan your inventory accordingly and account for this lead time when making ordering decisions. Orders placed today will arrive within 7 days, but the exact arrival time is variable. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ - Inventory items depreciate in value over time and may require disposal when approaching expiration. ,→ ,→ - Supplier heterogeneity affects product quality perceptions and customer reviews, leading to supplier-dependent demand outcomes. ,→ ,→ ,→ 16 - Daily rent: The store incurs a fixed daily rent cost of daily_rent that must be paid each day. This daily operating cost is automatically deducted at the end of each day and makes cash-flow management critical for long-term survival and profitability. You must ensure sufficient funds are available to cover the daily rent. The daily rent amount is fixed and must be paid every single day, regardless of sales performance or other factors. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ # Your Role in Strategy Phase Each day starts with a STRATEGY PHASE where you:,→ 1. Review the current strategy (provided at the start of this phase),→ 2. Use data analysis tools to gather information about:,→ - Current inventory status - Recent sales history (last 30-60 days),→ - Customer reviews and ratings - Supplier prices and quality news_data_point - Current financial status 3. Compare current situation with previous days to identify significant changes ,→ ,→ 4. Set the strategy using the three separate tools (set_macro_strategy, set_execute_strategy, set_action) to set the three strategy components ,→ ,→ ,→ # Strategy Format The strategy consists of three components:,→ 1. **macro_strategy**: A list of broad strategic guidelines (array of strings) ,→ ,→ - Examples: ["Focus on high-margin products", "Maintain competitive pricing", "Prioritize inventory turnover"] ,→ ,→ ,→ 2. **execute_strategy**: An object with seven fields, all values are arrays:,→ - **focus_skus**: Array of SKU IDs that need attention (e.g., ["SKU_001", "SKU_002"]) ,→ ,→ - **sku_supplier_mapping**: Array of mapping objects (e.g., ["sku_id": "SKU_001", "supplier_id": "supplier_A", "sku_id": "SKU_002", "supplier_id": "supplier_B"]) ,→ ,→ ,→ ,→ ,→ news_to_monitor_field - **skus_to_reorder**: Array of SKU IDs that need reordering (e.g., ["SKU_003", "SKU_004"]) ,→ ,→ - **price_adjustments**: Array of price adjustment objects (e.g., ["sku_id": "SKU_001", "adjustment": "increase by 10%", "sku_id": "SKU_002", "adjustment": "decrease by 5%"]) ,→ ,→ ,→ ,→ ,→ ,→ - **sku_to_monitor**: Array of SKU IDs that should be closely monitored (e.g., ["SKU_005", "SKU_006"]) ,→ ,→ ,→ - **other**: Array of other strategy notes or metadata (e.g., ["comment": "...", "risk_level": "high"]) ,→ ,→ ,→ 3. **today_action**: An array of action objects, each representing a concrete action using the parameter format of`place_order` or `modify_sku_price`. ,→ ,→ ,→ ,→ - Each action MUST be an object of the form:,→ - "tool": "place_order", "arguments": <place_order arguments> ,→ ,→ - OR "tool": "modify_sku_price", "arguments": <modify_sku_price arguments> ,→ ,→ ,→ - Example: [ "tool": "place_order", "arguments": "sku_id": "SKU_001", "supplier_id": "supplier_A", "quantity": 100, ,→ ,→ ,→ ,→ 17 "tool": "modify_sku_price", "arguments": "sku_id": "SKU_002", "new_price": 9.99 ,→ ,→ ,→ ] # Strategy Setting Tools Use three separate tools to set different parts of the strategy:,→ - **set_macro_strategy**: Set the macro_strategy (array of strings),→ - Parameter:`macro_strategy` (array of strings),→ - Example: set_macro_strategy(macro_s ⌋ trategy=["Focus on high-margin products", "Maintain competitive pricing"]) ,→ ,→ ,→ - **set_execute_strategy**: Set the execute_strategy (object with seven fields, all arrays) ,→ ,→ - Parameter:`execute_strategy` (object with fields: focus_skus, sku_supplier_mappingnews_to_mon ⌋ itor_param, skus_to_reorder, price_adjustments, sku_to_monitor, other) ,→ ,→ ,→ ,→ ,→ - All field values must be arrays - Example: set_execute_strategy(execu ⌋ te_strategy="focus_skus": ["SKU_001"], "sku_supplier_mapping": ["sku_id": "SKU_001", "supplier_id": "supplier_A"], ...) ,→ ,→ ,→ ,→ ,→ ,→ - **set_action**: Set the today_action (array of action objects),→ - Parameter:`action` (array of objects, each with "tool" and "arguments" fields) ,→ ,→ - Each action object: "tool": "place_order" | "modify_sku_price", "arguments": ... ,→ ,→ ,→ - Example: set_action(action=["tool": "place_order", "arguments": "sku_id": "SKU_001", "supplier_id": "supplier_A", "quantity": 100]) ,→ ,→ ,→ ,→ ,→ You can call these tools multiple times to build or modify the strategy. After your analysis, set all three components to reflect your decisions. ,→ ,→ ,→ ,→ # Available Tools for Analysis The available function signatures are provided within <tools></tools> XML tags: ,→ ,→ <tools> tool_definitions </tools> For each function call, return a JSON object with function name and arguments inside <tool_call></tool_call> XML tags: ,→ ,→ ,→ <tool_call> "name": <function-name>, "arguments": <args-json-object>,→ </tool_call> # Important Analysis Tools Use these tools to gather data: - view_funds_and_date: Check current funds and date,→ - view_inventory: Check current inventory levels,→ - view_sku_sales_history: Analyze sales trends (use last 30-60 days),→ - view_sku_avg_ratings: Check customer satisfaction,→ - view_current_date_supplier_prices: Check supplier availability and prices ,→ ,→ news_tools_list - view_current_orders: Check pending orders,→ - set_macro_strategy: Set the macro strategy (array of strings),→ - set_execute_strategy: Set the execute strategy (object with seven fields, all arrays) ,→ ,→ - set_action: Set the today action (array of action objects),→ 18 Note: You CANNOT use place_order or modify_sku_price in the strategy phase. These tools are only available in the execution phase. ,→ ,→ ,→ After completing your analysis and updating the strategy, the system will transition to the EXECUTION PHASE. ,→ ,→ ,→ C.2 Evolving Strategy & Execution Execution Framework Phase System Prompts You are a retail operations agent executing daily operations based on the current strategy. ,→ ,→ # Your Role in Execution Phase In the EXECUTION PHASE, you will receive the **final strategy** determined in the Strategy Phase. This strategy includes: ,→ ,→ ,→ - **macro_strategy**: Broad strategic guidelines (array of strings),→ - **execute_strategy**: Specific operational details (object with seven fields, all arrays) ,→ ,→ - **today_action**: Concrete actions to take today (array of action objects),→ # Important Operational Constraints - **Daily rent**: The store incurs a fixed daily rent cost of daily_rent that must be paid each day. This daily operating cost is automatically deducted at the end of each day. Ensure you have sufficient funds to cover this daily expense. The daily rent amount is fixed and must be paid every single day, regardless of sales performance or other factors. This makes cash-flow management critical for long-term survival and profitability. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ - **Order delivery time**: When you place an order using place_order, the items will not arrive immediately. The delivery time varies and can take up to 7 days (within 7 days). Orders placed today will arrive within 7 days, but the exact arrival time is variable. Plan your inventory and ordering decisions accordingly, considering the lead time for items to arrive. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ # Strategy Usage Guidelines **The strategy is provided as REFERENCE, but you can and should make additional actions based on real-time data:** ,→ ,→ ,→ 1. **Reference the strategy** to understand priorities and planned actions: ,→ ,→ - Use macro_strategy for overall decision-making direction,→ - Use execute_strategy fields (focus_skus, sku_supplier_mappin ⌋ gnews_to_monitor_ref, skus_to_reorder, price_adjustments, sku_to_monitor, other) as guidance ,→ ,→ ,→ ,→ ,→ ,→ - Consider today_action as suggested actions to take,→ 2. **Perform additional data queries** to validate and refine decisions:,→ - Check current inventory levels, sales history, supplier pricesnews_impacts_ref, funds, etc. ,→ ,→ ,→ - Use tools like view_inventory, view_sku_sales_history, view_current_date_supplier_price ⌋ snews_tools_ref, etc. ,→ ,→ ,→ 3. **Execute actions flexibly**: - You can execute actions from today_action when they still make sense given the latest data ,→ ,→ 19 - You can **adjust, skip, or modify** actions from today_action if your analysis shows better alternatives ,→ ,→ ,→ - You can **add additional actions** beyond today_action if needed (e.g., unexpected inventory changes, new supplier pricesnews_impacts_example) ,→ ,→ ,→ ,→ - You can use information from execute_strategy (like focus_skus, sku_supplier_mapping) to make decisions even if not explicitly in today_action ,→ ,→ ,→ ,→ 4. **End the day** by calling end_today when you've completed all operations for today. ,→ ,→ # Important Constraints - You MUST NOT modify the stored strategy itself in this phase (strategy can only be changed in the Strategy Phase) ,→ ,→ ,→ - You CANNOT call any tool that changes macro_strategy / execute_strategy / today_action ,→ ,→ - You SHOULD use the strategy as guidance but make final decisions based on current data and analysis ,→ ,→ # Available Tools The available function signatures are provided within <tools></tools> XML tags: ,→ ,→ <tools> tool_definitions </tools> For each function call, return a JSON object with function name and arguments inside <tool_call></tool_call> XML tags: ,→ ,→ ,→ <tool_call> "name": <function-name>, "arguments": <args-json-object>,→ </tool_call> # Ending the Day When you have completed all reasonable operations for the day (especially those in today_action, adjusted as needed by current data), you MUST call end_today to advance to the next day. This will trigger a new Strategy Phase for the next day. ,→ ,→ ,→ ,→ ,→ ,→ C.3 Reflection Framework Reflection Phase Prompts This prompt is used to generate reflections after each day in run\_reflection.py. ,→ ,→ verbatim You are a retail operations analyst reflecting on the day's performance.,→ # Task Goal task_spec # Day day End Result end_today_result.get('formatted', safe_dump(end_today_result)),→ # Day day Interaction History interaction_summary memory_context # Your Task Generate a comprehensive reflection on today's performance. This reflection should be a complete, detailed analysis that will replace previous reflections. Include: ,→ ,→ ,→ ,→ 1. **Performance Summary**: Overall assessment of today's operations, including key metrics (funds, inventory, sales, etc.) ,→ ,→ ,→ 2. **Issue Identification**: What specific problems or challenges occurred? Be specific about what went wrong. ,→ ,→ ,→ 20 3. **Root Cause Analysis**: Why did these problems happen? Analyze the interaction history to understand what actions or decisions led to the issues. Trace back through the day's operations. ,→ ,→ ,→ ,→ ,→ 4. **What Worked Well**: Identify any successful strategies or decisions that should be continued. ,→ ,→ 5. **Actionable Improvements**: What should be done differently next time? Provide specific, actionable recommendations for future operations. ,→ ,→ ,→ ,→ 6. **Key Learnings**: What are the most important lessons learned from today that should guide future decision-making? ,→ ,→ ,→ Format your reflection as a comprehensive, detailed analysis (multiple paragraphs, not just a few sentences). This reflection will be the complete memory used for future days, so it should be thorough and cover all important aspects. ,→ ,→ ,→ ,→ ,→ ,→ Reflection: C.4 Macro Similiarity Judge Prompt Please compare the similarity of the following two macro strategies.,→ Strategy 1's macro_strategy: macro1_formatted Strategy 2's macro_strategy: macro2_formatted Please evaluate the similarity between these two strategies and provide a score between 0 and 1, where: ,→ ,→ - 1.0 means identical or almost identical,→ - 0.8-0.9 means very similar with only minor differences,→ - 0.6-0.7 means somewhat similar with some common points,→ - 0.4-0.5 means somewhat similar but with significant differences,→ - 0.2-0.3 means not very similar - 0.0-0.1 means completely different Please return only a floating-point number between 0 and 1, without any additional text or explanation. ,→ ,→ D Agent Framework D.1 Illustration of Agent Framework See in Figure 7 21 System: You are a retail strategy analyst. Your task is to analyze current business data and determine whether the current strategy needs adjustment. Agent: According to the history, I need to check the status of store and sku's price Action: <tool_call>"name": 'view_current_date_supplier_prices'</tool_call> Execution Phase Evolving Strategy Phase Action: <tool_call>"name": 'view_funds_and_date'</tool_call> Env: Current date: 1992-03-05, funds balance: 23261.66. Current Price of SKU 37003239 is ..... Agent: According to the response, I need to update the strategy ...... Action: <tool_call>"name": 'update_macro_strategy'</tool_call> System: You are a retail operations agent executing daily operations based on the current strategy. Agent: According to the strategy, I need to place the order and modify the price of SKU 37003239 Action: <tool_call>"name": 'modify_sku_price'</tool_call> Action: <tool_call>"name": 'place_order'</tool_call> Env: Updated SKU 37003239 price: 2.245232962110558 -> 2.492209044275104 Agent: According to the response, I think today's actions is enough Action: <tool_call>"name": 'end_today'</tool_call> Day 2: Evolving Strategy & Execution Day 1: Evolving Strategy & Execution System: You are a retail strategy analyst. Your task is to analyze current business data and determine whether the current strategy needs adjustment. Figure 7: Illustration of Evolving Strategy and Execution Framework 22 Strategy TypeStrategy Value macro_strategyTissues (SKU 3700060511) face an imminent stockout with current inventory below one day of coverage; incoming replenishments scheduled for 10-08 and 10-10 imply an unavoidable short-term gap and an estimated loss of approximately 1,200 units. Soups (SKU 5100000011) exhibit declining sales and persistently high return rates following a recent price increase, indicating substantial margin erosion. Cigarettes and ginger ale remain operationally stable with high customer ratings and low return rates. Overall operational risk is assessed as high due to anticipated revenue loss and return-driven inefficiencies. execute_strategyfocus_skus: 3700060511, 5100000011, 1254612128 sku_supplier_mapping: (3700060511, supplier_4), (5100000011, sup- plier_1), (1230000014, supplier_1), (1690000012, supplier_3), (1254612128, supplier_3) news_to_monitor: [ ] skus_to_reorder: [ ] price_adjustments: sku_id = 5100000011, adjustment = decrease to 0.45 or by 20% to stimulate demand sku_to_monitor: 3700060511, 5100000011 other: Tissues — stockout gap 10-05/06/07 estimated 3 days, approximately 1200 units lost at price 0.65 (∼780 revenue); incoming supply covers post- 10-08; no further reorder feasible Soups — return rate approximately 25% persists despite supplier_1; reviews indicate supplier_4 dominant issues but effects persist; sales declined after price hike; monitor next 3 days sales, returns, and supplier quality Mints — low stock but incoming sufficient Core — stable operations excluding tissues and soups risks; customer traffic approximately 30k/day steady risk_level: high today_actionplace_order: (3700060511, supplier_4, quantity = 800), (1254612128, sup- plier_3, quantity = 600) modify_sku_price: (sku_id = 5100000011, new_price = 0.45), (sku_id = 1690000012, new_price = 0.78) Table 5: Example Strategy 23 ModelAvg. Days↑Avg. Daily Sales↑Avg. Daily Income↑Expiry Ratio↓Return Ratio↓Max Days↑ Difficulty: EASY (5 Categories) DeepSeek-V3.2 (Exp.)58.33229.19183.260.08890.112266 Gemini-3 (Fast)50.67399.39294.710.07990.131159 GLM-4.652.40174.34124.670.07730.129358 Grok-4.1 Fast61.75508.08336.940.04170.084788 Kimi-K2 (Thinking)54.25260.68168.720.02390.117958 OpenAI-5.1 Mini51.75192.90122.460.03600.123755 Qwen-235B37.50420.31236.500.07450.084148 Average (7 models)52.38301.98203.300.05900.106861.71 Hand-crafted Policy180.00674.18729.460.02660.0070180 Difficulty: MIDDLE (20 Categories) DeepSeek-V3.2 (Exp.)54.67263.68255.300.10640.181863 Gemini-3 (Fast)42.67439.05449.040.01880.163048 GLM-4.654.33182.55131.900.02370.127456 Grok-4.1 Fast51.00720.69398.230.04070.116059 Kimi-K2 (Thinking)37.00347.78356.100.00370.167750 OpenAI-5.1 Mini56.33336.24223.600.13070.123557 Qwen-235B29.33216.06179.420.06390.133345 Average (7 models)46.48364.24282.480.04910.139954.00 Hand-crafted Policy180.001870.212809.390.02720.0074180 Difficulty: HARD (20 Categories) DeepSeek-V3.2 (Exp.)56.67165.86175.790.07870.197859 Gemini-3 (Fast)35.33513.10650.940.09060.154244 GLM-4.653.33205.76203.070.06450.164855 Grok-4.1 Fast33.67504.57364.520.04240.179750 Kimi-K2 (Thinking)43.67248.64312.820.01130.188957 OpenAI-5.1 Mini55.33331.36335.320.19490.192357 Qwen-235B40.33433.64268.060.17460.183448 Average (7 models)45.48320.96311.280.09600.179152.86 Hand-crafted Policy180.001667.842748.940.05070.0075180 Table 6: Performance comparison of seven large language models under three difficulty levels. Lower Expiry and Return ratios indicate better operational stability. A hand-crafted policy is included as an approximate upper bound for each difficulty. 24 ModelContext EasyMiddleHard SKU↑Cat↑SKU↑Cat↑SKU↑Cat↑ DeepSeek-V3.2 (Exp.)128K7.804.076.434.254.713.65 Gemini-3 (Fast)1M8.503.9511.8911.1813.328.09 GLM-4.6200K6.613.556.013.488.47 6.05 OpenAI-5.1 Mini400K6.802.646.826.817.875.19 Grok-4.1 Fast200K7.884.356.513.537.693.90 Kimi-K2 (Thinking)256K5.922.945.534.096.033.22 Qwen-235B (Thinking)256K4.783.666.494.055.774.23 Heuristic Strategy–25596209620 Table 7: SKU and Category counts observed during the strategy stage across difficulties. Context denotes the maximum supported context length of each model. Bolded indicates the best model;underlinedindicates the second best (Heuristic Strategy excluded). 25 ModelAvg. Days↑Avg. Daily Sales↑Avg. Daily Income↑Expiry Ratio↓Return Ratio↓Max Days↑ Framework: Evolving Strategy & Execution DeepSeek-V3.2 (Exp.)58.33229.19183.260.08890.112266 Gemini-3 (Fast)50.67399.39294.710.07990.131159 GLM-4.652.40174.34124.670.07730.129358 Grok-4.1 Fast61.75508.08336.940.04170.084788 Kimi-K2 (Thinking)54.25260.68168.720.02390.117958 OpenAI-5.1 Mini51.75192.90122.460.03600.123755 Qwen-235B (Thinking)37.50420.31236.500.07450.084148 GPT-5.281.00457.21358.270.06600.114181 Average (8 models)55.96330.26228.190.06100.112164.13 Framework: Reflection (Day-Level) DeepSeek-V3.2 (Exp.)53.00235.01170.040.03820.120066 Gemini-3 (Fast)45.67447.74255.380.06820.135050 GLM-4.655.00160.70125.670.01940.117662 Grok-4.1 Fast48.33297.94197.540.14600.092554 Kimi-K2 (Thinking)58.33216.51184.010.09640.125571 OpenAI-5.1 Mini53.3393.0492.110.10620.133159 Qwen-235B (Thinking)30.00371.90197.530.19190.155143 GPT-5.264.00283.88260.880.17740.088764 Average (8 models)50.96263.34185.400.10550.120958.63 Framework: Reflection (Step-Level) DeepSeek-V3.2 (Exp.)– Gemini-3 (Fast)– GLM-4.651.6792.3577.180.01480.107257 Grok-4.1 Fast– Kimi-K2 (Thinking)51.67181.19111.290.03530.136253 OpenAI-5.1 Mini– Qwen-235B (Thinking)– GPT-5.256.00398.71324.010.15360.104856 Average (8 models)53.11224.08170.830.06790.116155.33 Framework: Plan-and-Act DeepSeek-V3.2 (Exp.)– Gemini-3 (Fast)– GLM-4.648.33231.35113.010.01980.109655 Grok-4.1 Fast– Kimi-K2 (Thinking)48.67170.26105.360.00000.105651 OpenAI-5.1 Mini– Qwen-235B (Thinking)– GPT-5.264.00323.88193.020.01520.109664 Average (8 models)53.67241.83137.130.01170.108356.67 Heuristic Policy (Upper Bound, Easy) Hand-crafted Policy180.00674.18729.460.02660.007180 Table 8: Performance comparison of eight large language models under four agent frameworks in the EASY environment. A hand-crafted heuristic policy is included as an approximate upper bound. 26