Paper deep dive
Active Offline-to-Online Reinforcement Learning
Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 2:58:15 PM
Summary
This paper introduces an Active Offline-to-Online Reinforcement Learning (O2O-RL) framework to address the challenge of selecting and fine-tuning policies under a limited online interaction budget. The authors propose a method that balances policy evaluation and fine-tuning by actively selecting candidate policies based on upper-confidence bounds (UCB) derived from locally linear performance forecasts. This approach aims to mitigate the risks associated with distributional shift and hyperparameter sensitivity in offline RL, demonstrating superior performance over existing baselines in simulated robotics benchmarks.
Entities (8)
Relation Signals (6)
Offline RL → suffersfrom → Distributional Shift
confidence 95% · A central challenge in offline RL is that the performance of a pretrained policy becomes unpredictable as its behavior diverges from that of the policy used to collect the dataset
Active O2O-RL → uses → Upper-Confidence Bound (UCB)
confidence 95% · We propose an approach that balances this trade-off by actively selecting policies for fine-tuning based on upper-confidence bounds on their future performance.
Upper-Confidence Bound (UCB) → derivedfrom → Locally Linear Performance Forecasts
confidence 92% · These bounds are derived from locally linear performance forecasts fitted to observations obtained through online evaluation.
Active O2O-RL → addresses → Distributional Shift
confidence 90% · This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous.
Active O2O-RL → outperforms → Standard O2O-RL Baselines
confidence 88% · Across a diverse range of experiments, the proposed approach consistently outperforms existing O2O-RL baselines.
CalQL → istypeof → Offline RL
confidence 85% · Calibrated Q-learning (CalQL) (Nakamoto et al. 2023), building on CQL, learns conservative Q-values better suited for fine-tuning.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the policy with the highest estimated value. However, as in offline pretraining, fine-tuning performance is highly sensitive to the choice of algorithm and hyperparameters, making it risky to commit to a single policy. Objectives: We study active policy selection for fine-tuning under a limited interaction budget in O2O-RL settings. To our knowledge, this is the first work to address this problem. Methods: We formulate the problem by identifying a fundamental trade-off between allocating online interactions to policy evaluation, which helps identify high-performing policies, and allocating them to fine-tuning, which improves policy performance. We then propose an approach that balances this trade-off by actively selecting policies for fine-tuning based on upper-confidence bounds on their future performance. These bounds are derived from locally linear performance forecasts fitted to observations obtained through online evaluation. Results: Across a diverse range of experiments, the proposed approach consistently outperforms existing O2O-RL baselines. Conclusions: Actively selecting and fine-tuning policies uses limited online interaction budgets more effectively than either committing to a single policy or dividing the budget equally among all policies. Our framework also advances offline RL toward practical deployment in real-world systems where online interaction is costly or risky.
Tags
Links
- Source: https://arxiv.org/abs/2607.11720v1
- Canonical: https://arxiv.org/abs/2607.11720v1
Trouble viewing inline? Open PDF directly →
Full Text
59,928 characters extracted from source content.
Expand or collapse full text
Active Offline-to-Online Reinforcement Learning ALPER KAMIL BOZKURT * , Virginia Commonwealth University, USA SHANGTONG ZHANG, University of Virginia, USA YUICHI MOTAI, Virginia Commonwealth University, USA Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policies offline, evaluate them using off-policy or online evaluation, and then deploy and fine-tune the policy with the highest estimated value. However, as in offline pretraining, fine-tuning performance is highly sensitive to the choice of algorithm and hyperparameters, making it risky to commit to a single policy. Objectives: We study active policy selection for fine-tuning under a limited interaction budget in O2O-RL settings. To our knowledge, this is the first work to address this problem. Methods: We formulate the problem by identifying a fundamental trade-off between allocating online interactions to policy evaluation, which helps identify high-performing policies, and allocating them to fine-tuning, which improves policy performance. We then propose an approach that balances this trade-off by actively selecting policies for fine-tuning based on upper-confidence bounds on their future performance. These bounds are derived from locally linear performance forecasts fitted to observations obtained through online evaluation. Results: Across a diverse range of experiments, the proposed approach consistently outperforms existing O2O-RL baselines. Conclusions: Actively selecting and fine-tuning policies uses limited online interaction budgets more effectively than either committing to a single policy or dividing the budget equally among all policies. Our framework also advances offline RL toward practical deployment in real-world systems where online interaction is costly or risky. 1 Introduction Reinforcement learning (RL) is becoming a key ingredient in the autonomy of modern robotic systems operating in unstructured, dynamic environments (Singh et al. 2022). By learning to make decisions and derive control actions directly from onboard sensing and perception, RL can substantially reduce human workload and the likelihood of human error (Zhang et al. 2022), thereby facilitating widespread real-world deployment. Deep RL has proven effective in synthesizing control policies for high-dimensional, nonlinear physical systems for which manual controller design is infeasible, leading to numerous successful applications (Tang et al. 2025). Despite the flexibility and power of this framework, standard RL methods typically require extensive direct interaction with the physical environment for exploration (Ladosz et al. 2022). They are therefore impractical for training policies from scratch when such interactions are costly, risky, or time-consuming (Dulac-Arnold et al. 2021). Offline RL (Levine et al. 2020; Prudencio et al. 2023) has emerged as an alternative to online RL, enabling policies to be trained from large, previously collected datasets. These datasets are usually collected under safe, controlled conditions (G. Zhou et al. 2023), often, though not exclusively, by human operators. A central challenge in offline RL is that the performance of a pretrained policy becomes unpredictable as its behavior diverges from that of the policy used to collect the dataset (Kostrikov et al. 2021). Due to this distributional shift, a policy pretrained via offline RL can perform arbitrarily poorly in the real environment (Qin et al. 2022). To partially mitigate this issue, offline policy selection, usually performed via off-policy evaluation (OPE) (Paine et al. 2020; Uehara et al. 2025), is used to identify high-performing policies among candidates pretrained using different hyperparameter * Corresponding Author. Authors’ Contact Information: Alper Kamil Bozkurt, bozkurta@vcu.edu, Virginia Commonwealth University, Richmond, Virginia, USA; Shangtong Zhang, xdm2bt@virginia.edu, University of Virginia, Charlottesville, Virginia, USA; Yuichi Motai, ymotai@vcu.edu, Virginia Commonwealth University, Richmond, Virginia, USA. , Vol. 0, Article 0. Publication date: 0. arXiv:2607.11720v1 [cs.LG] 13 Jul 2026 0:2 • Bozkurt, Zhang & Motai configurations. However, OPE estimates are generally not sufficiently reliable to determine which policy should be deployed, primarily because they are also vulnerable to distributional shift (Brandfonbrener et al. 2021). This difficulty is further exacerbated by nonstationary real-world environments that require online adaptation (Julian et al. 2021). The need to evaluate and fine-tune policies pretrained via offline RL through online interaction has motivated the offline-to-online RL (O2O-RL) paradigm, which treats the entire pipeline as a single joint problem. By combining offline pretraining with a small amount of online interaction, O2O-RL can efficiently produce high-performing policies, thereby enabling the broader deployment of RL in physical domains. As a result, O2O-RL has attracted substantial attention and led to the development of numerous methods. Most O2O-RL approaches (e.g., Ball et al. (2023), Lee et al. (2022), and Nair et al. (2020)) focus on improving the offline stage to produce pretrained policies that adapt more effectively to real environments during fine-tuning. Although these methods improve overall performance, they do not address the fundamental sensitivity of policy performance to hyperparameter choices and environmental conditions. More recently, several methods have examined policy selection in O2O-RL (e.g., Konyushova et al. (2021) and Kurenkov and Kolesnikov (2022)), aiming to identify the best candidate by evaluating pretrained policies under a limited online interaction budget. However, these methods do not incorporate fine-tuning into the policy-selection process and overlook the trade-off arising from allocating the same limited interaction budget to both policy evaluation and fine-tuning. In this work, we address the problem of obtaining high-performing, deployable policies in O2O-RL from a broader and more practical perspective. Prior work and our experiments show that pretrained policies can perform arbitrarily poorly in real environments and that fine-tuning may require substantial interaction to improve performance. Fine-tuning can even cause performance regressions, with outcomes varying across algorithms, hyperparameters, and environments. To manage this performance volatility under a limited interaction budget, we propose an O2O-RL approach (Fig. 1) that jointly performs active policy selection and fine-tuning. During the offline stage, we first train a diverse set of candidate policies using offline RL across a representative range of algorithms and hyperparameter configurations. During the novel online stage, we actively select and fine-tune candidate policies using an upper-confidence-bound (UCB) criterion based on their predicted future performance until the interaction budget is exhausted. After each online episode, we fine-tune the selected policy, evaluate its performance, and then fit a local linear model to forecast its future performance and construct a corresponding UCB. We switch policies whenever the bound of the selected policy falls below that of another candidate. In this manner, we efficiently allocate the interaction budget to both identifying promising policies and improving them through fine-tuning. We summarize our main contributions as follows: • To the best of our knowledge, we are the first to study policy selection for fine-tuning under a limited interaction budget in O2O-RL settings. •We propose a new approach that actively selects and fine-tunes policies using a UCB criterion based on future performance predicted through local regression. • We evaluate our approach on a suite of simulated robotics benchmarks and demonstrate consistent gains over standard O2O-RL baselines. This article substantially extends our earlier conference paper (Bozkurt et al. 2026) by providing additional methodological and technical discussion, a more comprehensive analysis of the proposed framework, and con- siderably expanded experiments. In particular, this version includes results on a broader collection of navigation, classic-control, and locomotion tasks, as well as ablation studies evaluating the method’s sensitivity to its design choices and its generality across environments and experimental settings. , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:3 휋 1 휋 2 ... ... 휋 푘−1 휋 푘 � 흅 ∗ 흅 ∗ (a) Dataset Collection (c) Candidates (b) Offline RL (d) Selection & Tuning Fig. 1. Proposed O2O-RL Framework. (a) Datasets are typically collected in controlled but imperfect settings. (b) Offline RL is used to train a diverse set of candidate policies using different algorithms and hyperparameter configurations. (c) A local linear model predicts the future performance of each policy and constructs an upper confidence bound (UCB). (d) The policy with the highest UCB is selected and fine-tuned, after which its predicted future performance and UCB are updated. Whenever its UCB falls below that of another policy, it is replaced by the policy with the higher UCB. 2 Related Work 2.1 Offline Reinforcement Learning 2.1.1 Basic Approaches. A naive approach to offline RL is to train policies via standard online, off-policy algorithms such as deep deterministic policy gradient (DDPG) (Lillicrap et al. 2016), twin delayed DDPG (TD3) (Fujimoto, Hoof, et al. 2018), and soft actor-critic (SAC) (Haarnoja et al. 2018) on batches drawn from previously collected datasets. However, such policies can fail catastrophically, as they may lead to states outside the dataset due to substantially inaccurate value estimates arising from distributional shift. In contrast, behavior cloning (BC); e.g., (Florence et al. 2022; Torabi et al. 2018), which aims to imitate the dataset behavior and thereby largely prevent distributional shift, often cannot surpass the performance of the behavior policy, which is undesirable for non-expert datasets. 2.1.2 Offline RL for Distributional Shift. The central objective in offline RL is to improve upon the behavior policy of the dataset while avoiding severe failures caused by distributional shift. Most offline RL methods therefore balance this trade-off by permitting off-policy updates with additional regularization or constraints. Some approaches, such as batch-constrained deep Q-learning (BCQ) (Fujimoto, Meger, et al. 2019), bootstrapping error accumulation reduction (BEAR) (Kumar, Fu, et al. 2019), and policy in the latent action space (PLAS) (W. Zhou et al. 2021), focus on constraining actions to prevent policies from taking actions outside the dataset. Other approaches, such as conservative Q-learning (CQL) (Kumar, A. Zhou, et al. 2020) and TD3+BC (Fujimoto and Gu 2021), incorporate, respectively, Q-value and BC regularization terms into the policy updates. We refer readers to (Levine et al. 2020) and (Prudencio et al. 2023) for comprehensive reviews of offline RL approaches. 2.2 Offline-to-Online Reinforcement Learning 2.2.1 Offline RL for Fine-Tuning. The fundamental limitation that the performance of offline RL methods is upper-bounded by dataset quality necessitates fine-tuning pretrained policies via online interactions. This insight has motivated methods that treat the entire offline-to-online framework as a unified problem, where the main objective is to improve the final performance of policies after fine-tuning. Advantage weighted actor critic (AWAC) (Nair et al. 2020) improves the efficiency of fine-tuning with off-policy data via dynamic programming. Implicit Q-learning (IQL) (Kostrikov et al. 2021) utilizes a policy-extraction step during training that aids fine-tuning. Calibrated Q-learning (CalQL) (Nakamoto et al. 2023), building on CQL, learns conservative Q-values better suited , Vol. 0, Article 0. Publication date: 0. 0:4 • Bozkurt, Zhang & Motai for fine-tuning. Revisited behavior regularized actor-critic (ReBRAC) (Tarasov et al. 2023) extends TD3+BC with simple hyperparameter choices that benefit both offline training and fine-tuning. Hybrid RL (Hy-Q) (Song et al. 2023) augments offline and online data and applies a fitted Q-iteration procedure. Robust O2O (Wen et al. 2024) uses an ensemble of Q-networks and smoothness regularization to mitigate distributional shift. In this work, we do not attempt to introduce or improve any particular O2O-RL algorithm. Instead, we treat the choice of algorithm as a high-level hyperparameter when training candidate policies and fine-tuning, since no single O2O-RL method is uniformly best in all environments. 2.2.2 Online Policy Selection. Recent studies, (Konyushova et al. 2021) and (Kurenkov and Kolesnikov 2022), studied the policy selection problem in O2O-RL. In a setting similar to ours, these studies use a limited online- interaction budget to identify the best-performing pretrained policies, rather than relying on OPE estimates as in prior offline RL approaches (Paine et al. 2020; Uehara et al. 2025). Since evaluating all candidates online would quickly exhaust the budget, they propose active selection strategies to decide which policies to evaluate, thereby allocating the budget efficiently. In contrast, our focus is to allocate online interactions to identify the policies that will perform best after fine-tuning under the given budget. 3 Problem Formulation 3.1 Background and Setup Following the standard RL framework, we model interactions with an environment to perform tasks as Markov decision processes (MDPs). Formally, an MDPMis a tuple(푆,퐴,푃,푝 0 ,푅,훾), where푆is the state set,퐴is the action set,푃and푝 0 are the transition and initial-state distributions,푅is the reward function, and훾is the discount factor. The ultimate objective in RL is to obtain a policy휋 : 푆 ↦→ 퐴that maximizes the expected cumulative discounted reward; i.e., 휋 ∗ : = argmax 휋 E " ∞ ∑︁ 푡=0 훾 푡 푟 푡 # .(1) In offline RL, this objective must be achieved using a datasetDof transitions(푠 푖 ,푎 푖 ,푟 푖 ,푠 ′ 푖 ) 퐷 푖=1 collected by a typically unknown behavior policy 휋 푏 inM, without further interaction with the environment. In this work, we extend the offline RL setting by allowing푀online interactions in addition to the datasetD; i.e., a set of at most푀new transitions may be collected from the environment. Unlike in (Konyushova et al. 2021) and (Kurenkov and Kolesnikov 2022), in our O2O-RL setting, these online interactions can be used not only for evaluation but also for fine-tuning, introducing a novel trade-off. For simplicity, we represent the budget as푁 fine-tuning & evaluation iterations, each using푀/푁transitions. Our aim is then, given an MDPM, a datasetD collected fromM, and an interaction budget푁, to obtain the highest-performing policy achievable, in the sense of minimizing regret. 3.2 Problem Statement We formalize policy selection as the following simple regret-minimization problem: PROBLEM 1. Consider a set of퐾candidate policies pretrained offline on a dataset with different hyperparameter settings, and let휋 푖 푗 denote the policy that is obtained from the푖th pretrained policy after푗fine-tuning iterations, and let 푣 푖 푗 : =E 휋 푖 푗 " ∞ ∑︁ 푡=0 훾 푡 푟 푡 # (2) , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:5 denote its corresponding value. For a given interaction budget푁, the optimal selection is then a policy휋 푖 ∗ 푗 ∗ such that 푖 ∗ , 푗 ∗ : =argmax 1≤푖≤퐾,0≤푗≤푁 푣 푖 푗 .(3) Our objective is to devise a procedure that fine-tunes each pretrained policy휋 푖 0 for ̄ 푗 푖 iterations and estimates the values ˆ 푣 푖 0 , . . ., ˆ 푣 푖 ̄ 푗 푖 under the constraint Í 퐾 푖=1 ̄ 푗 푖 ≤ 푁 such that the value of the selected policy ˆ 푣 ∗ = max 1≤푖≤퐾 max 0≤푗≤ ̄ 푗 푖 ˆ 푣 푖 푗 (4) minimizes regret : = 푣 푖 ∗ 푗 ∗ − ˆ 푣 ∗ .(5) In our setting, each policy휋 푖 푗 is specified by trained parameters휃 푖 푗 , an RL algorithmA 푖 , and a set of hyperpa- rameters휆 푖 that determine the offline training, fine-tuning, and action-selection procedures. Each configuration 휓 푖 = (A 푖 ,휆 푖 ), indexed by푖 ∈ 1, . . .,퐾, is predetermined and fixed, and is used to produce a policy lineage 휋 푖 0 ,휋 푖 1 , . . .with corresponding parameters휃 푖 0 ,휃 푖 1 , . . .. Here,휃 푖 0 is obtained through offline training, and for푗 ≥ 1, 휃 푖 푗 is obtained from휃 푖 푗−1 after one fine-tuning iteration. Henceforth, we omit the explicit notation forA 푖 ,휆 푖 , and 휃 푖 푗 , and instead treat them as implicit in the indices of휋 푖 푗 . The optimal selection휋 푖 ∗ 푗 ∗ intuitively represents the best policy that could have been obtained if the policy sequences and their corresponding values were known a priori. It therefore constitutes an upper bound on the improvement achievable by any policy-selection procedure during fine-tuning. We can establish a lower bound on the regret of this formulation in a general setting as a function of the number of candidates퐾and the budget푁. Suppose that the estimates of the policy values follow a normal distribution with mean휇and standard deviation휎for every configuration index푖and fine-tuning iteration푗, independently of all previous values. The expected maximum valueE[푣 푖 ∗ 푗 ∗ ] that can be obtained by fine-tuning all of these policies for푁 iterations can be lower-bounded as follows: E[푣 푖 ∗ 푗 ∗ ]=E max 1≤푖≤퐾,0≤푗≤푁 ˆ 푣 푖 푗 (6) ≥ 휎 √︄ log(퐾(푁 + 1)) 휋 log(2) .(7) In this setting, all selection procedures are equivalent to choosing the first configuration for fine-tuning at every iteration, and they yield the same expected value, which is upper-bounded by E[ ˆ 푣 ∗ ]=E max 0≤푗≤푁 푣 1 푗 (8) ≤ 휎 √︁ 2 log(푁 + 1).(9) Thus, we obtain the following lower bound on the regret: E[regret]=E[푣 푖 ∗ 푗 ∗ − ˆ 푣 ∗ ](10) =E[푣 푖 ∗ 푗 ∗ ]−E[ ˆ 푣 ∗ ](11) ≥ 휎 √︄ log(퐾(푁 + 1)) 휋 log(2) − √︁ 2 log(푁 + 1) ! (12) , Vol. 0, Article 0. Publication date: 0. 0:6 • Bozkurt, Zhang & Motai 0 1 2 0 1 2 3 maximum value at 푗 maximum value at 푗+1 1− 푝 푝 0.8 0.1 0.1 1 1 2 0 1 2 3 maximum value at 푗 maximum value at 푗+1 1 0.8 0.1 0.1 2 1 2 0 1 2 3 maximum value at 푗 maximum value at 푗+1 1 0.9 0.1 3 1 2 0 1 2 3 maximum value at 푗 maximum value at 푗+1 1 1 Fig. 2. Illustration of a simple policy-selection problem for fine-tuning, formulated as a Markov decision process. The circles denote states representing the maximum policy value achieved at the current fine-tuning iteration; the rectangles denote choices among training configurations; and the numbers above the arrows indicate the corresponding transition probabilities. This lower bound becomes positive for퐾=(푁 + 1) 푞 with푞> 2휋 log(2)− 1≈ 1.95and increases as퐾increases. This issue stems from the fact that the number of pretrained policies퐾is given rather than chosen in the problem formulation, which causes the regret to increase inherently with퐾. However, we implicitly assume that퐾 ≤ 푁, which alleviates this limitation. 3.3 Relation to Bandit and Low-Switching-Cost Formulations Our formulation is closely related to pure-exploration multi-armed bandits (Kaufmann et al. 2016), where the objective is to identify the best arm under a fixed budget without considering cumulative regret. In our formulation, arms correspond to configurations, playing an arm corresponds to selecting the most recent fine-tuned policy from the corresponding configuration, and the payoff is the value of the resulting policy after fine-tuning. The key distinction is that, in our setting, the arms are nonstationary, since the policy lineages induced by configurations evolve over time through fine-tuning. Moreover, the objective is to obtain the highest-payoff play, namely, to identify the fine-tuned policy with the highest value. Our formulation is also related to the notions of low switching cost (Bai et al. 2019) and deployment complexity (Huang et al. 2022), where the objective is to minimize the number of policy updates required to obtain a near- optimal policy. However, these formulations differ from ours in a key respect: we do not directly address the design of a new RL algorithm, nor do we aim to find a near-optimal policy with respect to cumulative reward. Instead, we focus on selecting, from a given set of candidate policies, the policy that achieves the highest performance after fine-tuning under a limited interaction budget. 3.4 Illustrative Example We illustrate two key factors induced by this problem formulation that affect the policy selection process: (i) the size of the remaining budget and (i) the highest value among previously obtained fine-tuned policies. Consider the example shown in Fig. 2, where there are two candidate policies with configurations indexed by푖 ∈ 1, 2. Each fine-tuning iteration푗, independently of previous fine-tuning iterations, produces a policy whose value푣is sampled from a static distribution associated with the corresponding training configuration. For simplicity, we assume that these distributions and their parameters are known a priori, and that policy values can be obtained directly without online evaluation. Fine-tuning the policy with the second configuration (푖= 2) exhibits higher variance than fine-tuning the policy with the first configuration (푖= 2). Specifically, for푖= 1, the policy value is1with probability (w.p.)푝and0w.p. 1− 푝; for푖= 2, the value is3w.p.0.1,2w.p.0.1, and0w.p.0.8. The expected values are therefore푝and0.5, respectively. Now consider the case푁= 3. The optimal strategy at any iteration푗can be specified as a function , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:7 050100150200 Fine-tuning Iteration 20 0 20 40 60 80 Evaluation Value Swimmer Expert Seed: 1 050100150200 Fine-tuning Iteration Swimmer Expert Seed: 2 050100150200 Fine-tuning Iteration Swimmer Expert Seed: 3 CalQL LR Scale: 1.0 CalQL LR Scale: 0.5 ReBRAC LR Scale: 1.0 ReBRAC LR Scale: 0.5 Fig. 3. Evolution of the return values of pretrained policies during fine-tuning on SWIMMER-EXPERT for three random seeds, each of which determines the initial state. Policies are pretrained with offline RL algorithms CalQL and ReBRAC for 200 K steps using default and half-default learning rates (LR). The value curves are highly irregular: they may improve, stall, regress after initial improvement, or exhibit high variance. Even with the same algorithms and hyperparameters, changing the initial state can significantly alter the outcome of offline training and the progression of the value curves, which necessitates an active policy selection approach during fine-tuning. of푣 ∗ 푗 , the value of the best-performing fine-tuned policy obtained up to iteration푗. If푣 ∗ 푗 = 3 , selecting either configuration is optimal because the highest possible value has already been obtained. If푣 ∗ 푗 ∈ 1, 2, the optimal strategy is to choose 푖= 2, since it is the only configuration that can produce a policy with a higher value. When푣 ∗ 푗 = 0, the optimal strategy depends strongly on both푗and푝. For푝< 0.5, the optimal strategy is to choose푖= 2at every iteration, yielding훼 ∗ =2, 2, 2for푗 ∈ 0, 1, 2, because its expected value is higher than that of푖= 1regardless of푣 푗 . When푝= 0.6, the optimal strategy becomes훼 ∗ = 1, 1, 1as the expected value of푖= 1increases. However, further increasing푝to0.9and0.99yields the optimal strategies훼 ∗ = 2, 1, 1and 훼 ∗ = 2, 2, 1, respectively. This occurs because larger values of푝make it increasingly likely that a fine-tuned policy with value 1 can be obtained in later iterations, which encourages greater risk taking in earlier iterations in pursuit of policies with values2and3. This example illustrates the distinctive nature of our problem formulation. 4 Active Offline-to-Online Reinforcement Learning Our overall procedure is shown in Algorithm 1. We now explain our approach in detail below. 4.1 Value Forecast We follow the offline initialization step in the standard O2O pipeline. We begin by training a diverse pool of퐾 candidate policies,휋 1 0 , . . .,휋 퐾 0 , using offline RL on the given datasetD. We sweep across multiple algorithmic families and hyperparameter ranges to cover the design space suitable for the environmentM. Each candidate policy휋 푖 0 is associated with a configuration휓 푖 that contains all information needed for pretraining, fine-tuning, and prediction. This includes the offline and online RL algorithm choices, as well as all hyperparameters, such as the neural network architecture, activation functions, optimizers, loss functions, number of pretraining steps, exploration rate, random seed, and other relevant details. A configuration휓 푖 , together with pretrained parameters 휃 푖 푗 obtained according to휓 푖 , completely defines the candidate policy 휋 푖 0 . At initialization, we use휏pseudo-estimates( ˆ 푣 푖 −휏 , . . ., ˆ 푣 푖 0 ) for each configuration, where휏is chosen according to the complexity of the local regression model used for forecasting. For example, since we use linear regression in , Vol. 0, Article 0. Publication date: 0. 0:8 • Bozkurt, Zhang & Motai Algorithm 1: Active Offline-to-Online Reinforcement Learning Input : MDPM, datasetD, interaction budget 푁 , window size 푤 Output : policy 휋 ∗ # Offline Stage 1 Estimate behavior policy value ˆ 푣 퐵 usingD 2 Train 퐾 candidate policies휋 푖 0 퐾 푖=1 onD # using diverse algorithms and hyperparameters 3 Initialize a policy lineage listΠ 푖 ←(휋 푖 0 ) for each 휋 푖 0 # Online Initialization 4 휋 ∗ ← null # best policy 5 푣 ∗ ←−∞ # best value estimate 6 for 푖= 1 to 퐾 do 7Initialize a value estimate list ˆ 푉 푖 = ˆ 푣 푅 , ˆ 푣 푅 8 휋 ∗ , ˆ 푣 ∗ ← finetune_eval(M,Π 푖 , ˆ 푉 푖 ,휋 ∗ , ˆ 푣 ∗ ) 9 ̃ 푣 푖 , ̃ 푠 푖 ← forecast( ˆ 푉 푖 ,푤,푑,푐) 10 end # Policy Selection & Fine-Tuning 11 for 푗=휏 · 퐾 to 푁 do 12 푖 ∗ ← min 푖 (푣 ∗ − ̃ 푣 푖 )/ ̃ 푠 푖 # pick the one requiring min scale to reach 푣 ∗ 13 휋 ∗ , ˆ 푣 ∗ ← finetune_eval(M,Π 푖 ∗ , ˆ 푉 푖 ∗ ,휋 ∗ , ˆ 푣 ∗ ) 14 ̃ 푣 푖 ∗ , ̃ 푠 푖 ∗ ,푢 푖 ∗ ← forecast( ˆ 푉 푖 ∗ ,푤, min(푑,푁 − 푗),푐) 15 end 16 return 휋 ∗ Algorithm 2: Fine-Tuning & Online Evaluation Input : MDPM, policy lineage listΠ, value estimate list ˆ 푉 , best policy 휋 ∗ , best value estimate 푣 ∗ Output : new best policy 휋 ∗ , new best value estimate 푣 ∗ 1 function finetune_eval(M,Π, ˆ 푉,휋 ∗ , ˆ 푣 ∗ ) 2 휋 푡+1 ← finetune(M,휋 푡 ) 3Π.append(휋 푡+1 ) 4 ˆ 푣 푡+1 ← evaluate(M,휋 푡+1 ) 5 ˆ 푉.append( ˆ 푣 푡+1 ) 6if ˆ 푣 푡+1 > 푣 ∗ then 7푣 ∗ ← ˆ 푣 푡+1 8휋 ∗ ← 휋 푡+1 9end 10 end 11 return 휋 ∗ , 푣 ∗ our experiments, we set휏= 1. These pseudo-estimates mitigate the initial identification and high-variance problems that arise when only a few early value estimates are available for fitting the model. We set all pseudo-estimates , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:9 Algorithm 3: Value Forecast Input : value estimate list ˆ 푉 , window size 푤 , future iterations 훿 , exploration degree 푐 Output : forecasted value ̃ 푣 , prediction interval ̃ 푠 1 function forecast( ˆ 푉,푤,훿,푐) 2 푡 ← ˆ 푉.length() # number of iterations used to fine-tune 3 휔 ← min푤,푡 # clipped window size 4Fit a local linear model ˆ 푏 푡 , ˆ 푚 푡 to ˆ 푉[−휔 :] # last 휔 value estimates 5 푡 ′ ← 푡 +훿 # predictor 6Forecast the future value ̃ 푣 푖 ← ˆ 푏 푡 + ˆ 푚 푡 푡 ′ with the prediction interval ̃ 푠 7 end 8 return ̃ 푣 , ̃ 푠 equal to the value of a random policy, ˆ 푣 푅 , which we assume to be known a priori. We do not use OPE values for initialization because their accuracy can be quite limited. The values of policies during fine-tuning do not always increase monotonically. They may stall or regress, exhibit sudden drops or jumps, and pass through both high- and low-variance regimes as we show in Fig. 3. We adopt a local linear regression model to capture short-term upward or downward trends, together with a local residual variance estimate to capture time-varying uncertainty. Although we favor this minimalist design to highlight the key aspects of our approach, richer forecasting models, such as higher-order or weighted local polynomial regression, can also be employed depending on the application. In particular, for each configuration푖, we assume that the evolution of the value estimates during fine-tuning can be locally approximated by a linear model around time 푡 : 푣 푖 푡 (푡 ′ ) ≈ 푏 푖 푡 +푚 푖 푡 · 푡 ′ +휖 푖 푡 (푡 ′ )(13) where휖 푖 푡 (푡 ′ ) ∼ 푁(0,휎 푖 푡 )is an independent and identically distributed normal error term. After obtaining each value estimate ˆ 푣 푖 푡 during fine-tuning, we estimate the model parameters as follows: ˆ 푏 푖 푡 , ˆ 푚 푖 푡 : = argmin 푏,푚 푡 ∑︁ 푡 ′ =푡−푤 ( ˆ 푣 푖 푡 ′ −푏−푚푡 ′ )(14) where푤is the bandwidth parameter controlling the size of the window over which linearity is assumed to hold. We then use these estimated parameters to predict the future value at time 푡 ′ : ̃ 푣 푖 푡 (푡 ′ )= ˆ 푏 푖 푡 + ˆ 푚 푖 푡 푡 ′ .(15) We also calculate the prediction interval ̃ 푠 푖 푡 (푡 ′ ) as ̃ 푠 푖 푡 (푡 ′ ) : =푠 푖 푡 √︄ 1+ 1 푤 + 1 + (푡 ′ − ̄ 푤) 2 푤푠 2 푤 (16) where푠 푖 푡 : = Í 푡 푡 ′ =푡−푤 ( ˆ 푣 푖 푡 ′ − ˆ 푏 푖 푡 ,− ˆ 푚 푖 푡 ,푡 ′ )/(푤− 1) is the standard deviation of the residuals, ̄ 푤 : = Í 푡 푡 ′ =푡−푤 푡 ′ /(푤+ 1)= 푡 − 푤/2is the window mean, and푠 2 푤 : = Í 푡 푡 ′ =푡−푤 (푡 ′ − ̄ 푤) 2 /푤= (푤 + 1)(푤 + 2)/12is the window variance. To understand how the window size and prediction horizon affect the interval, consider larger windows,푤 ≥ 10, and relatively long-horizon predictions with푑(푡 ′ ) : = (푡 ′ − 푡) ≥ 2푤. In this regime, we can further approximate the , Vol. 0, Article 0. Publication date: 0. 0:10 • Bozkurt, Zhang & Motai interval as ̃ 푠 푖 푡 (푡 ′ ) ≈ 푠 푖 푡 √︂ 1+ 12푑 2 (푡 ′ ) 푤 .(17) This expression indicates that the size of the prediction interval decreases as the window size increases, since more data provides greater confidence in the estimates, and increases as the relative prediction horizon grows. 4.2 Policy Selection Policy selection is performed through an iterative fine-tuning loop. At each iteration, we select a policy for fine- tuning according to the UCBs of future policy values forecasted by the fitted local linear models, fine-tune the most recent policy in the corresponding lineage, and evaluate the resulting policy through online interaction. This process is repeated until the interaction budget is exhausted, after which we return the policy with the highest obtained value estimate. The intuition behind this procedure is to efficiently allocate fine-tuning resources while maintaining a reasonable regret bound. At iteration푗, for each configuration index푖, we store the value estimates obtained through online evaluation in a listv 푖 푗 and fit a local linear model to predict future values using a fixed window size푤. During the initial stage of fine-tuning, however, the window size is capped by the number of previous value estimates. This produces larger prediction intervals, thereby encouraging exploration especially during the initial phase. We predict the value푤 iterations into the future whenever the remaining interaction budget permits; otherwise, we predict the value at the final iteration allowed by the remaining budget. Let푡 푖 푗 denote the number of policies in the lineage associated with the푖th configuration at fine-tuning iteration 푗. We project the future value ̃ 푣 푖 푡 ′ and its associated prediction interval ̃ 푠 푖 푡 ′ using the locally fitted linear model described above, where 푡 ′ : = 푡 + min푤,푁 − 푗, and compute the UCB as 푢 푖 푗 : = ̃ 푣 푖 푡 ′ +푐 ̃ 푠 푖 푡 ′ (18) where푐is a parameter that scales the prediction interval and thereby controls the exploration-exploitation trade-off. We then select the policy with the largest UCB, 푖 ∗ 푗 : = argmax 푖 푢 푖 푗 (19) fine-tune the most recent policy in the selected lineage, and append the estimated value of the resulting policy to the corresponding value list. Inspired by the illustrative example in Fig. 2, we propose a procedure for determining the exploration degree based on the highest policy value estimate obtained up to iteration 푗 , defined as 푣 ∗ 푗 : = max 푖 max 0≤푡≤푡 푖 푗 ˆ 푣 푖 푡 ,(20) Intuitively, our procedure increases the scale of the prediction intervals until one of the UCBs exceeds푣 ∗ 푗 and picks its corresponding policy. In this way, we avoid introducing푐as an additional tunable parameter. Specifically, we select 푖 ∗ 푗 : = argmin 푖 (푣 ∗ 푗 − ̃ 푣 푖 푡 ′ )/ ̃ 푠 푖 푡 ′ .(21) Overall, our active policy-selection strategy is based on two key principles: (i) short-horizon forecasting based on local trends to allocate the fine-tuning budget efficiently and (i) UCB-based, automatically calibrated exploration to avoid prematurely committing to suboptimal policy lineages. , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:11 5 Experiments 5.1 Implementation Details We employ four representative training algorithms designed for offline pretraining followed by online fine-tuning: AWAC (Nair et al. 2020), IQL (Kostrikov et al. 2021), CalQL (Nakamoto et al. 2023), and ReBRAC (Tarasov et al. 2023). We use the implementations of these algorithms provided by the offline RL library D3RLPY (Seno and Imai 2022). For each algorithm, we consider four hyperparameter settings obtained from the Cartesian product of two batch sizes (the default and half the default) and two learning rates (the default and half the default). We pretrain one candidate policy for each setting for 200 K iterations on the corresponding dataset, while fixing all remaining hyperparameters to the D3RLPY defaults. This procedure yields 16 pretrained candidate policies for each environment. We conduct the simulated experiments on a Linux machine using 16 AMD EPYC processor cores and 32 GB of RAM for each environment. We repeat each experiment using four random seeds. We fine-tune the pretrained policies through online interactions with the environments, using the same hyper- parameters as during pretraining. At each fine-tuning iteration, we use 4 K online transitions for training and an additional 1 K transitions for evaluation, resulting in a total cost of 5 K transitions per iteration. We consider window sizes of푤 ∈ 3, 4, 5, 6, 7for the value estimates used to fit the local linear model and generate forecasts. We use 푤= 3 for the main experiments and evaluate the remaining window sizes in the ablation study. 5.2 Baselines We compare our adaptive policy-selection and fine-tuning approach with several standard O2O-RL baselines. RANDOM: A pretrained policy is selected uniformly at random and used without fine-tuning. BEST: The pretrained policy with the highest estimated value is selected without fine-tuning. FTS (Fine-Tuning Single): A randomly selected pretrained policy is fine-tuned using the entire interaction budget, after which the policy with the highest value in its lineage is selected. FTA (Fine-Tuning All): All pretrained policies are fine-tuned by dividing the interaction budget equally among them, after which the policy with the highest estimated value is selected. ACTIVE (Ours): Our active O2O-RL approach; see Algorithm 1. 5.3 Environments We evaluate our approach on a diverse set of continuous-control environments from the Minari Python library (Farama 2021), spanning navigation, classic-control, and locomotion tasks. Minari provides datasets generated using MuJoCo (Todorov et al. 2012) in accordance with the principles of the D4RL benchmark (Fu et al. 2020). Further details on dataset generation are provided in (Farama 2021). For each environment, we construct the downstream fine-tuning task by fixing the initial state using a predetermined random seed. For the navigation tasks, we consider the Point Maze (MAZE) environments with dense, distance-based rewards. The observations include both state and goal information, whereas the actions correspond to planar forces. We use datasets for open, U-shaped, and medium mazes, each containing 1,000 K transitions. We additionally consider several classic-control environments. In InvertedDoublePendulum (PENDULUM), a cart must balance a two-link pendulum, with rewards encouraging upright and stable behavior. In Swimmer (SWIMMER), a three-link agent must move forward using joint torques, with rewards balancing forward progress against control costs. In Reacher (REACHER), a two-link robotic arm must move its fingertip toward a target, with rewards penalizing both the distance to the target and the control effort. In Pusher (PUSHER), a robotic arm must push an object toward a goal, with rewards penalizing the object-goal distance, the fingertip-object distance, and the control effort. For each environment, we use datasets collected by medium and expert policies. The PENDULUM datasets contain 100 K transitions; the REACHER and PUSHER datasets each contain 500 K transitions; and the SWIMMER datasets contain 1,000 K transitions. , Vol. 0, Article 0. Publication date: 0. 0:12 • Bozkurt, Zhang & Motai Finally, we consider four standard legged-robot locomotion environments that are widely used to evaluate offline RL algorithms: Hopper (HOPPER), HalfCheetah (CHEETAH), Walker2d (WALKER), and Ant (ANT). In these environments, a planar one-legged robot, a planar cheetah-like robot, a planar bipedal robot, and a quadrupedal robot, respectively, must move forward by applying joint torques. The observations describe the positions and velocities of the body parts, whereas the rewards combine forward velocity, a healthy-state bonus, and a control penalty. Each environment provides three datasets collected by simple, medium, and expert policies, with each dataset containing 1,000 K transitions. 5.4 Main Results We conduct experiments using an online interaction budget of 1,000 K transitions for each dataset. The main results are presented in Table 1. We compare the O2O-RL approaches based on the value of the final fine-tuned policy selected by each approach. For each environment, we rescale the mean returns using min-max normalization. The minimum is set to the estimated value of a random policy evaluated in the environment, while the maximum is set to the value of the best-performing policy (푣 푖 ∗ 푗 ∗ ; see Problem 1) that could be obtained if future fine-tuning values were known a priori. We repeat each experiment using four random seeds and report the mean values and corresponding standard deviations as percentages. Across the environments and datasets, offline RL approaches typically achieve lower scores and therefore incur greater opportunity loss, or regret, relative to the maximum achievable score of 100% obtained through fine-tuning. These results underscore the importance of effectively using the online interaction budget. For instance, in the MAZE environments, the pretrained policies perform worse on average (RANDOM) than a policy that takes random actions. Even the best pretrained policy (BEST) can perform worse than a random policy, as observed in the MAZE-MEDIUM-DENSE environment. This finding does not necessarily imply that offline training is ineffective. A small number of fine-tuning iterations can substantially improve the performance of pretrained policies, often more rapidly than training from scratch. Without fine-tuning, however, this potential may remain unrealized. The baseline fine-tuning approaches, FTS and FTA, substantially outperform the offline approaches. Committing to a single policy, as in FTS, often yields considerable performance gains because the entire interaction budget is devoted to fine-tuning that policy, as observed in the legged-robot locomotion environments. However, if the selected policy becomes trapped in a local optimum or improves slowly, the entire budget may be spent on a suboptimal policy, as observed in the SWIMMER environment. By contrast, FTA performs better when the interaction budget is relatively large for a given environment, as observed in PENDULUM, REACHER, and PUSHER. However, FTA struggles in environments such as the legged-robot locomotion tasks, where substantial performance gains emerge only after a relatively long fine-tuning warm-up period. Our approach, ACTIVE, actively selects and fine-tunes policies, effectively combining the strengths of both baselines. For example, in the MAZE environments, ACTIVE does not remain committed to policy lineages that fail to improve. Instead, it switches to and fine-tunes other policies, exploring a sufficient portion of the candidate pool and thereby identifying a promising policy lineage, unlike FTS. It then uses the remaining budget for further fine-tuning, ultimately achieving higher scores than FTA. In the legged-robot locomotion environments, ACTIVE avoids spending the entire budget on fine-tuning every candidate, particularly when the candidates are weakly pretrained and require long warm-up periods. Instead, it successfully identifies promising policy lineages that improve during fine-tuning and allocates most of the interaction budget to them. This allocation allows it to overcome the long warm-up period, unlike FTA. Moreover, when progress slows, ACTIVE switches among promising policies, enabling it to achieve higher scores than FTS. Nevertheless, ACTIVE struggles in environments such as SWIMMER and ANT, highlighting an important limitation. In these environments, the policy lineages show little or no indication of improvement for extended periods during the initial phase of fine-tuning, as illustrated for Seed 3 in Fig. 3. Consequently, our approach may exhaust the , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:13 Table 1. Main Results. Environment OfflineOnline RANDOMBESTFTSFTAACTIVE MAZE OPEN-DENSE-8.5± 5.944.5± 30.582.3± 1.294.8± 4.195.7± 4.5 UMAZE-DENSE-19.6± 16.41.5± 12.666.2± 15.892.9± 5.597.5± 2.0 MEDIUM-DENSE-11.9± 15.8-1.4± 28.452.5± 6.796.7± 2.198.6± 0.9 Average-13.3± 10.9 14.8± 15.367.0± 3.194.8± 2.397.3± 2.0 PEN . MEDIUM0.2± 0.00.8± 0.383.1± 6.7100.0± 0.0100.0± 0.0 EXPERT0.2± 0.00.6± 0.195.4± 5.1100.0± 0.0100.0± 0.0 Average0.2± 0.00.7± 0.289.3± 5.1100.0± 0.0100.0± 0.0 SWI . MEDIUM22.6± 2.236.9± 3.446.6± 3.248.8± 6.961.4± 17.7 EXPERT-11.6± 6.054.3± 27.151.2± 7.654.0± 6.667.8± 20.2 Average5.5± 3.545.6± 14.148.9± 3.351.4± 5.264.6± 11.6 REA . MEDIUM60.3± 29.681.4± 6.175.4± 36.698.8± 2.099.9± 0.0 EXPERT83.0± 11.686.2± 11.097.0± 2.0100.0± 0.099.9± 0.0 Average71.6± 11.383.8± 7.686.2± 17.499.4± 1.099.9± 0.0 PUS . MEDIUM80.7± 0.985.2± 1.791.3± 0.495.0± 1.195.0± 0.6 EXPERT74.7± 1.584.6± 1.892.8± 1.493.2± 1.794.2± 1.2 Average77.7± 1.284.9± 1.692.0± 0.794.1± 1.394.6± 0.5 HOPP . SIMPLE-0.3± 0.00.5± 0.544.9± 4.736.1± 9.496.5± 2.1 MEDIUM -0.0± 0.22.6± 2.658.9± 4.929.7± 0.672.9± 25.0 EXPERT-0.2± 0.00.1± 0.155.2± 1.635.5± 10.292.3± 3.3 Average-0.2± 0.11.1± 0.853.0± 1.733.8± 6.487.2± 9.0 CHEE . SIMPLE-0.2± 0.51.6± 0.846.2± 2.844.2± 3.483.8± 9.6 MEDIUM0.0± 0.12.1± 0.449.7± 2.839.6± 4.553.9± 12.8 EXPERT 1.3± 0.22.8± 0.454.8± 3.044.3± 2.287.1± 9.4 Average0.4± 0.22.2± 0.150.2± 1.442.7± 1.774.9± 10.2 WALK . SIMPLE2.6± 0.27.5± 0.546.9± 2.435.1± 8.170.9± 28.2 MEDIUM2.5± 0.48.7± 3.453.7± 5.722.9± 2.874.3± 12.1 EXPERT5.0± 0.77.9± 0.854.2± 2.331.0± 6.360.9± 21.0 Average3.4± 0.38.0± 1.451.6± 2.229.6± 1.468.7± 12.2 ANT SIMPLE-0.3± 3.611.5± 0.338.8± 1.227.0± 4.867.1± 18.1 MEDIUM-1.4± 1.916.6± 0.735.2± 2.516.5± 1.242.9± 25.0 EXPERT6.6± 0.919.3± 4.146.6± 8.923.1± 7.350.9± 11.1 Average1.6± 0.815.8± 1.140.2± 2.622.2± 2.753.6± 10.9 Overall Average16.3± 2.128.6± 2.464.3± 2.063.1± 0.782.3± 1.9 entire interaction budget while attempting to fine-tune all candidate policies in order to identify one that improves. This limitation could potentially be mitigated by incorporating a consistent risk-taking objective when none of the policy lineages are improving, while preserving the exploration needed to distinguish among candidate policies. However, incorporating such an objective could make the overall approach brittle and sensitive to algorithmic choices. We therefore leave this direction for future work. , Vol. 0, Article 0. Publication date: 0. 0:14 • Bozkurt, Zhang & Motai Table 2. Results for Smaller Budget. Environment OfflineOnline RANDOMBESTFTSFTAACTIVE MAZE OPEN-DENSE-8.5± 5.944.5± 30.581.9± 1.394.1± 4.295.3± 4.4 UMAZE-DENSE-19.6± 16.41.5± 12.766.0± 15.792.4± 5.297.4± 2.0 MEDIUM-DENSE-11.9± 15.8-1.4± 28.451.7± 6.796.3± 1.798.4± 0.7 Average-13.4± 10.9 14.9± 15.366.5± 3.194.3± 2.497.0± 2.1 PEN . MEDIUM0.2± 0.00.8± 0.383.1± 6.7100.0± 0.0100.0± 0.0 EXPERT0.2± 0.00.6± 0.191.3± 9.680.9± 33.1100.0± 0.0 Average0.2± 0.00.7± 0.287.2± 7.190.4± 16.5100.0± 0.0 SWI . MEDIUM23.0± 2.437.7± 4.045.1± 4.049.5± 7.351.1± 6.6 EXPERT-11.7± 5.955.0± 26.550.2± 6.655.3± 5.856.1± 6.7 Average5.7± 3.646.4± 13.747.7± 1.752.4± 5.153.6± 4.5 REA . MEDIUM60.3± 29.681.4± 6.175.3± 36.598.8± 2.099.9± 0.0 EXPERT83.0± 11.686.2± 11.097.0± 2.0100.0± 0.099.9± 0.1 Average71.7± 11.383.8± 7.686.2± 17.499.4± 1.099.9± 0.1 PUS . MEDIUM81.1± 1.285.6± 1.791.2± 0.594.8± 1.094.8± 1.0 EXPERT75.2± 1.685.1± 1.692.9± 1.293.6± 1.594.3± 1.0 Average78.1± 1.485.3± 1.592.0± 0.694.2± 1.194.5± 0.6 HOPP . SIMPLE-0.3± 0.00.5± 0.541.1± 4.931.1± 0.682.9± 24.4 MEDIUM-0.0± 0.22.7± 2.755.4± 3.630.1± 0.645.7± 22.6 EXPERT-0.2± 0.00.1± 0.153.1± 4.335.8± 10.278.1± 12.1 Average-0.2± 0.11.1± 0.849.9± 0.732.3± 3.168.9± 12.0 CHEE . SIMPLE-0.2± 0.61.7± 0.844.7± 2.939.7± 2.780.2± 12.4 MEDIUM 0.0± 0.12.3± 0.548.7± 4.335.1± 10.8 41.1± 14.5 EXPERT1.4± 0.23.0± 0.453.8± 2.941.6± 2.687.4± 9.2 Average0.4± 0.22.3± 0.149.1± 2.238.8± 3.669.5± 9.0 WALK . SIMPLE2.7± 0.37.8± 0.746.0± 1.424.1± 1.467.0± 25.8 MEDIUM2.6± 0.49.0± 3.551.8± 3.923.3± 3.059.4± 19.4 EXPERT5.3± 0.68.3± 0.854.0± 1.427.0± 4.155.5± 20.7 Average3.6± 0.38.4± 1.550.6± 1.424.8± 1.560.6± 17.0 ANT SIMPLE-0.3± 3.912.5± 0.637.2± 0.524.0± 3.761.1± 19.3 MEDIUM -1.6± 2.118.5± 0.735.5± 1.918.1± 1.135.5± 20.2 EXPERT7.2± 1.220.8± 4.146.0± 8.820.8± 4.237.9± 13.5 Average1.8± 0.917.3± 1.039.6± 2.521.0± 2.044.8± 5.5 Overall Average16.4± 2.128.9± 2.463.2± 1.760.8± 1.776.6± 1.5 5.5 Smaller Budget We also conduct experiments using a smaller interaction budget of 750 K transitions, with the results presented in Table 2. The offline stage is performed in the same manner as in the main experiments; variations in the normalized offline scores arise from changes in the maximum policy value achievable under the reduced budget. Overall, the offline scores are comparable to those reported in Table 1. , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:15 Table 3. Ablation Study of Window Size Environment Locality Window Size (푤 ) 푤= 3푤= 5푤= 6푤= 7 MAZE OPEN-DENSE0.6± 0.6-1.1± 3.50.8± 1.00.7± 1.1 UMAZE-DENSE0.0± 0.8-0.0± 0.40.2± 0.60.3± 0.5 MEDIUM-DENSE -0.1± 0.6-0.1± 0.40.1± 0.70.5± 0.5 Average0.2± 0.6-0.4± 1.20.4± 0.60.5± 0.5 PEN . MEDIUM-0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 EXPERT0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 Average-0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 SWI . MEDIUM-9.8± 14.4-0.8± 1.4-0.3± 0.5-4.7± 6.1 EXPERT-12.0± 15.7-0.8± 3.2-2.2± 5.1-2.9± 26.0 Average-10.9± 8.4-0.8± 1.7-1.2± 2.6-3.8± 13.0 REA . MEDIUM0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 EXPERT -0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 Average-0.0± 0.00.0± 0.00.0± 0.00.0± 0.0 PUS . MEDIUM-0.4± 1.00.0± 0.10.1± 0.20.1± 0.8 EXPERT -0.5± 0.8-0.3± 0.5-0.3± 0.6-0.3± 0.6 Average-0.4± 0.8-0.1± 0.2-0.1± 0.2-0.1± 0.5 HOPP . SIMPLE-22.3± 26.9-3.8± 5.4-6.6± 10.2-6.6± 10.2 MEDIUM-26.0± 31.6-8.3± 28.3-6.5± 7.6-5.8± 8.2 EXPERT-1.1± 3.41.6± 2.6-16.2± 25.1 -16.2± 25.1 Average-16.4± 17.7-3.5± 8.3-9.8± 6.5-9.5± 7.7 CHEE . SIMPLE0.8± 4.1-0.9± 1.54.7± 10.34.2± 9.5 MEDIUM6.2± 11.8-11.0± 15.0 -17.9± 17.6 -17.9± 17.6 EXPERT-2.6± 2.70.0± 0.00.4± 0.77.2± 11.5 Average1.5± 3.6-4.0± 4.9-4.3± 8.3-2.2± 11.5 WALK . SIMPLE-5.8± 6.2-3.1± 5.44.0± 9.55.3± 17.3 MEDIUM-18.1± 13.52.0± 9.44.1± 7.46.2± 7.7 EXPERT3.2± 15.310.4± 16.410.7± 13.310.0± 12.3 Average-6.9± 9.03.1± 9.26.3± 7.77.2± 9.9 ANT SIMPLE-21.7± 26.0-4.6± 4.1-4.2± 3.6-0.1± 5.3 MEDIUM -22.7± 28.4-1.5± 2.8-0.2± 5.92.1± 23.3 EXPERT -6.6± 18.8-11.0± 11.8 -20.1± 12.5 -20.0± 12.3 Average-17.0± 16.3-5.7± 4.1-8.2± 5.4-6.0± 10.1 Overall Average-5.6± 1.9-1.3± 1.4-1.9± 2.6-1.5± 3.5 As expected, the fine-tuning approaches achieve lower scores under the reduced interaction budget. The smaller budget prevents FTS from fine-tuning its selected policy for enough iterations to attain higher values. It also forces FTA to allocate fewer fine-tuning iterations to each candidate policy. For ACTIVE, the reduced budget further exacerbates the limitation discussed above by decreasing the number of interactions available to identify , Vol. 0, Article 0. Publication date: 0. 0:16 • Bozkurt, Zhang & Motai an improving policy lineage. This effect is particularly pronounced in the legged-robot locomotion environments, where policies require long fine-tuning warm-up periods. 5.6 Ablation Study of Window Size We perform an ablation study to investigate the sensitivity of our approach to the window-size meta-parameter. Specifically, we repeat the experiments using larger and smaller values of the window size,푤, used to fit the local linear model and generate forecasts, while holding all other parameters fixed. The results are presented in Table 3. Overall, our approach exhibits some sensitivity to the choice of 푤 . Increasing the window size has only a limited effect on performance in the low-dimensional environments, with the exception of SWIMMER. In contrast, the high-dimensional locomotion tasks are more sensitive to the choice of 푤, particularly HOPPER and ANT. The setting푤= 3is the smallest window size for which a prediction interval can be constructed. Because the variance of the estimated linear-regression parameters can be relatively high for such a small window, performance decreases, particularly in environments where the value estimates fluctuate substantially during fine-tuning, as shown in Fig. 3. Performance also generally decreases for larger window sizes. However, it improves for ANT, where larger windows produce less noisy forecasts that more reliably identify improving policy lineages. 6 Conclusion To our knowledge, we propose the first active O2O-RL framework that jointly performs policy selection and fine-tuning under a limited online interaction budget. We begin by training a diverse set of candidate policies with offline RL using multiple algorithms and hyperparameter configurations. Because pretrained policies may perform arbitrarily poorly and their performance during fine-tuning may stall or regress depending on the algorithm, hyperparameters, environment, and even random seed, we adaptively allocate scarce online interactions using a UCB criterion based on predicted future policy values. After each fine-tuning iteration, we evaluate the selected policy and update a local linear regression model to refine the corresponding value forecasts and confidence bounds. The framework switches policies whenever another policy attains a higher UCB, thereby allocating the interaction budget efficiently. Across multiple benchmarks, our adaptive policy-selection and fine-tuning procedure consistently outperforms existing O2O-RL baselines. Despite outperforming strong baselines, our approach leaves substantial room for improvement. One promising direction is to leverage exploration rollouts to monitor fine-tuning progress with little or no additional online evaluation. Another direction is to incorporate similarities among policies, as suggested by (Konyushova et al. 2021). An orthogonal avenue is to develop OPE methods tailored specifically to O2O-RL, with greater emphasis on eliminating poor pretrained policies and identifying policies with strong fine-tuning potential. We believe that our framework advances offline RL research by taking another step toward practical, deployable methods for real-world systems in which online interactions are costly or risky. Acknowledgments This work was supported by the Commonwealth Cyber Initiative HV-2Q25-035, HC-2Q25-033, and the Central Virginia Node under the award V-1Q26-001. References Y. Bai, T. Xie, N. Jiang, and Y.-X. Wang. 2019. “Provably efficient q-learning with low switching cost.” Advances in Neural Information Processing Systems, 32. P. J. Ball, L. Smith, I. Kostrikov, and S. Levine. 2023. “Efficient online reinforcement learning with offline data.” In: International Conference on Machine Learning. PMLR, 1577–1594. , Vol. 0, Article 0. Publication date: 0. Active O2O-RL • 0:17 A. K. Bozkurt, X. Xu, S. Zhang, M. Pajic, and Y. Motai. 2026. “Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning.” In: 8th Annual Learning for Dynamics & Control Conference. PMLR. D. Brandfonbrener, W. Whitney, R. Ranganath, and J. Bruna. 2021. “Offline rl without off-policy evaluation.” Advances in neural information processing systems, 34, 4933–4946. G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester. 2021. “Challenges of real-world reinforcement learning: definitions, benchmarks and analysis.” Machine Learning, 110, 9, 2419–2468. F. Farama. 2021. Dataset Reproducibility Guide. https://github.com/Farama-Foundation/d4rl/wiki/Dataset-Reproducibility-Guide. Accessed 10 November 2025. (2021). P. Florence et al.. 2022. “Implicit behavioral cloning.” In: Conference on robot learning. PMLR, 158–168. J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine. 2020. “D4rl: Datasets for deep data-driven reinforcement learning.” arXiv preprint arXiv:2004.07219. S. Fujimoto and S. S. Gu. 2021. “A minimalist approach to offline reinforcement learning.” Advances in neural information processing systems, 34, 20132–20145. S. Fujimoto, H. Hoof, and D. Meger. 2018. “Addressing function approximation error in actor-critic methods.” In: International conference on machine learning. PMLR, 1587–1596. S. Fujimoto, D. Meger, and D. Precup. 2019. “Off-policy deep reinforcement learning without exploration.” In: International conference on machine learning. PMLR, 2052–2062. T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. 2018. “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.” In: International Conference on Machine Learning. PMLR, 1861–1870. J. Huang, J. Chen, L. Zhao, T. Qin, N. Jiang, and T.-Y. Liu. 2022. “Towards Deployment-Efficient Reinforcement Learning: Lower Bound and Optimality.” In: International Conference on Learning Representations. https://openreview.net/forum?id=ccWaPGl9Hq. R. Julian, B. Swanson, G. Sukhatme, S. Levine, C. Finn, and K. Hausman. 2021. “Never Stop Learning: The Effectiveness of Fine-Tuning in Robotic Reinforcement Learning.” In: Conference on Robot Learning. PMLR, 2120–2136. E. Kaufmann, O. Cappé, and A. Garivier. 2016. “On the complexity of best-arm identification in multi-armed bandit models.” The Journal of Machine Learning Research, 17, 1, 1–42. K. Konyushova, Y. Chen, T. Paine, C. Gulcehre, C. Paduraru, D. J. Mankowitz, M. Denil, and N. de Freitas. 2021. “Active offline policy selection.” Advances in Neural Information Processing Systems, 34, 24631–24644. I. Kostrikov, A. Nair, and S. Levine. 2021. “Offline Reinforcement Learning with Implicit Q-Learning.” In: Deep RL Workshop NeurIPS. A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine. 2019. “Stabilizing off-policy q-learning via bootstrapping error reduction.” Advances in neural information processing systems, 32. A. Kumar, A. Zhou, G. Tucker, and S. Levine. 2020. “Conservative q-learning for offline reinforcement learning.” Advances in neural information processing systems, 33, 1179–1191. V. Kurenkov and S. Kolesnikov. 2022. “Showing your offline reinforcement learning work: Online evaluation budget matters.” In: International Conference on Machine Learning. PMLR, 11729–11752. P. Ladosz, L. Weng, M. Kim, and H. Oh. 2022. “Exploration in deep reinforcement learning: A survey.” Information Fusion, 85, 1–22. S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin. 2022. “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble.” In: Conference on Robot Learning. PMLR, 1702–1712. S. Levine, A. Kumar, G. Tucker, and J. Fu. 2020. “Offline reinforcement learning: Tutorial, review, and perspectives on open problems.” arXiv preprint arXiv:2005.01643. T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra. 2016. “Continuous control with deep reinforcement learning.” In: International Conference on Learning Representations. A. Nair, A. Gupta, M. Dalal, and S. Levine. 2020. “Awac: Accelerating online reinforcement learning with offline datasets.” arXiv preprint arXiv:2006.09359. M. Nakamoto, S. Zhai, A. Singh, M. Sobol Mark, Y. Ma, C. Finn, A. Kumar, and S. Levine. 2023. “Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.” Advances in Neural Information Processing Systems, 36, 62244–62269. T. L. Paine, C. Paduraru, A. Michi, C. Gulcehre, K. Zolna, A. Novikov, Z. Wang, and N. de Freitas. 2020. “Hyperparameter selection for offline reinforcement learning.” arXiv preprint arXiv:2007.09055. R. F. Prudencio, M. R. Maximo, and E. L. Colombini. 2023. “A survey on offline reinforcement learning: Taxonomy, review, and open problems.” IEEE Transactions on Neural Networks and Learning Systems, 35, 8, 10237–10257. R.-J. Qin, X. Zhang, S. Gao, X.-H. Chen, Z. Li, W. Zhang, and Y. Yu. 2022. “NeoRL: A near real-world benchmark for offline reinforcement learning.” Advances in Neural Information Processing Systems, 35, 24753–24765. T. Seno and M. Imai. 2022. “d3rlpy: An Offline Deep Reinforcement Learning Library.” Journal of Machine Learning Research, 23, 315, 1–20. http://jmlr.org/papers/v23/22-0017.html. B. Singh, R. Kumar, and V. P. Singh. 2022. “Reinforcement learning in robotic applications: a comprehensive survey.” Artificial Intelligence Review, 55, 2, 945–990. , Vol. 0, Article 0. Publication date: 0. 0:18 • Bozkurt, Zhang & Motai Y. Song, Y. Zhou, A. Sekhari, D. Bagnell, A. Krishnamurthy, and W. Sun. 2023. “Hybrid RL: Using both offline and online data can make RL efficient.” In: The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=yyBis80iUuU. C. Tang, B. Abbatematteo, J. Hu, R. Chandra, R. Martín-Martín, and P. Stone. 2025. “Deep reinforcement learning for robotics: A survey of real-world successes.” In: Proceedings of the AAAI Conference on Artificial Intelligence 27. Vol. 39, 28694–28698. D. Tarasov, V. Kurenkov, A. Nikulin, and S. Kolesnikov. 2023. “Revisiting the minimalist approach to offline reinforcement learning.” Advances in Neural Information Processing Systems, 36, 11592–11620. E. Todorov, T. Erez, and Y. Tassa. 2012. “MuJoCo: A physics engine for model-based control.” In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 5026–5033. doi:10.1109/IROS.2012.6386109. F. Torabi, G. Warnell, and P. Stone. 2018. “Behavioral cloning from observation.” In: International Joint Conference on Artificial Intelligence, 4950–4957. M. Uehara, C. Shi, and N. Kallus. 2025. “A review of off-policy evaluation in reinforcement learning.” Statistical Science. X. Wen, X. Yu, R. Yang, H. Chen, C. Bai, and Z. Wang. 2024. “Towards robust offline-to-online reinforcement learning via uncertainty and smoothness.” Journal of Artificial Intelligence Research, 81, 481–509. R. Zhang, Q. Lv, J. Li, J. Bao, T. Liu, and S. Liu. 2022. “A reinforcement learning method for human-robot collaboration in assembly tasks.” Robotics and Computer-Integrated Manufacturing, 73, 102227. G. Zhou, L. Ke, S. Srinivasa, A. Gupta, A. Rajeswaran, and V. Kumar. 2023. “Real World Offline Reinforcement Learning with Realistic Data Source.” In: 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 7176–7183. W. Zhou, S. Bajracharya, and D. Held. 2021. “Plas: Latent action space for offline reinforcement learning.” In: Conference on Robot Learning. PMLR, 1719–1735. , Vol. 0, Article 0. Publication date: 0.