Paper deep dive
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
Viet Dung Nguyen, Yuhang Song, Anh Nguyen, Jamison Heard, Reynold Bailey, Alexander Ororbia
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 2:13:21 AM
Summary
The paper introduces the 'Master Your Own Expertise' (MYOE) framework, a self-imitation learning approach for robotic agents designed to operate under limited demonstration data. It utilizes a novel 'Queryable Mixture-of-Preferences State Space Model' (QMoP-SSM) to estimate future goal trajectories and compute 'preference regret', which guides policy optimization and mitigates cascading errors common in imitation learning.
Entities (5)
Relation Signals (3)
MYOE ā utilizes ā QMoP-SSM
confidence 100% Ā· The MYOE framework uses the QMoP-SSM to estimate desired goals and compute preference regret.
QMoP-SSM ā computes ā Preference Regret
confidence 95% Ā· QMoP-SSM estimates desired goals used in computing the preference regret.
MYOE ā optimizes ā Robot Control Policy
confidence 95% Ā· The framework enables robotic agents to learn complex behaviors and optimize control policy.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Robot reinforcement learning from demonstrations (RLfD) assumes that expert data is abundant; this is usually unrealistic in the real world given data scarcity as well as high collection cost. Furthermore, imitation learning algorithms assume that the data is independently and identically distributed, which ultimately results in poorer performance as gradual errors emerge and compound within test-time trajectories. We address these issues by introducing the "master your own expertise" (MYOE) framework, a self-imitation framework that enables robotic agents to learn complex behaviors from limited demonstration data samples. Inspired by human perception and action, we propose and design what we call the queryable mixture-of-preferences state space model (QMoP-SSM), which estimates the desired goal at every time step. These desired goals are used in computing the "preference regret", which is used to optimize the robot control policy. Our experiments demonstrate the robustness, adaptability, and out-of-sample performance of our agent compared to other state-of-the-art RLfD schemes. The GitHub repository that supports this work can be found at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.03523v1
- Canonical: https://arxiv.org/abs/2604.03523v1
Trouble viewing inline? Open PDF directly ā
Full Text
51,806 characters extracted from source content.
Expand or collapse full text
OPTIMIZING NEUROROBOT POLICY UNDER LIMITED DEMONSTRATION DATA THROUGH PREFERENCE REGRET Viet Dung Nguyen Rochester Institute of Technology vn1747@rit.edu Yuhang Song Advanced Micro Devices, Inc. sgyson10@liverpool.ac.uk Anh Nguyen University of Liverpool anh.nguyen@liverpool.ac.uk Jamison Heard Rochester Institute of Technology jrheee@rit.edu Reynold Bailey Rochester Institute of Technology rjbvcs@rit.edu Alexander Ororbia Rochester Institute of Technology ago@cs.rit.edu ABSTRACT Robot reinforcement learning from demonstrations (RLfD) assumes that expert data is abundant; this is usually unrealistic in the real world given data scarcity as well as high collection cost. Furthermore, imitation learning algorithms assume that the data is independently and identically distributed, which ultimately results in poorer performance as gradual errors emerge and compound within test-time trajectories. We address these issues by introducing the āmaster your own expertiseā (MYOE) framework, a self-imitation framework that enables robotic agents to learn complex behaviors from limited demonstration data samples. Inspired by human perception and action, we propose and design what we call the queryable mixture-of-preferences state space model (QMoP-SSM), which estimates the desired goal at every time step. These desired goals are used in computing the āpreference regretā, which is used to optimize the robot control policy. Our experiments demonstrate the robustness, adaptability, and out-of-sample performance of our agent compared to other state-of-the-art RLfD schemes. TheGitHubrepositorythatsupportsthisworkcanbefoundat: https://github.com/NACLab/neurorobot-preference-regret-learning. Keywords Active inferenceĀ·Free energy principleĀ·NeuroroboticsĀ·Cognitive controlĀ· Generative world modelsĀ· Reinforcement learning 1 Introduction Typically, popular RLfD algorithms treat expert data as independent and identically distributed (i.i.d.), training in a hybrid manner between supervised learning [1,2,3] and reinforcement learning (RL) [4,5] (and adversarial imitation learning [6,7]). However, imitation learning (IL) and behavior cloning (BC) suffer from a problem known as cascading errors [2,8,9,10,11,12], where small errors accumulate over time in actual trajectories. This causes the robot to deviate from demonstrated examples, leading to policy collapse during online execution. Furthermore, less attention has been given to online RL, in the context of neurorobotics [13, 12], where the robot learns through real-time interaction, given only a limited number of demonstrations. In this work, we overcome the problem of cascading errors and optimize the robot policy in the constrained expert data RLfD setting via the adaptation of a generative model that dynamically produces āpreferredā state trajectories; this results in the learning of a policy through āpreference regretā minimization. Specifically, our contributions are as follows. 1) We leverage model-based reinforcement learning and active inference (AIF) process theory to model perception and action, enabling online learning under limited expert demonstrations. 2) We introduce the queryable arXiv:2604.03523v1 [cs.RO] 4 Apr 2026 Preprint Figure 1: Our Proposed Agent Framework. The agent learns internal representations via encoders, the QMoP- SSM, and a decoder. Imagined future states and preferences guide both policy learning and value estimation. mixture-of-preferences state space model (QMoP-SSM), a goal / preference predictor that estimates future desired trajectories to guide policy learning. 3) We propose a novel policy / behavior learning method that makes use of this future desired trajectory by computing what we call the āpreference regretā. We refer to our approach as the āmaster your own expertiseā (MYOE) framework. Finally, 4) we provide empirical results that shows that our framework outperforms state-of-the-art imitation-learning-augmented model-free and model-based RL baselines. 2 Related Work Perception Modeling. This line of work focuses on biologically-inspired design patterns that aim to provide an agent with an understanding of its niche, or ālivedā world [14,15,16]. This concept has been studied across different research domains such as model-based reinforcement learning [17], active inference [18,19], and variational inference [20]. Central to these approaches is the āgenerative world modelā, which encapsulates two key concepts: 1) inferring internal beliefs from observations, and 2) predicting future observations given those beliefs. To achieve this, world models and recurrent state space models aim to minimize the surprise with respect to model estimations as well as maximize the observation prediction accuracy [21,22], an idea often referred to as the free energy principle [23, 24]: F = D KL [q(s t |o t )ā„ p(s t |s tā1 ,a tā1 )]ā E q(s t ) [lnq(o t |s t )] (1) whereo t ,s t , anda t are the observation, hidden state, and action / control taken at time stept, respectively. Similar to minimizing the evidence lower bound (ELBO) [20], the first term minimizes the KullbackāLeibler (KL) divergence between the approximated posterior and the prior over the hidden state whereas the second term maximizes the systemās prediction accuracy [24]. Later work integrates recurrence into the planning process, e.g., using a āstate space modelā [21,25], which can enhance an agentās ability to form internal beliefs through a perception model. In AIF studies, some agents solve tasks by employing abstract representations of desired goals, a notion referred to as learning prior preferences. In this work, we use the term āpreferenceā to denote the learned / predicted trajectory of prior preferences. Imitation Learning. Imitation learning (IL) methodology originated from ALVINN, an early self-driving system that was designed to mimic human driving behavior [1]. This led to the development of behavioral cloning (BC), which frames IL as a supervised learning problem [26], where observations are inputs and expert actions are labels used to train a āclonedā policy. However, BC suffers from instability due to the problem of ācascading errorsā, where small planning mistakes accumulate during real-world execution, driving the agent into trajectory distributions that have not been observed in the training data [2,8]. This issue, also referred to as ācovariate shiftā [27] in the underlying dataset distribution, has been mitigated through algorithms such as the dataset 2 Preprint Figure 2: QMoP-SSM Learning Architecture. QMoP separates latent learning into two interconnected parallel processes: ārepresentationā and āpreferenceā learning. While the model predicts future latent given previous actions, future preference state trajectories are guided by provided and learned goals. aggregation (DAgger) method [3], inverse reinforcement learning [27], generative adversarial methods [6,7], and self-imitation learning [28]. However, relatively little work has explored scenarios where robotic agents must learn from a limited number of demonstrations, a more practical setting with a lower cost in terms of time, effort, and computational resources. In this work, we tackle this challenge by leveraging the concept of perception modeling as well as AIF process theory to construct agents that learn from limited expert demonstrations while mitigating the impact of cascading errors in the context of online learning. 3 Robot Learning through Preference Regret Our agent operates on partially-observable Markov decision processes (POMDPs) [17] where the observation space includes visual sensory inputs, proprioceptive states, object states, and desired goals. Agents must learn in real-time through environmental interactions and are only provided with a small number (five) episodes of expert demonstrations. To solve robotic tasks, our approach leverages perception modeling, enabling the ability to deduce both future observations and goal trajectories; see Section 3.1. Based on the derived future representations and preferences, we train our systemās behavioral policy by minimizing a term called āpreference regretā. This term is computed from the resulting goal trajectory from our proposed perception model; it is built on the advantage estimation of proximal policy optimization (PPO) frameworks [17,29] and the behavioral learning element of AIF process theory [30, 31, 32]; see Section 3.2. The training process can be viewed in Figure 1. 3.1 The Queryable Mixture-of-Preferences State Space Model Motivated by efforts in active inference and prior preference learning [33,34], we construct an agent that predicts future preference trajectories based on its own learned goal representations. Consequently, agent policy adaptation is influenced by both āimaginedā future preferences as well as the expert data, which mitigates the cascading errors effect. Concretely, provided any goal, (e.g., language instructions, goal images / states, or learnable goal vectors), the QMoP-SSM āqueriesā the agentās goals and then predicts a roll-out of future (latent) representations and (prior) preferences. These trajectories then drive the agentās policy toward desired goal states while maximizing the utility values contained within expert data as well as optimizing final rewards; see Figures 1 and 2 for details of the agent and its learning process. Formulation. Our framework integrates world model building [21,35,36], (recurrent) state space models [37,38, 39,22], and active inference [24], yielding an agent model that continuously predicts future states to guide its ability to plan actions. Our perception model, QMoP-SSM, which also predicts future āpreferencesā (or representations of 3 Preprint optimized states), is expressed formally as follows: Representation Posterior: q(s o t |o t ,h t ) Representation Prior: p(s o t |s o tā1 ,h tā1 ,a tā1 ) Representation Likelihood: p(o t |s o t ,h t ) Representation Reward: p(r t |s o t ,h t ) Preference Posterior: q(s p t |o t ,h t ) Preference Prior: p(s p t |s p tā1 ,g Īø tā1 ,h tā1 ) Preference Reward: p(r t |s p t ,h t ). (2) whereo t is the observation (at timet),h t is the recurrent hidden state,s o t is the world (representation) state,s p t is the preference state, andr t is the (extrinsic) reward.g Īø is the learnable goal, which is optimized via preference learning in our perception model. In this work, we use the term g Īø to denote the tensor that encapsulates all provided goal modalities. These modalities may include (but are not limited to) learnable tensors, natural language instructions, and goal/target images. Finally, in line with deep AIF [40] and recurrent state space modeling [21,22] methodology, we parameterize the representation and preference modules of our agent with artificial neural networks (ANNs) that are conditioned on a gated recurrent unit (GRU) recurrent networkās hidden neural states [41]. Figure 2 illustrates the inference and learning components of our QMoP-SSM. The agent first learns the representation model by continuously adjusting its states o so as to minimize the surprisal between the approximated posteriorq(s o t |o t ,h t ) and the priorp(s o t |s o tā1 ,h tā1 ,a tā1 ) while furthermore maximizing its observation prediction accuracy from the world representation likelihoodp(o t |s o t ,h t ) . This process resembles maximizing the ELBO [20,42] as well as minimizing what is known as the marginal free energy [43,44,24,45]. Next, our system learns to infer the preference states by predicting the successful trajectories, which are known from the expert demonstrations or collected from the agentās interaction with the environment. To achieve this, we minimize the distance between the preference state distribution and the actual representation state distribution. Formally, given that the perception model QMoP-SSM is parameterized byĪø, the free energy (learning) objective can be described as follows: arg min Īø F t (Īø) =F o +F o,KL +F r +F p,KL +F dist (3a) F o =āE q Īø (s o t ) [ln(p Īø (o t |s o t ,h t )](3b) F o,KL = D KL [q Īø (s o t |o t ,h t )ā„ p Īø (s o t |s o tā1 ,a tā1 ,h tā1 )](3c) F r =āE q Īø (s o t ) [ln(p Īø (r t |s o t ,h t )](3d) F p,KL = D KL [q(s p t |o t ,h t )ā„ p(s p t |s p tā1 ,g Īø ,h tā1 )]ā m t (3e) F dist =D[q(s p t |o t ,h t )ā„ q(s o t |o t ,h t )]ā m t .(3f) In the above, the terms in Equations 3b and 3d seek to maximize prediction accuracy for observations and rewards. The terms in Equations 3c and 3e minimize the KL divergence, (i.e., the amount of surprise), between the approximate posterior and the prior distribution. Finally, the term in Equation 3f aims to minimize the distance between the posterior over the states of the preference and the representation models. This helps the future predicted preference state distributions align with the successful state distributions of the representation model. Choices for distance functionDcan include the mean (squared) Euclidean distance, negative cosine similarity, and Manhattan distance. Note that the (binary) maskm t represents whether the state attis part of a successful trajectory (or not) and is element-wise multiplied (ā) with the corresponding preference objectives in order to constrain the preference model to only predict the next goal states. Finally, when learning how to predict rewards, the preference reward objectiveāE q Īø (s p t ) [ln(p Īø (r t |s p t ,h t )]can be omitted since the weights may be shared among the reward modelsp Īø (r t |s o t ,h t )andp Īø (r t |s p t ,h t )(i.e., the preference state aims to estimate the representation state within successful trajectories). Mixture of Preferences. As discussed earlier, our proposed state space model (SSM) predicts future preference trajectories that are guided by goal information. However, multiple feasible trajectories can lead to the same goal (state); this introduces noise into the preference learning process, potentially causing mode collapse [46]. To address this challenge, we extend the SSM to operate as a mixture of stochastic distributions for both the prior and posterior. Specifically, given a mixture of sizeM ā N, we then estimate a collection of preference state distributions p i Īø (s p t |s p tā1 ,q Īø ,h tā1 ) M i=1 . This formulation provides three key advantages: 1) varied-goal environments ā an individual mixture component can specialize for a specific trajectory type, facilitating comprehensive coverage of all preferred behaviors; 2) multiple preference trajectories expand the desired behaviorsā state space, which prevents mode collapse in the estimated preference distributions; and 3) a diverse mixture of preferences yields 4 Preprint Figure 3: Cumulative reward (y-axis) across 1 million interaction/training steps (x-axis) for different agents. high entropy ā this promotes exploration as predicted future preferences are used as the learning signal to drive policy adaptation. Note that this also aligns with past RL effort that has demonstrated that entropy maximization enhances exploration [47, 48, 49, 50]. Formally, given a mixture ofMpredicted preference distributionsp i Īø (s p t |s p tā1 ,q Īø ,h tā1 ) M i=1 , we designate / choose one by computing a weighted linear combination over all of the distributionsā parameters within the mixture: Ģp Īø (s p t ) = M X i=1 h Ļ(z)ā s p,i t i , z = W q Īø ⤠+ b (4) wheres p,i t ā¼ p i Īø (s p t |s p tā1 ,q Īø ,h tā1 ) andq Īø ā R d q is the (learnable) encoded goal / query,W ā R MĆd q and bā R MĆ1 produce the logitz, andĻ(z)is the softmax function applied to the mapped logit component so as to produce the mixture probability. For each preference trajectory, we seek high modality coverage, avoidance of mode collapse, and exploratory signals for downstream policy learning. To achieve these objectives, we maximize the entropy of the predicted mixture of preferences in the following manner: arg max Īø L t (Īø) = H [ Ģp Īø (s p t )]ā α℠Ģs p t ā„ 2 , Ģs p t ā¼ Ģp Īø (s p t ) (5) whereH [ Ģp Īø (s p t )]is the entropy of the predicted preference distribution andā„ Ģs p t ā„ 2 is the L2 regularization term used to stabilize the preference prediction. Finally,α > 0is the coefficient to stabilize the learning objective, e.g., α = 0.1. This last objective enlarges preference mode coverage while maintaining the smoothness of the combined distribution. 3.2 MYOE: Behavioral Learning through Preference Regret In order to enable a form of effective behavior learning that leverages expert demonstrations while further still maximizing the final rewards obtained via environmental interactions, we employ a preference-guided policy optimization framework based on regret minimization principles [51]. As a result, we integrate both the predicted representationp(s o ) and preferencep(s p ) distributions produced by the QMoP-SSM to guide the adaptation of our systemās policy. Specifically, to train the behavioral model, we utilize policy gradients [52,53,54], actor- critic [55,17,47,22,25], and generalized advantage estimation (GAE-Ī») [29], leveraging state-action-value estimation to nudge the agent toward trajectories with higher rewards. Formally, we consider the following modules: Policy: Ļ Ļ (a Ļ |s o Ļ ,s p Ļ ,h Ļ ) Value: V ν (v Ļ |s o Ļ ,h Ļ ); Target Value: V Ⲡν (v Ļ |s o Ļ ,h Ļ ). whereĻ,ν, andν ā² denote the parameters of the policy network, the value network, and the target value network, respectively. Given the future planning stepĻ, imagination horizonH, and the learned goalg Īø , the agent works to predict sequences of future representations of observation and preference state distributions as follows: (p Īø (s o Ļ |a Ļā1 ,h tā1 ), Ģp Īø (s p Ļ |s p Ļā1 ,g Īø ,h tā1 )) H Ļ=1 (6) wherea Ļ ā¼ Ļ Ļ (a Ļ |s o Ļ ,s p Ļ ,h Ļ ) represents actions sampled from the current policy and Ģp Īø (s p Ļ |s p Ļā1 ,g Īø ,h tā1 )) denotes the QMoP-SSMās mixture preference distribution. This formulation allows downstream behavior learning to benefit from both environmental dynamics and the learned preferences. The Preference-Regret. To effectively leverage the preference trajectories yielded by our QMoP-SSM, we introduce a novel intrinsic reward augmentation formulation to compute the target Q-value, inspired by the concept 5 Preprint of āregret minimizationā in multi-armed bandit problems [56,17,57]. Classical regret quantifies the difference between rewards obtained by an agent and those retrieved from optimal actions that could have been executed to interact with the environment. In our context, we extend the concept of āregretā by defining āpreference regretā as the difference between actual rewards and rewards estimated using preference state trajectories. We also incorporate intrinsic motivation [49,58] into our framing based on the agentās state distribution entropy; this is similar to AIF process theory where the agent seeks to reduce the epistemic uncertainty within future estimated state distributions [30,32,31,59]. Thus, our agentās intrinsic reward, in line with expected free energy [30,15,24], is: R Ļ = r o Ļ |z reward ā (r p Ļ ā r o Ļ ) | z preference regret +α H p(s o Ļ |s o Ļā1 ,a Ļā1 ,h Ļā1 ) |z representation state entropy (7) wherer o Ļ ā¼ p Īø (r Ļ |s o Ļ ,h Ļ )andr p Ļ ā¼ p Īø (r Ļ |s p Ļ ,h Ļ )represent predicted rewards conditioned on estimated future representations and preferences distributions, respectively.a Ļā1 ā¼ Ļ Ļ (a Ļā1 |s o Ļā1 ,s p Ļā1 ,h Ļā1 ) denotes the policyās sampled actions.α ā 0.0003is the scaling coefficient for stability, with its value chosen according to the soft actor-critic algorithm convention [47,48]. The āpreference regretā(r p Ļ ā r o Ļ )penalizes deviations from preferred trajectories while preserving online task performance. Therefore, the behavioral model learns to solve tasks while still following the preferred trajectories guided by the learned goal(s). Following typical temporal difference (TD) learning frameworks [29,17,60], we train the value functionV ν to estimate the expected discounted returnG Ļ (which is computed from discounted, augmented rewardsR Ļ ). In order to achieve this, we first compute the TD error as follows: Ī“ Ļ = R Ļ + γV ν (v Ļ+1 |s o Ļ+1 ,h Ļ+1 )ā V ν (v Ļ |s o Ļ ,h Ļ ) (8) whereγ ā [0, 1], the discount factor, determines the relative importance of future rewards. Finally, to obtain stable low-variance policy estimates, we employ GAE-Ī»[29] to compute the advantage valueA Ļ and the target returns G Ļ as below: A Ļ = H X Ļ=1 (γλ) Ļ Ī“ Ļ ; G Ļ =A Ļ + V ν (v Ļ |s o Ļ ,h Ļ ) (9) whereĪ»controls the bias-variance trade-off in advantage estimation. WhenĪ» = 0, the estimates reduce to a one-step TD method (low variance, high bias) whereas Ī» = 1 approaches Monte Carlo estimation (high variance, low bias). This term weights the temporal difference errorsĪ“ Ļ across the trajectory, enhancing optimization stability. Value and Policy Learning. After computing the targetĪ»-weighted return valueG Ļ from Equations 7, 8, and 9, we train the value estimator f ν in the following manner: arg min ν L Ļ (ν) =ā„ V ν (v Ļ |s o Ļ ,h Ļ )ā sg(G Ļ )ā„ 2 (10a) +α℠V ν (v Ļ |s o Ļ ,h Ļ )ā sg V Ⲡν (v Ļ |s o Ļ ,h Ļ ) ā„ 2 .(10b) āsgā denotes the stop gradient operator and scalarαinduces training stability. Equation 10a minimizes the mean squared error between the predicted value and theĪ»-return; Equation 10b reduces the weight gap between the value and target networks (to improve training stability [61,60]). Target networkV Ⲡν weights are updated (per step) using the source value network V ν via an exponential moving average [62]. Next, policy learning is performed through advantage value maximization via policy gradients as follows: arg max Ļ L Ļ (Ļ) =L adv +L ac +L exp (11a) L adv = lnĻ Ļ (a Ļ |s o Ļ ,s p Ļ ,h Ļ ) sg(A Ļ )(11b) L ac =αH [Ļ Ļ (a Ļ |s o Ļ ,s p Ļ ,h Ļ )](11c) L exp =ā βm Ļ āā„ a Ļ ā a ā Ļ ā„ 2 (11d) wherea Ļ ā¼ Ļ Ļ (a Ļ |s o Ļ ,s p Ļ ,h Ļ )is the imagined action(s) anda ā Ļ is the actual action(s) taken.m Ļ is a mask that specifies whether the actiona ā is from the expert or not. The (stability) coefficient values areα = 0.0003and β = 0.5. In addition to advantage (value) maximization ā of the termL adv in Equation 11b ā our approach draws inspiration from the fusion of model-free RL, BC [4], and the actor refreshing mechanism used in active predictive coding [63]. Specifically, the termL exp (Equation 11d) provides the agent with a small reward when its actions follow the expert demonstrations. Finally, the termL ac (Equation 11c) denotes the rewards supplied to the agent for entering states where its policy incurs high entropy. This quantity is often referred to as ācuriosityā [31] and has been widely used to enhance exploration. 6 Preprint Task/AgentMYOEDrm-BCPPO-BCLAIfO K Kettle0.97± 0.060.90± 0.110.01± 0.020.27± 0.42 K Light Switch1.00± 0.010.51± 0.490.00± 0.000.00± 0.00 K Microwave0.75± 0.430.67± 0.340.00± 0.000.00± 0.00 K Slide Cabinet0.99± 0.030.74± 0.430.28± 0.420.57± 0.44 R Lift1.00± 0.010.62± 0.260.00± 0.000.00± 0.01 R Door0.99± 0.020.98± 0.030.00± 0.000.00± 0.00 M Button Press1.00± 0.000.25± 0.430.04± 0.050.00± 0.00 M Button Press Wall1.00± 0.000.00± 0.000.00± 0.000.00± 0.00 M Coffee Button1.00± 0.010.99± 0.030.07± 0.120.54± 0.36 M Coffee Pull0.77± 0.130.00± 0.000.00± 0.000.00± 0.00 M Coffee Push0.85± 0.100.00± 0.000.00± 0.000.01± 0.02 M Door Close1.00± 0.001.00± 0.000.25± 0.430.25± 0.43 M Door Lock1.00± 0.020.98± 0.040.11± 0.130.06± 0.11 M Door Unlock1.00± 0.001.00± 0.000.00± 0.000.00± 0.00 M Drawer Close1.00± 0.001.00± 0.000.95± 0.091.00± 0.00 M Drawer Open1.00± 0.001.00± 0.000.01± 0.030.02± 0.04 M Faucet Close1.00± 0.021.00± 0.000.00± 0.000.00± 0.00 M Handle Press1.00± 0.000.99± 0.020.30± 0.230.24± 0.19 M Handle Pull Side1.00± 0.000.00± 0.000.00± 0.000.00± 0.00 M Handle Pull1.00± 0.000.49± 0.490.02± 0.040.04± 0.06 M Peg Insert Side0.08± 0.140.00± 0.000.00± 0.000.00± 0.00 M Reach0.36± 0.170.02± 0.040.00± 0.000.00± 0.00 M Plate Slide0.97± 0.050.00± 0.000.02± 0.040.18± 0.33 M Plate Slide Back Side0.75± 0.430.00± 0.000.00± 0.000.00± 0.00 M Plate Slide Side1.00± 0.010.99± 0.020.12± 0.210.00± 0.00 M Plate Slide Back1.00± 0.000.46± 0.460.00± 0.000.00± 0.00 M Peg Unplug Side1.00± 0.010.00± 0.000.00± 0.010.02± 0.04 M Soccer0.30± 0.110.00± 0.000.00± 0.010.00± 0.00 M Reach Wall0.94± 0.100.00± 0.000.00± 0.000.00± 0.00 M Sweep Into0.47± 0.170.00± 0.010.00± 0.000.00± 0.01 M Window Open1.00± 0.001.00± 0.000.01± 0.020.10± 0.18 M Window Close1.00± 0.001.00± 0.000.00± 0.000.00± 0.00 Table 1: We report the evaluation success rate of the last100episodes of MYOE (ours) as compared to other model/RL baselines. Cyan cells represent the best performing agents within the corresponding task. Note that āDrmā is short for āDreamerV3.ā The Effect of Preference Regret. Lemma 1. Minimizing preference regret as an internal reward in the advantage computation guides the agent toward preferred trajectories while maintaining the ability to maximize the final reward when preferences are sub-optimal. Proof. We divide the proof into two cases. Case 1: High-quality Preferences. Assume that the reward model predicts high reward for preference states. When the agentās imagined trajectory deviates from the predicted preference trajectory,r p Ļ > r o Ļ , it yields negative preference regretā(r p Ļ ā r o Ļ ). The internal rewardR Ļ is then penalized for deviation from the preference. Conversely, when the imagined trajectory aligns with the preference,ā(r p Ļ ā r o Ļ )ā 0, maintaining or increasing the advantage value, it reinforces actions that follow the preferred trajectory. Case 2: Sub-optimal Preferences. Assume that the preference trajectory yields lower reward than the actual policy, e.g., due to suboptimal expert demonstrations. In this case,r p Ļ < r o Ļ results in a positive preference regret (ā(r p Ļ ā r o Ļ ) > 0). This contributes to the increased advantage value and thus reinforces the agent to pursue higher environmental reward (r o Ļ ) trajectories rather than blindly imitating the expert. Consequently, this mechanism enables the agent to adaptively balance self / expert-imitation and real-time optimization: avoiding naive behavior cloning ā which suffers from cascading errors due to distributional shift ā while maximizing the final rewards in online learning. 7 Preprint Task/AgentMYOEMBCMBC-VAE MBC-RNN K Kettle0.97± 0.060.65± 0.460.66± 0.470.67± 0.47 K Light Switch1.00± 0.010.23± 0.220.12± 0.200.24± 0.17 K Microwave0.75± 0.430.50± 0.500.49± 0.490.50± 0.50 K Slide Cabinet0.99± 0.030.08± 0.120.16± 0.170.08± 0.15 R Lift1.00± 0.010.38± 0.290.70± 0.230.46± 0.25 R Door0.99± 0.020.96± 0.060.81± 0.260.97± 0.04 M Button Press1.00± 0.000.00± 0.000.08± 0.100.29± 0.24 M Button Press Wall1.00± 0.000.41± 0.410.49± 0.420.00± 0.00 M Coffee Button1.00± 0.010.62± 0.310.82± 0.130.15± 0.16 M Coffee Pull0.77± 0.130.10± 0.170.05± 0.060.07± 0.09 M Coffee Push0.85± 0.100.14± 0.110.24± 0.130.10± 0.13 M Door Close1.00± 0.000.04± 0.080.36± 0.350.27± 0.42 M Door Lock1.00± 0.020.21± 0.140.57± 0.090.36± 0.24 M Door Unlock1.00± 0.000.33± 0.240.36± 0.150.34± 0.09 M Drawer Close1.00± 0.001.00± 0.000.99± 0.030.99± 0.02 M Drawer Open1.00± 0.000.31± 0.340.80± 0.170.02± 0.04 M Faucet Close1.00± 0.020.71± 0.290.77± 0.120.72± 0.24 M Handle Press1.00± 0.000.65± 0.200.72± 0.260.64± 0.15 M Handle Pull Side1.00± 0.000.03± 0.060.00± 0.000.02± 0.03 M Handle Pull1.00± 0.000.28± 0.220.27± 0.180.50± 0.26 M Peg Insert Side0.08± 0.140.04± 0.070.01± 0.020.06± 0.06 M Reach0.36± 0.170.19± 0.130.20± 0.130.21± 0.23 M Plate Slide0.97± 0.050.00± 0.000.08± 0.140.16± 0.19 M Plate Slide Back Side0.75± 0.430.24± 0.410.66± 0.320.49± 0.40 M Plate Slide Side1.00± 0.010.44± 0.450.71± 0.220.24± 0.41 M Plate Slide Back1.00± 0.000.77± 0.150.86± 0.110.42± 0.36 M Peg Unplug Side1.00± 0.010.20± 0.140.23± 0.230.17± 0.18 M Soccer0.30± 0.110.03± 0.040.20± 0.110.23± 0.15 M Reach Wall0.94± 0.100.18± 0.110.09± 0.080.15± 0.14 M Sweep Into0.47± 0.170.08± 0.120.17± 0.080.06± 0.11 M Window Open1.00± 0.000.30± 0.270.45± 0.190.34± 0.15 M Window Close1.00± 0.000.81± 0.180.93± 0.090.88± 0.13 Table 2: We report the evaluation success rate of the last100episodes of MYOE (ours) in comparison to other baselines. Cyan cells represent the best performance agents within the corresponding task. 4 Experimental Results Baselines. We seek to demonstrate that our agent outperforms state-of-the-art RLfD baselines in a limited expert data scenario. Specifically, we compare our agent (MYOE) to relevant model-based and model-free RL with IL-augmentation and adversarial IL schemes, namely: 1) DreamerV3 [60] with BC (Drm-BC); 2) PPO [64,29] with BC (PPO-BC); and 3) PPO with adversarial IL [7,6] (LAIfO). Additionally, we also compare to self-supervised IL systems, i.e., BC agents that train on their successful data in addition to the expert data, to demonstrate the robustness of āpreference regretā optimization over the cascading error. These particular baseline systems includes multimodal BC (MBC), MBC with recurrent neural network (MBC-RNN), and MBC with variational autoencoder [20] (MBC-VAE). All of the agentsā total number of parameters are constrained toā 12million and each agent trains with the environment for1million steps. For each agent and environment, we conduct4trials and measure performance by calculating the mean and standard deviation of the success rate (SR) from the last100 non-learning episodes and their corresponding cumulative rewards. Benchmarks. We evaluate our proposed agent and relevant baselines on three sets of simulated environments consisting of32tasks in total:4Franka Kitchen [65] (K) tasks,26Meta-World [66] (M) tasks, and2Robosuite [67] (R) tasks. All of the environments have a sparse reward setting and use visual frames and states as observation, with the Franka Kitchen providing language as an additional modality. For each environment, we collect only5 episodes of expert demonstration data to use in the sparse reward setting and all agents / baselines benefit from this data pool (each episode consists ofā 250,200, and1000steps in Franka Kitchen, Meta-World, and Robosuite, respectively). All environments have been modified to encourage cascading errors, (i.e., Franka Kitchen adds noise to the robot actions whereas Robosuite and Meta-World change their goal in each episode). 8 Preprint Figure 4: Our proposed MYOE solving the āreachā task when integrated into the ā7botā robot (top) and our agent model solving the āblock pickingā task when integrated into the PX100 robot (bottom). MYOEDrm-BCPPO-BC MBC-RNN 0.34± 0.110.17± 0.010.29± 0.110.00± 0.00 Table 3: āReachā task success rate on the ā7botā. MYOEDrm-BCPPO-BCLAIfO 0.89± 0.050.81± 0.050.78± 0.060.79± 0.06 80.78± 4.4872.34± 4.9366.78± 6.3067.96± 6.19 Table 4: āBlock pickingā task success rate (top) and episodic cumulative reward (bottom) on the PX100 robot. Simulation Results. Overall, our agent converges faster for most tasks in comparison to other baselines; see Figure 3. This result demonstrates that our proposed method is capable of utilizing potentially suboptimal and limited demonstrations while still being able to optimize for the task at hand. This is due to the fact that our agent continuously optimizes its preference, which guides its policy toward expert policies when they are optimal and toward overall reward signal values regardless of the expert given that some expert policies might be suboptimal. Baseline RLfD algorithms can also converge in many tasks because of the direct expert action signal; however, some expert policies are not optimal (such as M Window Close), affecting the training performance of the baseline RLfD algorithms; see Figure 3, Table 1, and 2. Additionally, sparse reward settings reduce the effectiveness of traditional RLfD algorithms, hindering their learning ability (such as in LAIfO and PPO-BC). Finally, we observe that baseline RLfD algorithms and self-supervised IL systems does not optimize at all in some tasks (such as in M Button Press Wall), suggesting that they suffer from cascading errors, resulting in the acquisition of policies that have diverged. Real-World Neurorobot Online Learning. To further investigate our methodās robustness compared to state-of- the-art RLfD baselines, we conducted two real-world online RLfD experiments using the ā7botā [68] (6-DOF) and the PX100 robot [69] (4-DOF). The first task requires the ā7botā to reach a position given a target joint configuration and the second task requires the PX100 robot to pickup a block. Due to the computational and logistical constraints inherent to real-world in-robot online learning ā including the non-parallelizable nature of physical systems ā we focus our comparison on4representative agents over4trials: our MYOE model, DRM-BC, PPO-BC, and MBC-RNN (for ā7botā) / LAIfO (for PX100 robot). To simulate realistic learning scenarios with limited expert data, we constrain each agent to train for only125episodes while providing5episodes of demonstrations (10 steps per episode). The expert data is recorded to encourage cascading errors (such as shaking the robot arm before picking up the block). We evaluate performance by recording the mean and standard deviation of the task success rate (and task cumulative reward for the block picking task) over the last20episodes. The results show that, across all tasks, MYOE outperforms all other RLfD and self-supervised IL baselines (see Table 3, 4, and Figure 4). This validates that optimizing preference regret provides notable advantages for neurorobotic task optimization, especially in resource-constrained RLfD scenarios where cascading errors is mostly present. 9 Preprint 5 Conclusions In this work, we constructed an RLfD agent framework that optimizes (neuro)robotic task solvers using few expert demonstration samples. We proposed a perception learning paradigm, QMoP-SSM, that makes use of goal learning and preference trajectory estimation. Furthermore, we formulated an agent behavioral learning framework that leverages preference information from the perception model. Specifically, by introducing a concept called āpreference regretā, our agentās policy is capable of learning tasks quickly and accurately as compared to other state-of-the-art model-based and model-free RL (with IL-augmentation and self-supervised IL) systems on both simulation and real-world robot platforms. Limitations and Future Work. Although MYOE outperforms other state-of-the-art baselines, for more complex tasks such as āM Soccer,ā our approachās performance is not likely to surpass that of experts. We assume that these scenarios are too difficult (due to small object pixel areas) such that our preference learning and imagination scheme would struggle to effectively utilize the image observation and properly extract useful temporal dependencies within the data. Future work should consider improving MYOEās preference trajectory estimation, utilizing different autoregressive models such as deep Kalman filters [38,39]. Finally, one could considering enhancing our architectureās ability to encode information in its posterior from subtle pixel changes in order to better capture small objects within a (neuro)robotās workspace. References [1] D. A. Pomerleau, āAlvinn: An autonomous land vehicle in a neural network,ā in Advances in Neural Information Processing Systems, D. Touretzky, Ed., vol. 1. Morgan-Kaufmann, 1988. [Online]. Available: https://proceedings.neurips.c/paper/1988/file/812b4ba287f5e0bc9d43bbf5bbe87fb-Paper.pdf [2]S. Ross and D. Bagnell, āEfficient reductions for imitation learning,ā in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, Y. W. Teh and M. Titterington, Eds., vol. 9. Chia Laguna Resort, Sardinia, Italy: PMLR, 13ā15 May 2010, p. 661ā668. [Online]. Available: https://proceedings.mlr.press/v9/ross10a.html [3]S. Ross, G. Gordon, and D. Bagnell, āA reduction of imitation learning and structured prediction to no-regret online learning,ā in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, G. Gordon, D. Dunson, and M. DudĆk, Eds., vol. 15.Fort Lauderdale, FL, USA: PMLR, 11ā13 Apr 2011, p. 627ā635. [Online]. Available: https://proceedings.mlr.press/v15/ross11a.html [4]S. Fujimoto and S. S. Gu, āA minimalist approach to offline reinforcement learning,ā in Advances in Neural Information Processing Systems, vol. 34, 2021, p. 20 132ā20 145. [5] A. Mandlekar, D. Xu, R. MartĆn-MartĆn, S. Savarese, and F.-F. Li, āLearning to generalize across long-horizon tasks from human demonstrations,ā arXiv preprint arXiv:2003.06085, 2020, revised version (v2) released June 23, 2021. [6]J. Ho and S. Ermon, āGenerative adversarial imitation learning,ā in Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.c/paper_files/paper/2016/file/ c7e2b878868cbae992d1fb743995d8f-Paper.pdf [7]V. Giammarino, J. Queeney, and I. Paschalidis, āAdversarial imitation learning from visual observations using latent information,ā Transactions on Machine Learning Research, 2024. [Online]. Available: https://openreview.net/forum?id=ydPHjgf6h0 [8]J. A. D. Bagnell, āAn invitation to imitation,ā Carnegie Mellon University, Pittsburgh, PA, Tech. Rep. CMU-RI-TR-15-08, March 2015. [9]F. Codevilla, E. Santana, A. Lopez, and A. Gaidon, āExploring the limitations of behavior cloning for autonomous driving,ā in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, p. 9328ā9337. [10]J. D. Chang, M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun, āMitigating covariate shift in imitation learning via offline data with partial coverage,ā in Advances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021. [Online]. Available: https://openreview.net/forum?id=7PkfLkyLMRM 10 Preprint [11]S. Seo, B.-J. Lee, J. Lee, H. Hwang, H. Yang, and K.-E. Kim, āMitigating covariate shift in behavioral cloning via robust stationary distribution correction,ā in The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Online]. Available: https://openreview.net/forum?id=lHcvjsQFQq [12]V. D. Nguyen, Z. Yang, C. L. Buckley, and A. Ororbia, āSr-aif: Solving sparse-reward robotic tasks from pixels with active inference and world models,ā in 2025 IEEE International Conference on Robotics and Automation (ICRA), May 2025, p. 6510ā6518. [13]H. J. Chiel and R. D. Beer, āThe brain has a body: adaptive behavior emerges from interactions of nervous system, body and environment,ā Trends in neurosciences, vol. 20, no. 12, p. 553ā557, 1997. [14]R. C. Conant and W. Ross Ashby, āEvery good regulator of a system must be a model of that system,ā International journal of systems science, vol. 1, no. 2, p. 89ā97, 1970. [15]K. Friston, āLife as we know it,ā Journal of the Royal Society, Interface / the Royal Society, vol. 10, p. 20130475, 06 2013. [16]A. Ororbia and K. Friston, āMortal computation: A foundation for biomimetic intelligence,ā 2024. [Online]. Available: https://arxiv.org/abs/2311.09589 [17]R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction, 2nd ed.The MIT Press, 2018. [Online]. Available: http://incompleteideas.net/book/the-book-2nd.html [18] T. Parr, G. Pezzulo, and K. J. Friston, Active Inference: The Free Energy Principle in Mind, Brain, and Behavior. The MIT Press, 03 2022. [Online]. Available: https://doi.org/10.7551/mitpress/12441.001.0001 [19] A. G. Ororbia and A. Mali, āBackprop-free reinforcement learning with active neural generative coding,ā Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 1, p. 29ā37, Jun. 2022. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/19876 [20]D. P. Kingma and M. Welling, āAuto-encoding variational bayes,ā in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014. [21]D. R. Ha and J. Schmidhuber, āWorld models,ā ArXiv, vol. abs/1803.10122, 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:4807711 [22]D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, āDream to control: Learning behaviors by latent imagination,ā in International Conference on Learning Representations, 2020. [Online]. Available: https://openreview.net/forum?id=S1lOTC4tDS [23] K. J. Friston, āThe free-energy principle: a unified brain theory?ā Nature Reviews Neuroscience, vol. 11, p. 127ā138, 2010. [Online]. Available: https://api.semanticscholar.org/CorpusID:5053247 [24] K. Friston, T. FitzGerald, F. Rigoli, P. Schwartenbeck, and G. Pezzulo, āActive inference: a process theory,ā Neural computation, vol. 29, no. 1, p. 1ā49, 2017. [25]N. Hansen, H. Su, and X. Wang, āTD-MPC2: Scalable, robust world models for continuous control,ā in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=Oxh5CstDJU [26]D. A. Pomerleau, āEfficient training of artificial neural networks for autonomous navigation,ā Neural Computation, vol. 3, no. 1, p. 88ā97, March 1991. [27]M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi, āA survey of imitation learning: Algorithms, recent developments, and challenges,ā IEEE Transactions on Cybernetics, vol. 54, no. 12, p. 7173ā7186, 2024. [28]J. Ferret, O. Pietquin, and M. Geist, āSelf-imitation advantage learning,ā in Proceedings of the 20th Inter- national Conference on Autonomous Agents and MultiAgent Systems, ser. AAMAS ā21.Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, 2021, p. 501ā509. [29] J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel, āHigh-dimensional continuous control using generalized advantage estimation,ā CoRR, vol. abs/1506.02438, 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:3075448 [30]K. Friston, F. Rigoli, D. Ognibene, C. Mathys, T. FitzGerald, and G. Pezzulo, āActive inference and epistemic value,ā Cognitive neuroscience, 02 2015. [31]K. J. Friston, M. Lin, C. D. Frith, G. Pezzulo, J. A. Hobson, and S. Ondobaka, āActive Inference, Curiosity and Insight,ā Neural Computation, vol. 29, no. 10, p. 2633ā2683, 10 2017. [Online]. Available: https://doi.org/10.1162/neco_a_00999 11 Preprint [32]B. Millidge, A. Tschantz, and C. Buckley, āWhence the expected free energy?ā Neural Computation, vol. 33, p. 1ā36, 01 2021. [33]J. Y. Shin, C. Kim, and H. J. Hwang, āPrior preference learning from experts: Designing a reward with active inference,ā Neurocomputing, vol. 492, p. 508ā515, 2022. [34]N. Sajid, P. Tigas, and K. Friston, āActive inference, preference learning and adaptive behaviour,ā IOP Conference Series: Materials Science and Engineering, vol. 1261, no. 1, p. 012020, oct 2022. [Online]. Available: http://dx.doi.org/10.1088/1757-899X/1261/1/012020 [35]L. Buesing, T. Weber, S. RacaniĆØre, S. M. A. Eslami, D. J. Rezende, D. P. Reichert, F. Viola, F. Besse, K. Gregor, D. Hassabis, and D. Wierstra, āLearning and querying fast generative models for reinforcement learning,ā CoRR, vol. abs/1802.03006, 2018. [Online]. Available: http://arxiv.org/abs/1802.03006 [36]D. Hafner, T. P. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, āLearning latent dynamics for planning from pixels,ā in Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97.PMLR, 2019, p. 2555ā2565. [Online]. Available: http://proceedings.mlr.press/v97/hafner19a.html [37]A. Doerr, C. Daniel, M. Schiegg, N.-T. Duy, S. Schaal, M. Toussaint, and T. Sebastian, āProbabilistic recurrent state-space models,ā in Proceedings of the 35th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, J. Dy and A. Krause, Eds., vol. 80.PMLR, 10ā15 Jul 2018, p. 1280ā1289. [Online]. Available: https://proceedings.mlr.press/v80/doerr18a.html [38] R. E. Kalman, āA new approach to linear filtering and prediction problems,ā Transactions of the ASMEā Journal of Basic Engineering, vol. 82, no. Series D, p. 35ā45, 1960. [39]R. G. Krishnan, U. Shalit, and D. A. Sontag, āDeep kalman filters,ā ArXiv, vol. abs/1511.05121, 2015. [Online]. Available: https://api.semanticscholar.org/CorpusID:15578083 [40]P. Mazzaglia, T. Verbelen, O. Ćatal, and B. Dhoedt, āThe free energy principle for perception and action:A deep learning perspective,ā Entropy, vol. 24, no. 2, 2022. [Online]. Available: https://w.mdpi.com/1099-4300/24/2/301 [41]J. Chung, Ćaglar GülƧehre, K. Cho, and Y. Bengio, āEmpirical evaluation of gated recurrent neural networks on sequence modeling,ā ArXiv, vol. abs/1412.3555, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:5201925 [42]M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley, āStochastic variational inference,ā Journal of Machine Learning Research, vol. 14, no. 40, p. 1303ā1347, 2013. [Online]. Available: http://jmlr.org/papers/v14/hoffman13a.html [43]T. Parr, D. Markovi Ģ c, S. J. Kiebel, and K. J. Friston, āNeuronal message passing using mean-field, bethe, and marginal approximations,ā Scientific Reports, vol. 9, 2019. [44]K. Friston, T. FitzGerald, F. Rigoli, P. Schwartenbeck, G. Pezzulo, et al., āActive inference and learning,ā Neuroscience & Biobehavioral Reviews, vol. 68, p. 862ā879, 2016. [45]R. Smith, K. J. Friston, and C. J. Whyte, āA step-by-step tutorial on active inference and its application to empirical data,ā Journal of Mathematical Psychology, vol. 107, p. 102632, 2022. [Online]. Available: https://w.sciencedirect.com/science/article/pii/S0022249621000973 [46]Y. Kossale, M. Airaj, and A. Darouichi, āMode collapse in generative adversarial networks: An overview,ā in 2022 8th International Conference on Optimization and Applications (ICOA). IEEE, 2022, p. 1ā6. [47]T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, āSoft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,ā in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, StockholmsmƤssan, Stockholm, Sweden, July 10-15, 2018, ser. Proceedings of Machine Learning Research, J. G. Dy and A. Krause, Eds., vol. 80.PMLR, 2018, p. 1856ā1865. [Online]. Available: http://proceedings.mlr.press/v80/haarnoja18b.html [48]T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine, āSoft actor-critic algorithms and applications,ā 2018. [Online]. Available: https://arxiv.org/abs/1812.05905 [49]A. Barto, M. Mirolli, and G. Baldassarre, āNovelty or surprise?ā Frontiers in Psychology, vol. 4, 2013. [Online]. Available: https://w.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2013.00907 [50]P.-Y. Oudeyer and F. Kaplan, āWhat is intrinsic motivation? a typology of computational approaches,ā Frontiers in Neurorobotics, vol. 1, 2007. [Online]. Available: https://w.frontiersin.org/articles/10.3389/ neuro.12.006.2007 12 Preprint [51]M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione, āRegret minimization in games with incomplete information,ā Advances in neural information processing systems, vol. 20, 2007. [52]R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, āPolicy gradient methods for reinforcement learning with function approximation,ā in Advances in Neural Information Processing Systems, S. Solla, T. Leen, and K. Müller, Eds., vol. 12.MIT Press, 1999. [Online]. Available: https://proceedings.neurips.c/paper/1999/file/464d828b85b0bed98e80ade0a5c43b0f-Paper.pdf [53] O. Ćatal, J. Nauta, T. Verbelen, P. Simoens, and B. Dhoedt, āBayesian policy selection using active inference,ā CoRR, vol. abs/1904.08149, 2019. [Online]. Available: http://arxiv.org/abs/1904.08149 [54]B. Millidge, āDeep active inference as variational policy gradients,ā Journal of Mathematical Psychology, vol. 96, p. 102348, 2020. [Online]. Available: https://w.sciencedirect.com/science/article/pii/ S0022249620300298 [55]V. R. Konda and J. N. Tsitsiklis, āOn actor-critic algorithms,ā SIAM J. Control Optim., vol. 42, no. 4, p. 1143ā1166, apr 2003. [Online]. Available: https://doi.org/10.1137/S0363012901385691 [56]P. Auer, N. Cesa-Bianchi, and P. Fischer, āFinite-time analysis of the multiarmed bandit problem,ā Machine Learning, vol. 47, no. 2, p. 235ā256, May 2002. [57] T. Lattimore and C. SzepesvĆ”ri, Bandit Algorithms. Cambridge University Press, 2020. [58]E. L. Deci and R. M. Ryan, Intrinsic Motivation and Self-Determination in Human Behavior.Springer US, 1985. [Online]. Available: http://dx.doi.org/10.1007/978-1-4899-2271-7 [59]P. Schwartenbeck, J. Passecker, T. U. Hauser, T. H. FitzGerald, M. Kronbichler, and K. J. Friston, āComputational mechanisms of curiosity and goal-directed exploration,ā eLife, vol. 8, p. e41703, may 2019. [Online]. Available: https://doi.org/10.7554/eLife.41703 [60]D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, āMastering diverse control tasks through world models,ā Nature, vol. 640, no. 8059, p. 647ā653, Apr 2025. [Online]. Available: https://doi.org/10.1038/s41586-025-08744-2 [61] D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba, āMastering atari with discrete world models,ā in International Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/forum?id=0oabwyZbOu [62]T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, āContinuous control with deep reinforcement learning,ā 2015. [Online]. Available: https://arxiv.org/abs/1509.02971 [63]A. Ororbia and A. Mali, āActive predictive coding: Brain-inspired reinforcement learning for sparse reward robotic control problems,ā in 2023 IEEE International Conference on Robotics and Automation (ICRA), 2023, p. 3015ā3021. [64]J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, āProximal policy optimization algorithms,ā CoRR, vol. abs/1707.06347, 2017. [Online]. Available: http://arxiv.org/abs/1707.06347 [65]A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman, āRelay policy learning: Solving long-horizon tasks via imitation and reinforcement learning,ā in Proceedings of the Conference on Robot Learning, ser. Proceedings of Machine Learning Research, L. P. Kaelbling, D. Kragic, and K. Sugiura, Eds., vol. 100.PMLR, 30 Octā01 Nov 2020, p. 1025ā1037. [Online]. Available: https://proceedings.mlr.press/v100/gupta20a.html [66]T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, āMeta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,ā in Proceedings of the Conference on Robot Learning, ser. Proceedings of Machine Learning Research, L. P. Kaelbling, D. Kragic, and K. Sugiura, Eds., vol. 100.PMLR, 30 Octā01 Nov 2020, p. 1094ā1100. [Online]. Available: https://proceedings.mlr.press/v100/yu20a.html [67] Y. Zhu, J. Wong, A. Mandlekar, R. MartĆn-MartĆn, A. Joshi, S. Nasiriany, Y. Zhu, and K. Lin, ārobosuite: A modular simulation framework and benchmark for robot learning,ā in arXiv preprint arXiv:2009.12293, 2020. [68]7Bot, ā7bot:A powerful desktop robot arm for future inventors,ā Kickstarter, 2015, ac- cessed:March 3, 2026. [Online]. Available:https://w.kickstarter.com/projects/1128055363/ 7bot-a-powerful-desktop-robot-arm-for-future-inven [69]Trossen Robotics, āPincherx 100 robot arm,ā Trossen Robotics Product Page, 2018, accessed: March 3, 2026. [Online]. Available: https://w.trossenrobotics.com/pincherx100 13