Paper deep dive
Human-in-the-Loop Pareto Optimization: Trade-off Characterization for Assist-as-Needed Training and Performance Evaluation
Harun Tolasa, Volkan Patoglu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/26/2026, 2:11:35 AM
Summary
The paper introduces a Human-in-the-Loop (HiL) Pareto optimization framework using Bayesian multi-criteria optimization to characterize the trade-off between task performance (quantitative) and perceived challenge level (qualitative) in motor skill training. This approach enables the design of Assist-as-Needed (AAN) protocols and provides a rigorous method for evaluating individual and group-level training progress and performance comparisons.
Entities (6)
Relation Signals (3)
Harun Tolasa → authored → Human-in-the-Loop Pareto Optimization
confidence 100% · Harun Tolasa and V. Patoglu are with the Faculty of Engineering and Natural Sciences at Sabanci University
HiL Pareto Optimization → utilizes → Bayesian multi-criteria optimization
confidence 95% · We adapt Bayesian multi-criteria optimization to systematically and efficiently perform HiL Pareto characterizations.
USeMO → implements → HiL Pareto Optimization
confidence 90% · We utilize a customized version of the wrapper method, called Uncertainty-aware Search framework for optimizing Multiple Objectives (USeMO)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:During human motor skill training and physical rehabilitation, there is an inherent trade-off between task difficulty and user performance. Characterizing this trade-off is crucial for evaluating user performance, designing assist-as-needed (AAN) protocols, and assessing the efficacy of training protocols. In this study, we propose a novel human-in-the-loop (HiL) Pareto optimization approach to characterize the trade-off between task performance and the perceived challenge level of motor learning or rehabilitation tasks. We adapt Bayesian multi-criteria optimization to systematically and efficiently perform HiL Pareto characterizations. Our HiL optimization employs a hybrid model that measures performance with a quantitative metric, while the perceived challenge level is captured with a qualitative metric. We demonstrate the feasibility of the proposed HiL Pareto characterization through a user study. Furthermore, we present the utility of the framework through three use cases in the context of a manual skill training task with haptic feedback. First, we demonstrate how the characterized trade-off can be used to design a sample AAN training protocol for a motor learning task and to evaluate the group-level efficacy of the proposed AAN protocol relative to a baseline adaptive assistance protocol. Second, we demonstrate that individual-level comparisons of the trade-offs characterized before and after the training session enable fair evaluation of training progress under different assistance levels. This evaluation method is more general than standard performance evaluations, as it can provide insights even when users cannot perform the task without assistance. Third, we show that the characterized trade-offs also enable fair performance comparisons among different users, as they capture the best possible performance of each user under all feasible assistance levels.
Tags
Links
- Source: https://arxiv.org/abs/2603.23777v1
- Canonical: https://arxiv.org/abs/2603.23777v1
Trouble viewing inline? Open PDF directly →
Full Text
81,396 characters extracted from source content.
Expand or collapse full text
1 Human-in-the-Loop Pareto Optimization: Trade-off Characterization for Assist-as-Needed Training and Performance Evaluation Harun Tolasa, Student Member, IEEEVolkan Patoglu, Member, IEEE Abstract—During human motor skill training and physical re- habilitation, there is an inherent trade-off between task difficulty and user performance. Characterizing this trade-off is crucial for evaluating user performance, designing assist-as-needed (AAN) protocols, and assessing the efficacy of training protocols. In this study, we propose a novel human-in-the-loop (HiL) Pareto optimization approach to characterize the trade-off between task performance and the perceived challenge level of motor learning or rehabilitation tasks. We adapt Bayesian multi-criteria optimization to systematically and efficiently perform HiL Pareto characterizations. Our HiL optimization employs a hybrid model that measures performance with a quantitative metric, while the perceived challenge level is captured with a qualitative metric derived from preference-based user feedback. We demonstrate the feasibility of the proposed HiL Pareto characterization through a user study. Furthermore, we present the utility of the framework through three use cases in the context of a manual skill training task administered to healthy individuals with haptic feedback. First, we demonstrate how the characterized trade- off can be used to design a sample AAN training protocol for a motor learning task and to evaluate the group-level efficacy of the proposed AAN protocol relative to a baseline adaptive assistance protocol. Second, we demonstrate that individual-level comparisons of the trade-offs characterized before and after the training session enable fair evaluation of training progress under different assistance levels. This evaluation method is more general than standard performance evaluations, as it can provide insights even when users cannot perform the task without assistance. Third, we show that the characterized trade-offs also enable fair performance comparisons among different users, as they capture the best possible performance of each user under all feasible assistance levels. Index Terms—Pareto optimization, Bayesian multi-criteria op- timization, human-in-the-loop optimization, qualitative perfor- mance, assist-as-needed paradigms, interaction control, force- feedback devices, motor learning, robot-assisted rehabilitation. I. INTRODUCTION Physical human-robot interaction (pHRI) is commonly used for training tasks, such as human motor skill training or robot-assisted physical rehabilitation. In such applications, the efficacy of the training, typically measured by performance evaluations administered after the training without any assis- tance, is of utmost importance. Haptic assistance is provided only during the training sessions, when the user is coupled to a force-feedback robot, to ensure that the execution of the H. Tolasa and V. Patoglu are with the Faculty of Engineering and Natural Sciences at Sabanci University, Istanbul, Turkiye. harun.tolasa, volkan.patoglu@sabanciuniv.edu. This work has been par- tially supported by TUBITAK Grants 120N523 and 23AG003. task can be completed with sufficient success. For instance, consider a stroke patient going through physical rehabilitation, for whom task execution may not be feasible without a proper level of assistance. For such a patient, some level of assistance is necessary for task completion. On the other hand, too much assistance during training is known to be detrimental, as users may learn to rely on the existence of this support and slack [1], [2]. Too much assistance may also cause the task to be per- ceived as not sufficiently challenging by the patient, resulting in a lack of engagement. As a consequence of excessive assistance, users may not learn how to successfully execute the task when no assistance is available. Accordingly, there exists an inherent trade-off between the level of assistance provided and the level of challenge perceived by the user for any given motor control/rehabilitation task. This trade-off has been widely acknowledged in the litera- ture, and it has been established that an optimal assistance level consists of the least assistance that allows a user to execute a task with a sufficient level of success. Controllers that aim to provide such assistance are commonly referred to as assist-as- needed (AAN) controllers. AAN controllers aim to keep the assistance level at a proper level to maximally challenge but not to overwhelm users with task difficulty or demotivate them with continual failures. The literature on AAN control mostly focuses on the design of interaction controllers that can administer assistance forces safely and naturally, without overriding users’ intent. In par- ticular, path-following controllers, such as velocity field con- trollers [3] and controllers with guaranteed coupled stability properties, such as passive velocity field controllers [4], have been proposed and adopted in many related works [5]–[9]. Although these studies on interaction control are indispensable for the safe and natural delivery of force feedback, these control approaches necessitate the proper assistance level to be provided as an input, typically by a domain expert. In AAN controllers, the proper level of assistance is com- monly decided empirically or heuristically, based on thresholds imposed on the measured/estimated performance of the user. Most methods utilize sensor inputs or biosignals [10]–[14], such as EMG and EEG, to adjust the level of assistance to promote voluntary participation based on some pre-determined thresholds. Adaptive controllers that utilize the dynamic model of the user and the device to minimize a cost function [15] or rely on statistical estimates of psychophysical thresholds [16] have also been proposed. Reinforcement learning algorithms have been used to develop AAN controllers that do not depend arXiv:2603.23777v1 [cs.RO] 24 Mar 2026 2 Fig. 1. Proposed HiL Pareto characterization scheme and its application to AAN training, within- and between-participant performance evaluations. on user- or device-specific parameters [17]. Furthermore, in addition to quantitative measures, more qualitative aspects, such as the psychological states of the user, have been es- timated by neural networks to adjust assistance levels [18], [19]. Interested readers are referred to the review articles [20], [21] for recent implementations of AAN control techniques for lower- and upper-extremity rehabilitation. While these studies are quite promising, the characterization of the underlying trade-off between the performance versus the perceived task difficulty and the determination of a customized level of assistance to be provided to a user are still open chal- lenges, commonly delegated to a domain expert. The problem is challenging as each individual is unique; furthermore, user preferences and perceptions undergo continuous changes as learning/recovery takes place with training. For instance, the perceived challenge level of a task under a certain level of assistance is likely to decrease as a user gets better at the task. User-dependent metrics that are not easy to quantify, such as motivation and comfort level, may directly affect the perceived challenge level. Hence, determining the ideal assistance level necessitates personalization, possibly through a characterization of the underlying trade-off between the perceived challenge level and the performance for each user. Characterization of the trade-off between the perceived challenge level and the task performance can guide training by helping designers to establish a proper level of assistance for AAN controllers. Moreover, such characterizations can be performed at different stages of the training, such that changes in user preferences/performance can be captured as learning takes place and the assistance can be adjusted accordingly. Furthermore, such trade-off characterizations can serve as a novel means for evaluating performance at a group- or individual-level. While typical evaluations of manual skill training or physical rehabilitation are performed when no assistance is provided, as this represents the real-life use case, such evaluations cannot capture improvements in performance during the early phases of training/rehabilitation. For instance, if a task is excessively hard for a patient, then it may not be possible for the patient to execute it successfully without a proper level of assistance. In such cases, evaluations with no assistance will not capture any improvement with training, as the patient’s success rate will remain low. On the other hand, a comparison of the characterized trade-offs between the perceived challenge level versus the performance of a user, or a group of users, at different stages of training provides a feasible alternative: the time evaluation of the trade-off provides a rigorous means to fairly assess the progress, under all feasible assistance levels. In addition, such trade-off characterizations also enable comparisons among various users or groups of users. It is important to emphasize that ensuring the rigor and fairness of such performance comparisons under assistance necessitates the characterization of the trade-off via Pareto optimization, since multiple variables, such as assistance and the perceived challenge levels, need to be considered simultaneously. In this study, we propose a human-in-the-loop (HiL) Pareto optimization approach to characterize the trade-off (i.e., Pareto optimal results forming a set of non-dominated solutions) between the user’s performance and the perceived challenge level for a motor learning task. During the optimization, the user performance is measured by a quantitative metric, while the perceived challenge level is captured by a qualitative metric gathered through preference-based qualitative feedback, as depicted in Figure 1. A multi-criteria Bayesian optimization technique is utilized for HiL Pareto optimization of this hybrid model with quantitative and qualitative metrics. The sample- 3 efficiency of the underlying Bayesian optimization is a crucial aspect, as it enables the trade-off to be characterized system- atically and efficiently by focusing on the relevant portion of the search space and without inducing fatigue to participants. Once the trade-off is characterized via HiL Pareto opti- mization, first, we demonstrate how this trade-off can be used to guide the training sessions with Pareto optimal assistance levels. In particular, we show how a subset of optimal solutions can be selected from the set of non-dominated solutions to guide AAN training sessions. Second, we show that the trade- off evolves in time as learning takes place, and Pareto-front curves capturing the trade-off can be used to fairly evaluate both the individual- and group-level progress. Third, we show that the characterized trade-offs can also be used for fair performance comparisons among different users, by capturing the best possible performance of each user under all feasible assistance levels. Contributions The main contribution of this study is a novel HiL Pareto op- timization framework with hybrid (quantitative and qualitative) performance measures, as depicted in the first row of Figure 1. We not only formulate the HiL trade-off characterization framework but also show its feasibility through a user study by demonstrating that the proposed HiL Pareto optimization can be used to systematically and efficiently characterize the trade-off between the performance and the perceived challenge level of a motor skill learning task. Our second contribution is the demonstration of the useful- ness of the HiL trade-off characterization through three use cases, as depicted in the second row of Figure 1. Within the context of a case study involving a manual skill training task administered to healthy individuals with haptic feedback, we demonstrate that - the non-dominated solutions characterizing the trade-off between the performance and the perceived challenge level can be used to design AAN controllers, - the comparisons of the trade-off curves characterized at different stages of training can be used to establish a novel and rigorous means of individual- and group-level performance evaluation across assistance levels, and - the trade-off curves of different users can be used for fair comparisons among various users, as they capture the best possible performance of each user under all feasible assistance levels. I. RELATED WORK Bayesian optimization is a sample-efficient derivative-free global optimization approach [22] that has been employed in several studies involving HiL optimization. For instance, HiL applications of these algorithms have been used to optimize wearable robotic assistive devices, where the evaluation of optimization metrics is costly, or the number of experiments is constrained by human involvement [23]–[26]. HiL Bayesian optimization studies focus on the optimization of a single-criterion cost function. These techniques are mostly based on quantitative metrics, such as metabolic cost [23]– [26]. Recently, qualitative metrics, such as perceived realism, have been captured by probabilistic latent functions and used to implement HiL Bayesian optimizations [27]–[31]. Bayesian optimization approaches have been extended to solve multi-criteria optimization problems [32]–[37]. One means to address a multi-criteria optimization problem is to use a weighted sum of the cost functions to form a single aggregate cost function. Such scalarization approaches enable the original multi-objective problem to be formulated as a single-criterion optimization problem. Scalarization ap- proaches have also been applied to the Bayesian optimization setting, either by utilizing an aggregate cost function [32] or a scalarized acquisition function for parameter sampling [33], [34]. However, in all scalarization methods, since the relative weights of cost functions need to be pre-determined, the design preferences among the objectives must be assigned a priori, before gaining sufficient knowledge of the trade-off involved. On the other hand, for multi-criteria optimization prob- lems, Pareto optimization methods compute a set of non- dominated solutions that correspond to optimal designs for different design preferences among the optimization metrics, and all such non-dominated solutions constitute the Pareto- front [38]–[40]. Unlike the single-shot scalarization-based optimization methods, Pareto methods fully characterize the trade-off among multiple objectives. Once the Pareto-front solutions are computed, the designer can study these solutions to get an insight into the underlying trade-offs and make an informed decision to finalize the design by selecting a subset of optimal solutions from the Pareto set. Pareto-optimization approaches have also been applied in the Bayesian optimization setting [35]–[37]. Existing tech- niques can be loosely classified as hyper-volume improve- ment approaches [36], information-theoretic methods [37], and wrapper methods via single-objective acquisition func- tions [35]. In hyper-volume improvement approaches, the ex- pected improvement acquisition function is extended for multi- objectives to produce an effective sampling strategy [36], while the information-theoretic methods derive a single acquisition function to maximize information gain for all objectives [37]. Information-theoretic methods have also been extended to work with multiple qualitative metrics [41]. Finally, the wrap- per methods utilize a multi-criteria optimizer to compute a surrogate Pareto-front characterizing the trade-off between the acquisition functions and sample parameters with the highest volumetric variance [35]. The previous studies of the authors focus on single-criterion HiL Bayesian optimization with qualitative metrics to improve perceived realism of haptic rendering [30], [42] and to explore visual-haptic cue integration during multi-modal haptic ren- dering under conflicting cues [31]. This study extends these works to HiL Pareto optimization for the characterization of the underlying trade-off between users’ performance and perceived challenge level of motor learning tasks. The under- lying concepts, mathematical formulations, and optimization approaches used to solve Pareto optimization problems are distinct from those of single-criterion optimization [38]; hence, the extension from single-criterion HiL optimization to HiL Pareto optimization is substantial. 4 Application of Pareto optimization approaches in a HiL setting has only recently been pursued. In addition to this study, a multi-objective HiL optimization problem has recently been addressed in [43] as a case study, utilizing a genetic algorithm-based solver to adjust the assistance profiles of an ankle exoskeleton by simultaneously minimizing two quanti- tative metrics capturing gait deviations. Our study considers hybrid (both quantitative and qualitative) metrics, relies on a sample-efficient Bayesian Pareto optimization approach, and applies the framework in a human motor skill learning setting. Specifically, our study extends the Bayesian Pareto optimization approach in [35] to hybrid models and applies it in the HiL context for a human motor skill learning task. Our study not only applies HiL Pareto optimization in a motor skill learning setting with hybrid metrics, but also demonstrates its usefulness for the evaluation of individual- and group-level performance and for the design of AAN controllers. Finally, the comparison of Pareto solutions to enable fair comparisons of performance among various designs is new to the motor skill learning context, while the idea has been successfully applied to comparisons of exoskeleton perfor- mance of musculoskeletal simulations [44] and comparisons of interaction controller performance in the context of pHRI [45]. I. HUMAN-IN-THE-LOOP PARETO OPTIMIZATION The goal of HiL Pareto optimization is to characterize the inherent trade-off between task difficulty and performance. In general, these two metrics are in direct conflict with each other; hence, a multi-objective optimization needs to be performed to consider them simultaneously. While it is easier to measure task performance directly through quantitative metrics, the task difficulty is much harder to capture, as it depends on many aspects, including but not limited to the sensorimotor skills, as well as engagement, (physical/mental) fatigue, and comfort level of the user. Along these lines, we measure the task performance (GP num ) quan- titatively based on users’ scores, as detailed in Section I-B, while we capture the perceived challenge level (GP qual ) of the users by directly acquiring their qualitative feedback, by querying for their ordinal classifications and pair-wise prefer- ences, as detailed in Section I-C. Given these two metrics, the following max-max optimization problem is solved over the design variable of assistance percentage, subject to the physical/mental constraints imposed by the user: max assistance %∈[0,100] GP num : task performance GP qual :perceived task difficulty subject to: physical and mental constraints of the user While performing HiL Pareto optimization, it is important to determine the trade-off among multiple conflicting objective functions, while also minimizing the total resource cost of the experiments. Among various approaches, multi-criteria Bayesian optimization techniques hold promise for use in HiL applications due to their inherent sample efficiency [46], as the central idea of all Bayesian optimization approaches is to minimize the number of observations while rapidly converging to the optimal solution(s). Accordingly, Bayesian optimiza- tion provides a class of sample-efficient global optimization methods, where a probabilistic model conditioned on previous observations is used to determine the future evaluations. The statistical modeling in multi-criteria Bayesian opti- mization techniques is typically handled by one Gaussian process (GP) model for each objective to ensure the tractability of the problem, while the selection of the acquisition function to capture the trade-off between multiple objectives results in various alternative techniques [32]–[34]. We utilize a customized version of the wrapper method, called Uncertainty-aware Search framework for optimizing Multiple Objectives (USeMO), which is based on single- objective acquisition functions [35]. USeMO utilizes a multi- criteria sampling strategy that allows one to leverage acqui- sition functions from single-objective Bayesian optimization to solve the multi-objective Bayesian optimization problem, as detailed in Section I-D. We have preferred USeMO as its application to optimizations with hybrid (qualitative and quantitative) metrics is more accessible. Please note that USeMO is a sample optimization approach appropriate for HiL Pareto optimization, and alternative methods, such as [36], [37], may also be adapted for HiL Pareto optimization. A. An Overview of HiL Multi-Criteria Bayesian Optimization The multi-criteria Bayesian optimization approach, summa- rized in Algorithm 1, relies on two GP regression models: one based on quantitative performance scores as detailed in Section I-B, and another one based on qualitative feedback collected from the participants as detailed in Section I-C. Both GP regression models possess their corresponding ac- quisition functions. During the initial iterations, parameters are selected via the Sobol sequence to ensure search-space exploration. For the rest of the iterations, new solutions are computed via the sampling strategy proposed in [35] and further detailed in Section I-D. Algorithm 1 HiL Bayesian Pareto Optimization Initiate A: Feasible parameter space, GP num : Quantitative GP re- gression model, GP qual : Qualitative GP regression model, N : # of iterations, N 0 : # of space-filling iterations 1: Assign acquisition function α num to GP num 2: Assign acquisition function α qual to GP qual 3: Create a sampling set x with size N 0 using Sobol sequence 4: for n = 1 to N do 5:if n≤ N 0 then 6:Select x n from set x 7:else 8:Compute surrogate Pareto set: 9:x p ← arg max x∈A (α num ,α qual ) 10:Select parameter: 11:x n ← arg max x∈x p (σ qual · σ num ) 12:Append x n to set x 13:Use x n in the task trial 14:Get the performance score s n 15:Append s n to s and re-calculate y 16:Get user’s qualitative feedback q n 17:Append q n to set q 18:Update GP num using x and standardized score set y 19:Update GP qual using x and qualitative feedback set q 20: Compute Pareto front using the mean predictions of the GP models 21: Plot the Pareto front and list the non-dominated solutions 5 In particular, the sampling strategy utilizes a multi-criteria optimizer to compute a surrogate Pareto front characterizing the trade-off between the conflicting acquisition functions. Then, from the set of non-dominated solutions on the surrogate Pareto front, a parameter with the highest volumetric variance is selected for the next sampling. At the end of the trade-off characterization session, Algorithm 1 utilizes a multi-criteria optimizer to compute the non-dominated solutions forming the Pareto front of the expected performance scores and expected perceived challenge level. Once the Pareto set is computed, it can be used for evaluations, or promising non-dominated solutions can be selected as in Section I-F to design AAN training protocols. B. Quantitative Gaussian Process Regression Model A GP regression model, GP num , is designed to learn the relationship between the level of assistance and the partici- pant’s performance. The maximum applicable assistance level during a task is represented by one, while zero represents the case without assistance. Let A = x ⊂R d : 0 ≤ x i ≤ 1 be the feasible parameter space and x i be the assistance level. Let x = x 1 ,x 2 ,..,x n be a set consisting of n assistance levels used in previous task trials. Then, let s = s 1 ,s 2 ,..,s n be a set consisting of n observed performance scores corresponding to x, and let y = y 1 ,y 2 ,..,y n be the statistically standardized version of s. The dataset used to train the numerical GP regression model is represented with D N =(x 1 ,y 1 ), (x 2 ,y 2 ),..., (x n ,y n ). Finally, let f N (x) be the black-box function representing the relationship between the assistance level and the standardized performance scores. To model the deviation of the score measurements, we assume that the standardized performance scores are affected by a Gaussian white noise with variance σ 2 w , and observe their noisy version y. Then, the prior black- box function f N for the performance scores can be modeled as f N (x)∼GP(0,K N + σ 2 w I)(1) with K N ∈R n×n ,k N i,j = k N (x i ,x j ), is a kernel function used for the correlation. The performance scores for an unknown assistance level can be predicted via Bayesian inference. Let x ∗ be any arbitrary assistance level, and let f N ∗ |D N be the corresponding performance score estimation based on previously acquired performance scores. Then, the Bayesian inference indicates f N ∗ |D N ∼ GP(μ N ∗ |D N ,σ 2 N ∗ |D N )(2) μ N ∗ |D N = k N ∗,1:n (K N + σ 2 w I) −1 y(3) σ 2 N ∗ |D N = k N ∗,∗ − k N ∗,1:n (K N + σ 2 w I) −1 k N 1:n,∗ (4) C. Qualitative Gaussian Process Regression Model A second GP regression model, GP qual , is designed to learn the relationship between the assistance level and the perceived challenge of the task as perceived by the participant. During the Pareto characterization session, after each task trial, the participant classifies the perceived challenge level of the task trial by selecting one of the following categories: easy, moderate, and hard. Then, excluding initialization, the participant compares the perceived challenge level of the task trial with that of the previous iteration. Based on the modeled probabilities of the responses, the qualitative GP regression model is updated. The range of assistance level is represented by A = x ⊂ R d : 0≤ x i ≤ 1, as in the quantitative GP regression model. Let f Q (x) be the latent function representing the participant’s perceived challenge level of the task. The prior distribution of f Q (x) is modeled with a normal GP distribution. f Q (x)∼GP(0,K Q )(5) where K Q ∈R n×n ,k Q i,j = k Q (x i ,x j ) is the noiseless kernel matrix of the qualitative GP regression model. It is worth noting that the kernel functions used in both models need not be identical. Let q be the set of qualitative feedback provided by the participant, where q consists of both ordinal classifications q o = q o 1 ,q o 1 ,...,q o n and pairwise preferences q p = q p 2 ,q p 3 ,...,q p n . The dataset including ordinal classifica- tions is defined as D O = (x 1 ,q o1 ), (x 2 ,q o2 ),..., (x n ,q on ), and the dataset including pairwise comparisons is defined as D P = (x 1 ,x 2 ,q p 1 ), (x 2 ,x 3 ,q p 2 ),..., (x n−1 ,x n ,q p n ). Then, the dataset consisting of all qualitative feedback can be defined as D Q = D O ∪ D P . The probability of the latent function based on the provided feedback is calculated from the proportionality P (f Q |D Q )∝ P (D Q |f Q )P (f Q ), where P (f Q ) is the prior unbiased prob- ability of the regression model and P (D Q |f Q ) is the proba- bility of the qualitative feedback of the participant is correct based on the given latent function. Ordinal classifications and pairwise comparisons are mod- eled as in [28]–[31], [47], as follows: Let O = o 1 ,o 2 ,o 3 be the finite set of three ordinal classifications representing the perceived challenge level from “easy” to “hard”. Let t be the set of ordered thresholds used to distinguish ordinal classifications t = t o 0 ,t o 1 ,t o 2 ,t o 3 and −∞ = t o 0 < t o 1 < t o 2 < t o 3 = ∞. Then, the probability for the participant to correctly classify a parameter x i with the ordinal class o th j is modeled as P (q o i =o j |f Q ) = Φ t o j − f Q (x i )) c o − Φ t oj−1 − f Q (x i ) c o (6) where Φ represents the cumulative distribution function of the Gaussian distribution and c o > 0 is used to capture the classification noise. The probability of the participant correctly identifying the harder one between two task trials is given by P (q p i = (x i ≻ x i−1 )|f Q ) = Φ (f Q (x i ) )− f Q (x i−1 ) ) c p (7) where c p > 0 captures the noise in pair-wise preferences. Under the assumption of the independence of provided qualitative feedback, P (D Q |f Q ) = P (D O |f Q )P (D P |f Q ) is calculated as P (D Q |f Q ) = n Y i=1 P (q o i |f Q ) n Y i=2 P (q p i |f Q ).(8) 6 The posterior distribution of the regression model is com- puted by the Laplace approximation [48]. Utilizing this com- monly adopted method for posterior approximation, it becomes possible to make predictions for Bayesian optimization and extend the probabilistic derivations to capture qualitative feed- back from humans [28], [29], [31], [49]. The perceived challenge level estimate for the participant f Q ∗ |D Q for any arbitrary assistance level x ∗ can be computed using the posterior model as follows: f Q ∗ |D Q ∼ GP (μ ∗|D Q ,σ 2 ∗|D Q )(9) μ ∗|D Q = k Q ∗,1:n K −1 Q ˆ f Q (10) σ 2 ∗|D Q = k Q ∗ − k Q ∗,1:n (W −1 + K Q ) −1 k Q 1:n,∗ (11) where K is the noiseless covariance matrix calculated with RBF kernel, W is the negative Hessian matrix W ij =− ∂ 2 log(P (D Q |f (x)) ∂f (x i )∂f (x j ) and ˆ f is the latent function that maximizes the log-likelihood ˆ f = argmax f(x) (log(P (D Q |f Q )P (f Q ))). D. Sampling Strategy for HiL Pareto Characterization Let α num be the acquisition function of GP num and let α qual be the acquisition function of GP qual . Given these acquisition functions, first, a computationally cheap Pareto optimization problem is solved to find the parameter set x p lying on the surrogate Pareto front. Next, a parameter is selected among x p based on the maximum volumetric uncertainty, where the volumetric uncertainty is calculated by multiplying the standard deviations as follows: x p = argmax x∈A (α num ,α qual )(12) x n = argmax x∈x p (σ (x)|D N n−1 × σ (x)|D Q n−1 ). (13) E. Acquisition Function and Hyper-Parameter Selection Acquisition functions and hyperparameters of the HiL op- timization utilized in this study are empirically determined, based on expertise gained through pilot experiments. Ten trials were observed to be sufficient for HiL Pareto characterizations. In particular, a radial basis function (RBF) kernel, k i,j =exp(−θ||x i − x j || 2 2 ), is used for modeling both GP num and GP qual . The kernel hyperparameters θ N for GP num and θ Q for GP qual are both selected as 5. The observation noise hyper-parameter σ 2 w of GP num is selected as 0.1. The hyperparameters c p and c o , capturing the pairwise preference and ordinal classification noise, are selected as 0.5 and 1, respectively. The ordinal classification thresholds are selected as t =−∞,−0.5, 0.5,∞. For the parameter sampling strategy, two upper confidence bound (UCB) acquisition functions are used. To sample an assistance level x n for n th trial, α num and α qual are used as α num (x ∗ ) n = μ ∗|D N n−1 + λ N σ ∗|D N n−1 (14) α qual (x ∗ ) n = μ ∗|D Q n−1 + λ Q σ ∗|D Q n−1 (15) where λ N and λ Q are the hyper-parameters of GP num and GP qual , respectively, and x ∗ is an any arbitrary feasible assistance level. In these equations, λ N is selected as 2, while λ Q is selected as 1. Once the human-subject experiments were completed, we compared the empirically selected GP hyperparameters with the MLE-optimized values determined from the data and ver- ified that the Pareto-optimal assistance levels obtained under both parameter sets remain close to each other. F. Sample Design Selection for Pareto-based AAN Training All Pareto solutions characterizing the trade-off between the user performance and the perceived challenge are optimal. Without additional preference information, all Pareto opti- mal solutions are mathematically equal. The Pareto selection process imposes additional preferences on the set of Pareto optimal solutions to select a subset among them. Pareto optimization methods allow the designer to impose additional constraints after the computation of the non-dominated so- lutions and inspection of the trade-off involved among the conflicting objectives. Unlike scalarization approaches, where the preferences need to be assigned a priori, the ability to impose additional constraints on the set of optimal solutions corresponding to different preferences in a multi-shot manner is among the most beneficial aspects of Pareto approaches. The trade-off between performance and task difficulty char- acterized via HiL Pareto optimization can provide insights to guide the design of AAN algorithms. To demonstrate how such protocols can be designed by the characterized trade-off, we propose and test the efficacy of a sample AAN training protocol. By imposing new constraints on the Pareto-front based on domain expertise, we form the AAN training protocol as shown in Figure 2, so that the controller keeps the task moderately challenging while allowing the user to achieve an acceptable level of performance. Perce � ved Challenge Level Easy Task Performance Moderate Hard Selected Non-Dominated Solutions 40% of max score 80% of max score Fig. 2.Selection of non-dominated solutions from the Pareto-front by introducing design constraints after characterizing the trade-off In particular, to facilitate a design selection among all non- dominated solutions, we introduce personalized thresholds to the set of non-dominated solutions to keep the perceived challenge level and quantitative performance sufficiently high, as decided by the domain expert. In the sample protocol, the performance range is limited to 40–80% of the capacity of the individual, while the challenge range is limited to 40– 80% of the perceived challenge level of the individual. All 7 solutions in this range have been selected to be included in the training session, to induce diversity in the training exercises. The proposed AAN training method emphasizes the importance of customization of the assistance based on individual performance characterizations. It is important to note that the resulting AAN training protocol based on HiL Pareto characterization is merely a sample design selected by the domain expert that utilizes a set of non-dominated solutions for diversity-focused personalized AAN training. The Pareto selection process is not unique, and alternative training methods may be devised from the same set of Pareto solutions. While each of these protocols is equally valid from a multi-criteria design perspective, their training efficacy is likely to vary significantly. To examine how the selected threshold affects the set of assistance levels used during AAN training, we evaluated alternative thresholds on the Pareto solutions after the human- subject experiments were completed. The analysis reveals that threshold ranges of 30–70%, 40–80% and 50–90% yield mean assistance levels of 30.6± 22.3%, 39.3± 22.6%, and 52.3±17.8%, respectively. As expected, shifting the thresholds to cover higher performance yields increased assistance levels. Given different preferences yield different AAN strategies with varying training outcomes, more sophisticated designs, such as an adaptive protocol for selecting the appropriate threshold ranges to be imposed on Pareto solutions to adjust the assistance levels based on instantaneous user performance on the fly, may perform better than the sample AAN protocol in terms of training efficacy. The sample AAN protocol is preferred due to its simplicity, as the main goal of this study is to illustrate the feasibility of such designs based on HiL Pareto characterizations. IV. HUMAN SUBJECT EXPERIMENT This section presents the application of the proposed HiL Pareto optimization approach to a manual skill training task with haptic feedback, administered to healthy volunteers. It demonstrates that HiL Pareto optimization can be used to (i) efficiently and systematically characterize the trade-off between the task performance and the perceived challenge level, (i) design AAN training protocols, and (i) establish a novel and rigorous means of performance evaluation under assistance, within and between participants and groups. A. Participants 34 participants, with an average age of 23.6 ± 3.6 years participated in the experiment. The participants performed a motor learning task using a force-feedback joystick with their dominant hand. All participants signed the informed consent form approved by the Institutional Review Board of Sabanci University (Protocol No: FENS-2025-18). Participants with any known sensory-motor disability or significant prior experience with haptic devices are excluded from the study. Similarly, ceiling pre-training performance of 80% under less than 40% assistance was used as an exclusion criterion to avoid saturation during training. Ac- cordingly, participants completed warm-up and pre-evaluation sessions to verify that the task is sufficiently challenging; two participants who achieved ceiling-level performance were excluded from the study before the group assignments. Eligible participants were assigned to experimental groups using a Latin square method to achieve balanced group allocations. Participants were blinded to group identity and protocol dif- ferences throughout the experiment. All assigned participants completed the full protocol without technical failures or with- drawal, resulting in no attrition or missing data. B. Task and Apparatus The motor learning task consists of a two-dimensional balancing of a virtual inverted pendulum on a cart, displayed on an LCD screen, as shown in Figure 3(a). A single-axis force-feedback joystick, called HandsOn-SEA, presented in Figure 3(b), renders the dynamics of the cart and pendulum model, so that users feel as if they are moving the cart while operating the joystick. Assistance forces are also provided through this interface. Two monsters, separated by a constant distance, provide continuous disturbances. The external distur- bances are generated using a stochastic process with identical parameters across participants and trials. The task fails if the pendulum rotates more than ±50 ◦ or the cart touches the monsters. The objective of the game is to survive for 25 s, with raw scores corresponding to the survival time and normalized scores divided by 25 s. Pendulum Balancing TaskHandsOn-SEA Haptic Interface (a)(b) Fig. 3. (a) A participant holding a force-feedback joystick (HandsOn-SEA) while interacting with the pendulum game, (b) HandsOn-SEA haptic interface The balancing task is intentionally designed to be challeng- ing to avoid performance saturation; however, due to the high sensitivity of the task dynamics, brief lapses of attention can easily lead to failures. Accordingly, participants are provided with three chances per trial, and a best-of-three score is used as a more robust and optimistic performance metric to reduce noise and avoid under-evaluating participant performance due to such momentary distractions. Interaction and assistance forces are rendered through HandsOn-SEA, a custom-built single-axis haptic interface with series elastic actuation (SEA) [50]. HandsOn-SEA features a coreless DC motor equipped with an encoder to actuate its sector pulley through a capstan-drive to impose desired mo- tions. A cross-flexure pivot, formed by crossing two symmetric leaf springs, acts as the compliant element located between the rigid sector pulley and the handle structures. A Hall- effect sensor constrained to move between the neodymium 8 Warm-Up 5-10 �terat�ons ~10 m�n Pre- Evaluat�on 3 �terat�ons ~2 m�n Pareto Character�zat�on I 10 �terat�ons ~7 m�n Tra�n�ng 20 �terat�ons ~18 m�n Post- Evaluat�on 3 �terat�ons ~2 m�n Pareto Character�zat�on I 10 �terat�ons ~7 m�n Fig. 4. Experimental procedure block magnets embedded in the sector pulley measures the deflections of the compliant element, thereby enabling esti- mation of the interaction force. HandsOn-SEA can provide 15 N continuous force output at its handle with a force control bandwidth of 12 Hz, within a workspace of ±55 ◦ . HandsOn- SEA works under velocity-sourced impedance control [51]– [54], implemented in real-time at 1 kHz, utilizing a TI C2000 microcontroller. The microcontroller communicates with the host computer, displaying the game via serial communication. To determine assistance forces, an LQR controller is imple- mented to determine the appropriate cart positions to stabilize the pendulum and to move the cart to the middle of the monsters. Along with the cart and pendulum dynamics, a virtual spring is rendered between the stabilizing controller- determined ideal position and the current position of the cart by HandsOn-SEA to provide assistance forces to participants. The assistance is adjusted through the stiffness of the coupling spring; the larger the spring constant, the larger the assistance. C. Experimental Conditions The training efficacy of the proposed AAN training method based on Pareto optimization (test group) is compared with a control group based on a commonly employed performance- based adaptive assistance controller. Each group had an equal number of participants. The following protocols are compared: i) Test Group – Pareto Approach: As detailed in Sec- tion I-F, the AAN training method based on Pareto optimization relies on HiL characterizations of the trade- off between the task performance and the perceived change level. Pareto-based approach naturally results in a diverse set of assistance levels. After a trade-off is characterized, thresholds are introduced to select the subset of non-dominated solutions that span from 40% to 80% of the performance and the perceived challenge level of individual users. Non-dominated solutions in this subset are used in the AAN training session by selecting the assistance levels in a randomized order for each trial. If the number of non-dominated solutions is less than the number of trials, then the random selection process is restarted from the original non-dominated solution subset, after all solutions in the subset are used. On average, 32.9± 24.0 non-dominated solutions were located within the 40–80% thresholds. i) Control Group – Adaptive Assistance: The control group is trained with a commonly used adaptive assistance controller. A performance-based adaptive assistance con- troller is implemented using an adaptive staircase ap- proach, as in [11], [55], [56]. The adaptive approach is an online local search method that seeks a single assistance level for a participant during trials. The assistance level of the control group is initialized at 50% and employs a two-up-one-down variation. In this adaptive approach, the assistance decreases by 10% after every two consecutive successful trials or increases by 10% after each failure. D. Experimental Procedure The experiment consists of 6 sessions, as in Figure 4: warm- up, pre-evaluation, pre-training HiL characterization, training, post-evaluation, and post-training HiL characterization. The experimental procedure was administered identically for the control and the test groups, except for the assistance levels provided to each group during the training session. The warm-up session includes a tutorial to help participants become familiar with the rules and mechanics of the game. Participants play the game with several levels of assistance. The warm-up session is concluded when a participant displays a clear understanding of the rules and can play the game without unpremeditated failure. Pre- and post-evaluations measure the performance of the participants under no assistance. A comparison of pre- versus post-evaluation sessions provides a measure of participants’ progress after the training sessions. HiL Pareto characterization sessions are used to learn the trade-off between a participant’s quantitative performance and the perceived challenge level of the game. For this purpose, two GP regression models are trained in each HiL session. During each iteration of the HiL learning, the participant plays the game with an assistance level determined by the optimization algorithm. Once the game ends, the best score of the participant is used to train a quantitative GP model. After each game, two questions are asked to the participant to determine the perceived difficulty of the game. First, the participant is asked to rate the challenge level of the game by selecting one option from “hard”, “moderate”, or “easy”, resulting in an ordinal classification. Next, (except for the first trial), the participant is asked to compare the challenge level of the last game with respect to the previous one in a pair- wise manner. The answers to these queries are used to train a qualitative GP regression model. Finally, utilizing a multi- criteria sampling strategy, the HiL optimization algorithm updates the assistance level provided for the next iteration. E. Hypotheses The human subject experiments aim to test the validity of the following hypotheses: H 1 The trade-off between the performance and the perceived challenge level of a task can be characterized by utilizing a HiL Pareto optimization approach. H 2 Comparisons of the trade-off curves characterized at the different stages of training provide a rigorous means of evaluating training performance within and between participants, across feasible assistance levels. H 3 The non-dominated solutions characterizing the trade-off between the performance and the perceived challenge can be used to design an AAN training protocol. 9 Participant 1Participant 2 Assistance LevelAssistance LevelAssistance Level Assistance Level Task PerformanceTask PerformanceTask PerformanceTask Performance Fig. 5. In the top rows, the orange and blue lines show the mean of the surrogate function for the perceived challenge and the task performance, respectively, while the shading depicts their standard deviation. The bottom rows present the Pareto solutions with dark red dots, while the dominated solutions in the feasible set are shown with orange dots. The Pareto solutions inside the red squares are used during the training of the participants in the test group. V. RESULTS AND DISCUSSION In this section, we elaborate on each hypothesis presented in Section IV-E in the view of the experimental evidence. We also demonstrate how the HiL Pareto optimization approach can be used to assess group-level progress under all feasible assistance levels. A. Hypothesis 1 Figure 5 presents the GP regression models and the Pareto fronts, before and after the training, for two sample par- ticipants with high and intermediate performance. The data collected during the experiments are available for download from our laboratory website 1 . In the top rows, the orange lines show the mean of the surrogate function for the perceived challenge level, while the orange shading depicts its standard deviation. Similarly, the blue lines represent the mean predic- tion for the game scores, while the blue shading depicts their standard deviation. The bottom rows present the Pareto plots characterized during pre- and post-training sessions, shown in the left and right columns. In the Pareto plots, the black points represent Pareto solu- tions, while the orange points depict all feasible solutions for the problem. Pareto solutions cover a non-trivial (and possibly disconnected, as in Figure 6) subset of the set of all feasible solutions, as they capture the non-dominated solutions of the trade-off between the expectations of the perceived challenge level and the task performance (game score). For ease of vi- sualization, the Pareto plots are divided into three regions; the red, yellow, and green shades represent the perceived challenge levels of hard, moderate, and easy, respectively. Finally, the red squares indicate the Pareto solutions used during the training of the participants in the test group, according to the Pareto selection criteria detailed in Section I-F. As hypothesized in H 1 , Figure 5 provides evidence that the trade-off between the performance and the perceived challenge level of a task can be characterized via HiL Pareto optimization 1 http://hmi.sabanciuniv.edu/HiL ParetoExperimentData.zip with hybrid performance measures. Thanks to the sample- efficiency inherited from the underlying Bayesian optimiza- tion approach [35], HiL Pareto characterization converged in 6.3± 1.8 min; hence, it can easily be performed before the training sessions without inducing fatigue to the participants. The top rows in Figure 5 verify that the Bayesian multi- criteria optimization technique samples from the regions of surrogate functions that conflict with each other. The efficiency of the sampling strategy in locating the conflicting regions and the uniformity of samples in these regions are key aspects of the efficiency of the Pareto optimization technique that make its use adequate for HiL optimization. The bottom rows in Figure 5 characterize the trade-off between the performance and the perceived challenge level of the task. In particular, in the sample Pareto plots, both participants can achieve their best scores when their perceived challenge level is low, and their performance decreases as their perceived challenge level increases. Furthermore, the trade-off curves also enable the evaluation of whether the task is too challenging or too easy for the participant, by checking their performance under high and low assistance levels, respectively. If such a case is detected, it may be preferable to change the task to better cater to the abilities of the participant. Overall, the experimental results provide strong evidence that the trade-off between the performance and the perceived challenge level of a task can be characterized by utilizing a HiL Pareto optimization approach. Consequently, the results support the first hypothesis H 1 . B. Hypothesis 2 The proposed Pareto approach emphasizes the customiza- tion of training sessions based on individual performances. Accordingly, this section focuses on individual-level results. In the bottom rows of Figure 5, the changes between the pre- and post-training trade-off characterization plots capture the progression of the participants under assistance. For Par- ticipants 1 and 2, the shifts in the Pareto plots explicitly indicate that participants not only perceive the game as less 10 challenging after the training, but also their scores improve significantly. Accordingly, their post-training Pareto curves have shifted towards the right and downwards, compared to their pre-training Pareto curves. For instance, for Participant 2, the perceived level of challenge for the most difficult trial has reduced from hard to moderate, indicating that the task became easier to perform as learning took place. The positive progress in the performance can also be observed by the shrinkage of the Pareto front towards the right side, which captures the region for higher game scores. This shift indicates the game scores for the most difficult perceived challenge level have increased from 0.30 to 0.55 for Participant 2. These improvements can also be observed in the GP re- gression models, before and after the training. For instance, one can observe by inspecting the surrogate function for the perceived challenge level and game scores of Participant 2 that the performance and challenge level saturate around 70% assistance before the training, while this saturation shifts to 45% assistance after the training. Hence, only the assistance levels from 0% to 40% belong to the post-training solutions. Comparison of the pre- and post-training surrogate functions is especially useful to understand Pareto plots that consist of multiple disconnected sections. Figure 6 presents the results for a low-performing Participant 3, whose Pareto plots are harder to interpret, but the surrogate functions for the perceived challenge level and the game scores can help understand the results. While the worst performance of Participant 3 did not significantly improve from pre- to post-training, the shifts in the surrogate functions of this participant indicate that Partic- ipant 3 requires less assistance to achieve a similar level of performance to the pre-training case. In particular, the increase in performance and the decrease in challenge level shifts from 70% assistance in the pre-training characterization results to 50% assistance in the post-training results. Accordingly, the trade-off characterization captures improvements in the performance that cannot be captured by simply observing the participants’ performance without any assistance. Finally, utilizing the HiL Pareto optimization, the trade-off curves of different participants can also be compared with Participant 3 Task Performance Assistance LevelAssistance Level Task Performance Fig. 6. Results for a sample participant with low performance. The presen- tation follows the same format as in Figure 5. each other. For instance, the post-training Pareto front of Participant 2 is slightly better than the pre-training Pareto front of Participant 1, indicating that Participant 2 reaches a more advanced stage after the training, compared to the pre-training performance of Participant 1. Note that rigorous comparisons of different participants with each other, in general, is a very challenging task; a fair comparison of different participants is possible by considering the non-dominated solutions of the Pareto optimization, as they capture the best possible performances for each challenge level [44], [45]. Overall, the results support that comparisons of the trade- off curves characterized at the different stages of training can provide a rigorous means of evaluating training performance, as hypothesized in H 2 . C. Hypothesis 3 To assess the overall improvement of the participants after the training session, we compare the unassisted game scores of the participants before and after the training sessions, as depicted by violin plots in Figure 7. Furthermore, a two-way mixed-design ANOVA is conducted to examine the effects of training phase (pre- vs. post-) and group (test vs. control) on the task performance. Prior to the analysis, assumptions for ANOVA were tested and verified as follows: The Shapiro- Wilk test confirms that the residuals were normally distributed (p = .420). Levene’s test indicates the homogeneity of vari- ances for both the pre- (p = .406) and post-training phases (p = .955), supporting the assumption of equal variance across the groups. p < 0.001 1.0 0.8 0.6 0.4 0.2 0.0 Pre-Tra�n�ng Control Group Test Group Control Group Test Group Post-Tra�n�ng Normal � zed Scores Fig. 7. Violin plots of the unassisted pre- and post-training performance There is a significant main effect of training phase (f(1,32) = 37.88, p < .001, partial η 2 = 0.542), indicating a statistically significant improvement in performance from pre- to post-training across all participants. The main effect of the group is not significant (f(1,32) = 0.13, p = .725, partial η 2 = 0.004), suggesting that overall performance did not differ between the test and control groups. Furthermore, the interaction between the group and the training phase is not significant (f(1,32) = 0.08, p = .784, partial η 2 = 0.002), indicating that both groups showed similar improvements from pre- to post-training. 11 Control Group (a) Test Group (b)(c) Perce � ved Challenge Task PerformanceTask PerformanceAss�stance level Perce � ved Challenge Pre-tra�n�ng Post-tra�n�ng Pre-tra�n�ng Post-tra�n�ng Overlapp�ng po�nts Improvement �n perce�ved challenge Improvement �n task performance Pre-post change �n pred�cted performance Test Group Control Group Fig. 8. Comparison of aggregate Pareto solutions of control (a) and test (b) groups during pre- and post-training. In these figures, a shift of the Pareto front from left to right indicates improvement in task performance, while a shift from top to bottom indicates the task being perceived as less challenging. (c) Percent improvement of the available performance of the control and test groups. Since no statistically significant differences are observed under no-assistance conditions, within the power and scope of the statistical analysis, there is no evidence that one protocol outperforms the other. Together with the successful implementation of the Pareto-based protocol, these results support hypothesis H 3 . Consequently, the results demonstrate that training based on non-dominated solutions can yield AAN protocols. D. Comparison of Group Performance under Assistance As discussed in Section I, typical evaluations of manual skill training are performed when no assistance is provided, as in Section V-C, since this represents the real-life use case. On the other hand, such evaluations cannot capture performance improvements during the early phases of training or under assistance. One of the important insights provided by the proposed HiL Pareto optimization framework is that trade-offs characterized at different stages of training provide a rigorous means to fairly assess the progress of individuals, under all feasible assistance levels, as discussed in Section V-B. In this section, we demonstrate how performance analyses can be performed at a group-level. In particular, we study the average improvements of the control and test groups through their aggregate GP models. The goal of such analyses is not necessarily to show that one group is superior to the other, but to gain further insights into the performance of both training protocols under all feasible assistance levels. To estimate group-level performance, we first aggregate trained GP models by statistically averaging them among the participants and use these aggregate GP models to de- rive aggregate Pareto curves for each group during pre- and post-training characterizations. Figures 8(a) and (b) present aggregate Pareto plots for the control and test groups, re- spectively. These comparisons across assistance levels and pre/post-training characterization instances are rigorous and fair, as the Pareto solutions capture the best achievable perfor- mance predicted for each group under all feasible assistance conditions at each characterization instance. The shifts of the aggregate Pareto curves from pre- to post-training in these figures indicate observable improvements in the average task performance of both groups, together with minor improve- ments in their average perception of task difficulty. Next, we investigate the difference in pre- and post-training performance predicted by the aggregate GP models of the two groups. Figure 8(c) presents, for each assistance level, the mean change in GP-predicted performance together with 95% confidence intervals obtained via non-parametric bootstrap- ping. For each bootstrap replicate, participants are resampled with replacement, the aggregate GP model is evaluated at each assistance level, and the group-level mean change is recomputed. A total of 5000 bootstrap samples are used, and percentile-based confidence intervals are reported. Figure 8(c) shows that both groups possess wide confi- dence intervals at each assistance level, reflecting substan- tial inter-participant variability in GP-predicted performance improvements. Given the confidence intervals for the two groups largely overlap, it can be concluded that the overall performance improvement is broadly similar for both training protocols across assistance levels. As a descriptive trend, the test (Pareto-based) group achieves slightly higher mean performance gains, especially within the 40%-90% assistance level range. While these differences are not statistically robust for this particular human-subject experiment, similar analyses may lead to stronger trends for other experiments, providing valuable insights to the designers for making informed deci- sions while further optimizing the training protocol. To have a better understanding of the assistance levels selected by the AAN protocols for an identical group of participants, Figure 9 presents the assistance level provided to Adapt�ve ass�stance prov�ded to the control group Prospect�ve ass�stance to be prov�ded to the control group �f they were �n the test (Pareto) group Iterat�on Ass � stance Level Fig. 9. Assistance provided to participants in the control group (adaptive assistance) versus the prospective assistance that would have been provided to these participants in the control group if they were in the test group (Pareto approach). Lines show the average assistance levels provided to participants, while shaded areas indicate one standard deviation. 12 the control group in comparison to the prospective assistance level that would have been provided to the same control group, if these participants were placed in the test group. Since the assistance levels used in the Pareto-based training protocol are independent of the participant’s instantaneous performance, they can be determined solely from their corresponding Pareto solutions. Consequently, the prospective assistance levels to be provided to the participants if they were in the test group are determined based on their pre-training HiL characterization. Figure 9 indicates that the adaptive staircase method pro- vided mean assistance levels around 60% to the control group, while the Pareto-based approach would have selected prospective assistance levels closer to 40% on average for the same set of individuals in the control group. The difference between the mean assistance levels of the two approaches provides a possible explanation for larger performance change trends observed for the Pareto-based group between 40%-90% assistance levels in Figure 8(c). Training under lower mean assistance for the Pareto-based training protocol may have exposed participants to conditions demanding greater active control more frequently, and this could have contributed to the observed trend between 40%-90% assistance levels. Since the confidence intervals for both the assistance level distributions in Figure 9 and the performance changes in Figure 8 have substantial overlap, the assistance experienced by the two groups is quite similar at the group-level, and the observed trends need to be interpreted with caution. Overall, the analysis in this subsection does not provide evidence of significant differences between the two tested training protocols across assistance levels for this particular human-subject experiment. On the other hand, the proposed analysis method, based on comparing aggregate Pareto plots and GP models captured during pre- and post-training, pro- vides a rigorous method to fairly compare the performance of the control and test groups across all assistance levels, thereby enabling the designer to capture potential differences among training protocols to gain useful insights. VI. CONCLUSION This study demonstrates the feasibility of HiL Pareto opti- mization, its potential to help with the design of new AAN controllers, and its novel use for individual and group-level performance evaluations and comparisons. Human subject experiments with qualitative and quantitative cost functions are provided to demonstrate the use of two different forms of feedback in HiL Pareto optimization. While the simplest models are used to promote ease of presentation, the proposed HiL Pareto optimization approach is generic and can be easily extended to Pareto solutions for any number and type of cost functions. Similarly, while this study has been conducted for the single decision variable of the assistance level, the proposed Pareto optimization method trivially extends to a larger number of decision variables. Furthermore, while we have utilized USeMO as an efficient optimizer, similar results may be achieved by utilizing other sample-efficient Pareto optimization approaches. A. Limitations of the Study While our results show that the Pareto-based training pro- tocol performs comparably with a commonly used adaptive method, the number of participants in this study is sufficient to detect medium effect sizes. As a result, smaller differ- ences between training strategies may have gone undetected. Similarly, group-level comparisons based on GP estimates provide only weak trends for our study, as the confidence regions of groups largely overlap. Further experiments with a larger number of participants, multiple training sessions, and long-term retention evaluations may offer deeper insights into performance differences not captured in this feasibility study. Additionally, our current experiment evaluates the effective- ness of Pareto-based AAN training using a simple protocol to facilitate the presentation of the underlying Pareto-based design concept. The design of an effective training protocol is an involved process that needs consideration of multiple aspects, such as scheduling of breaks and repetitions, that go beyond the determination of the proper level of assistance. Overall, it is important to re-emphasize that our focus in this feasibility study is to demonstrate that the characterized trade-off can help with the design of promising and effective AAN protocols by providing insights. Designs of more so- phisticated AAN protocols and claims of improved training efficacy require further optimizations, such as sensitivity and robustness analyses, which are beyond the scope of this study. B. Ongoing Works The determination and evaluation of more sophisticated AAN approaches are parts of our ongoing work. Exploring and validating these approaches through systematic experi- mentation will help us better understand how to optimize AAN training using the HiL Pareto optimization framework. Since the task difficulty-user performance trade-off also ex- ists in the rehabilitation context, the proposed HiL Pareto opti- mization approach is directly applicable to these applications. Figure 10 provides a snapshot of our ongoing studies, in which HiL Pareto optimization is applied to stroke patients with the AssistOn-Arm six DoF upper extremity exoskeleton [57]–[59]. For the utilization of the proposed framework in rehabilita- tion applications, the perceived challenge level cost function and the underlying HiL Pareto optimization method do not require any changes, while the rehabilitation task to be per- formed and the quantitative performance metric are modified with more clinically relevant ones. HiL Pareto optimization can still be performed over the design variable of the assistance level provided to the user, capturing the force-feedback applied to the patient within a virtual tunnel around the nominal path of the rehabilitation exercise. The multi-DoF nature of the exoskeleton becomes relevant only when the required assistance level in task space is mapped to the joint space of the robot for the appropriate actuator torques. As in the case of motor skill training, the number of design variables can be easily extended to include other relevant variables, such as the diameter of the virtual tunnel utilized during rehabilitation. Overall, our results indicate that the proposed HiL Pareto optimization approach holds promise for applications in both manual skill training and robot-assisted rehabilitation. 13 ACKNOWLEDGEMENT This work has been partially supported by TUBITAK Grants 120N523 and 23AG003. REFERENCES [1] Y. Li, V. Patoglu, and M. K. O’Malley, “Negative efficacy of fixed gain error reducing shared control for training in virtual environments,” ACM Transactions on Applied Perception, vol. 6, no. 1, p. 1–21, 2009. [2] A. Erdogan and V. Patoglu, “Slacking prevention during assistive contour following tasks with guaranteed coupled stability,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems, 2012, p. 1587–1594. [3] J. Moreno-Valenzuela, “Velocity field control of robot manipulators by using only position measurements,” Journal of the Franklin Institute, vol. 344, no. 8, p. 1021–1038, 2007. [4] P. Y. Li and R. Horowitz, “Passive velocity field control of mechanical manipulators,” IEEE Trans. on Robotics and Automation, vol. 15, no. 4, p. 751–763, 1999. [5] R. Colombo, I. Sterpi, A. Mazzone, C. Delconte, and F. Pisano, “Development of a progressive task regulation algorithm for robot-aided rehabilitation,” in Int. Conf. of the IEEE Engineering in Medicine and Biology Society, 2011, p. 3123–3126. [6] A. Erdogan and V. Patoglu, “Online Generation of Velocity Fields for Passive Contour Following,” in IEEE World Haptics Conference, 2011, p. 245–250. [7] U. Keller, G. Rauter, and R. Riener, “Assist-as-needed path control for the PASCAL rehabilitation robot,” in IEEE Int. Conf. on Rehabilitation Robotics, 2013, p. 1–7. [8] H. J. Asl, M. Yamashita, T. Narikiyo, and M. Kawanishi, “Field- Based Assist-as-Needed Control Schemes for Rehabilitation Robots,” IEEE/ASME Trans. on Mechatronics, vol. 25, no. 4, p. 2100–2111, 2020. [9] J. Lopes, C. Pinheiro, J. Figueiredo, L. Reis, and C. Santos, “Assist- as-needed Impedance Control Strategy for a Wearable Ankle Robotic Orthosis,” in IEEE Int. Conf. on Autonomous Robot Systems and Competitions, 2020, p. 10–15. [10] H. Krebs, J. Palazzolo, L. Dipietro, M. Ferraro, J. Krol, K. Rannekleiv, B. Volpe, and N. Hogan, “Rehabilitation Robotics: Performance-Based Progressive Robot-Assisted Therapy,” Autonomous Robots, vol. 15, p. 7–20, 2003. [11] Y. Li, J. C. Huegel, V. Patoglu, and M. K. O’Malley, “Progressive shared control for training in virtual environments,” in World Haptics, 2009, p. 332–337. [12] M. Sarac, E. Koyas, A. Erdogan, M. Cetin, and V. Patoglu, “Brain Computer Interface based robotic rehabilitation with online modification of task speed,” in IEEE Int. Conf. on Rehabilitation Robotics, 2013, p. 1–7. [13] O. Ozdenizci, M. Yalcın, A. Erdogan, V. Patoglu, M. Grosse-Wentrup, and M. Cetin, “Electroencephalographic identifiers of motor adaptation learning,” Journal of Neural Engineering, vol. 14, no. 4, p. 046027, 2017. Fig. 10. A snapshot during HiL Pareto characterization and AAN training with the AssistOn-Arm upper-extremity exoskeleton [14] R. Yang, Z. Shen, Y. Lyu, Y. Zhuang, L. Li, and R. Song, “Voluntary Assist-as-Needed Controller for an Ankle Power-Assist Rehabilitation Robot,” IEEE Trans. on Biomedical Engineering, vol. 70, no. 6, p. 1795–1803, 2023. [15] J. L. Emken, J. E. Bobrow, and D. J. Reinkensmeyer, “Robotic move- ment training as an optimization problem: designing a controller that assists only as needed,” Int. Conf. on Rehabilitation Robotics, p. 307– 312, 2005. [16] V. Squeri, A. Basteris, and V. Sanguineti, “Adaptive regulation of assistance ‘as needed’ in robot-assisted motor skill learning and neuro- rehabilitation,” in IEEE Int. Conf. on Rehabilitation Robotics, 2011, p. 1–6. [17] S. Pareek, H. J. Nisar, and T. Kesavadas, “AR3n: A Reinforcement Learning-Based Assist-as-Needed Controller for Robotic Rehabilita- tion,” IEEE Robotics & Automation Magazine, p. 2–10, 2023. [18] B. Zhong, W. Niu, E. Broadbent, A. McDaid, T. M. C. Lee, and M. Zhang, “Bringing Psychological Strategies to Robot-Assisted Phys- iotherapy for Enhanced Treatment Efficacy,” Frontiers in Neuroscience, vol. 13, 2019. [19] A. Koenig, X. Omlin, L. Zimmerli, M. Sapa, C. Krewer, M. Bolliger, F. Mueller, and R. Riener, “Psychological state estimation from physi- ological recordings during robot-assisted gait rehabilitation,” Journal of Rehabilitation Research and Development, vol. 48, p. 367–385, 2011. [20] R. Baud, A. Manzoori, A. Ijspeert, and M. Bouri, “Review of Control Strategies for Lower-limb Exoskeletons to Assist Gait,” Journal of NeuroEngineering and Rehabilitation, vol. 18, no. 119, 2021. [21] D. Mahfouz, O. Shehata, E. Morgan, and F. Arrichiello, “A Compre- hensive Review of Control Challenges and Methods in End-Effector Upper-Limb Rehabilitation Robots,” Robotics, vol. 13, no. 12, p. 181, 2024. [22] E. Brochu, V. M. Cora, and N. de Freitas, “A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning,” CoRR, 2010. [Online]. Available: https://arxiv.org/abs/1012.2599 [23] W. Felt, J. C. Selinger, J. M. Donelan, and C. D. Remy, “Body-in-the- loop: Optimizing device parameters using measures of instantaneous energetic cost,” PLOS One, vol. 10, no. 8, p. 1–21, 2015. [24] J. R. Koller, D. H. Gates, D. P. Ferris, and D. C. Remy, “Body-in-the- Loop Optimization of Assistive Robotic Devices: A Validation Study,” in Robotics: Science and Systems, 2015. [25] Y. Ding, I. Galiana, A. Asbeck, S. de Rossi, J. Bae, T. Santos, V. de Araujo, S. Lee, K. Holt, and C. Walsh, “Biomechanical and physiological evaluation of multi-joint assistance with soft exosuits,” IEEE Trans. on Neural Systems and Rehab. Eng., vol. 25, no. 2, p. 119–130, 2016. [26] J. Zhang, P. Fiers, K. A. Witte, R. W. Jackson, K. L. Poggensee, C. G. Atkeson, and S. H. Collins, “Human-in-the-loop optimization of exoskeleton assistance during walking,” Science Robotics, vol. 356, no. 6344, p. 1280–1284, 2017. [27] D. Sadigh, A. Dragan, S. Sastry, and S. Seshia, “Active Preference- Based Learning of Reward Functions,” in Robotics: Science and Systems, vol. 13, 2017. [28] K. Li, M. Tucker, E. Biyik, E. Novoseller, J. Burdick, Y. Sui, D. Sadigh, Y. Yue, and A. Ames, “ROIAL: Region of interest active learning for characterizing exoskeleton gait preference landscapes,” in IEEE Int. Conf. on Robotics and Automation, 2021, p. 3212–3218. [29] E. Biyik, N. Huynh, M. Kochenderfer, and D. Sadigh, “Active Preference-Based Gaussian Process Regression for Reward Learning,” in Robotics: Science and Systems, 2020. [30] B. Catkin and V. Patoglu, “https://ieeexplore.ieee.org/document/10102327 Preference-Based Human-in-the-Loop Optimization for Perceived Realism of Haptic Rendering,” IEEE Transactions on Haptics, vol. 16, no. 4, p. 470–476, 2023. [31] H. Tolasa, B. Catkin, and V. Patoglu, “Human-in-the-Loop Optimization of Perceived Realism of Multi-Modal Haptic Rendering under Conflict- ing Sensory Cues,” IEEE Trans. on Haptics, vol. 18, no. 2, p. 295–311, 2025. [32] A. Mathern, O. Steinholtz, A. Sj ̈ oberg, M. ̈ Onnheim, K. Ek, R. Rem- pling, E. Gustavsson, and M. Jirstrand, “Multi-objective constrained Bayesian optimization for structural design,” Structural and Multidis- ciplinary Optimization, vol. 63, 2021. [33] S. Suzuki, S. Takeno, T. Tamura, K. Shitara, and M. Karasuyama, “Multi-objective Bayesian Optimization using Pareto-frontier Entropy,” in Int. Conf. on Machine Learning, vol. 119, 2020, p. 9279–9288. [34] B. Paria, K. Kandasamy, and B. P ́ oczos, “A Flexible Framework for Multi-Objective Bayesian Optimization using Random Scalarizations,” 14 in Uncertainty in Artificial Intelligence Conf., vol. 115, 2020, p. 766– 776. [35] S. Belakaria, A. Deshwal, N. K. Jayakodi, and J. R. Doppa, “Uncertainty-aware search framework for multi-objective Bayesian op- timization,” in AAAI Conf. on Artificial Intelligence, vol. 34, no. 06, 2020, p. 10 044–10 052, An extended version is available online at https://arxiv.org/abs/2204.05944. [36] M. Emmerich, J. W. Klinkenberg, and N. Bohrweg, “The computation of the expected improvement in dominated hypervolume of Pareto front approximations,” Leiden University, The Netherlands, Technical Report LIACS-TR9-2008, 2008. [37] D. Hernandez-Lobato, J. Hernandez-Lobato, A. Shah, and R. Adams, “Predictive Entropy Search for Multi-objective Bayesian Optimization,” in Int. Conf. on Machine Learning, vol. 48, 2016, p. 1492–1501. [38] P. Y. Papalambros and D. J. Wilde, Principles of optimal design: modeling and computation. Cambridge University Press, 2000. [39] R. T. Marler and J. S. Arora, “The weighted sum method for multi- objective optimization: New insights,” Structural and Multidisciplinary Optimization, vol. 41, no. 6, p. 853–862, 2010. [40] R. Unal, G. Kiziltas, and V. Patoglu, “A multi-criteria design optimiza- tion framework for haptic interfaces,” in IEEE Haptics Symposium, 2008, p. 231–238. [41] R. Astudillo, K. Li, M. Tucker, C. X. Cheng, A. D. Ames, and Y. Yue, “Preferential Multi-Objective Bayesian Optimization,” CoRR, 2024. [Online]. Available: https://arxiv.org/abs/2406.14699 [42] H. Tolasa, G. Gemalmaz, and V. Patoglu, “Active Learning of Fractional- Order Viscoelastic Model Parameters for Realistic Haptic Rendering,” CoRR, 2025. [Online]. Available: https://arxiv.org/abs/2512.00667 [43] X. Zhang, A. Fredriksen, S. Palmcrantz, and E. M. Gutierrez-Farewik, “Biplanar Ankle Assistance for Dropfoot Gait Post-Stroke with Multi- Objective Human-In-the-Loop Optimization: A Case Study,” in Int. Conf. on Rehabilitation Robotics, 2025, p. 1132–1138. [44] A. K. Bonab and V. Patoglu, “Musculoskeletal Simulation-Based Multi- criteria Optimization Framework for Exoskeleton Design,” IEEE Trans. on Neural Systems and Rehabilitation Engineering, 2026, Early Access. [Online]. Available: https://ieeexplore.ieee.org/document/11367100 [45] Y. Aydin, O. Tokatli, V. Patoglu, and C. Basdogan, “A Computational Multicriteria Optimization Approach to Controller Design for Physical Human-Robot Interaction,” IEEE Trans. on Robotics, vol. 36, no. 6, p. 1791–1804, 2020. [46] R. Garnett, Bayesian Optimization. Cambridge: Cambridge University Press, 2023. [47] W. Chu and Z. Ghahramani, “Preference learning with Gaussian pro- cesses,” in Int. Conf. on Machine Learning, 2005, p. 137–144. [48] C. E. Rasmussen and C. Williams, Gaussian Processes for Machine Learning. MIT Press, 2005. [49] M. Tucker, M. Cheng, E. Novoseller, R. Cheng, Y. Yue, J. W. Bur- dick, and A. D. Ames, “Human Preference-Based Learning for High- dimensional Optimization of Exoskeleton Walking Gaits,” in IEEE Int. Conf. on Robotics and Systems, 2020, p. 3423–3430. [50] A. Otaran, O. Tokatli, and V. Patoglu, “Physical Human-Robot Interac- tion Using HandsOn-SEA: An Educational Robotic Platform With Series Elastic Actuation,” IEEE Trans. on Haptics, vol. 14, no. 4, p. 922–929, 2021. [51] F. E. Tosun and V. Patoglu, “Necessary and Sufficient Conditions for the Passivity of Impedance Rendering With Velocity-Sourced Series Elastic Actuation,” IEEE Trans. on Robotics, vol. 36, no. 3, p. 757–772, 2020. [52] C. U. Kenanoglu and V. Patoglu, “Passive Realizations of Series Elastic Actuation: Effects of Plant and Controller Dynamics on Haptic Render- ing Performance,” IEEE Trans. on Haptics, vol. 17, no. 4, p. 882–899, 2024. [53] C. U. Kenanoglu and V. Patoglu, “Effect of Inherent Damping of the Series Elastic Element on Rendering Performance and Passivity of Interaction Control,” ASME Journal of Dynamic Systems, Measurement, and Control, vol. 147, no. 5, p. 051008, 2025. [54] C. U. Kenanoglu and V. Patoglu, “Effect of Reduced-Order Modelling on Passivity and Rendering Performance Analysis of Series Elastic Actuation,” IEEE Robotics and Automation Letters, vol. 10, no. 6, p. 5745–5752, 2025. [55] J. Anguera, J. Boccanfuso, J. Rintoul, O. Claflin, F. Faraji, J. Janowich, E. Kong, Y. Larraburo, C. Rolle, E. Johnston, and A. Gazzaley, “Video game training enhances cognitive control in older adults,” Nature, vol. 501, p. 97–101, 09 2013. [56] R. Gray, “Transfer of Training from Virtual to Real Baseball Batting,” Frontiers in Psychology, vol. 8, 2017. [57] H. Argunsah, B. Yalcin, M. A. Ergin, G. Coruhlu, M. Yalcin, V. Patoglu, and Z. G ̈ uven, “Advancing Precision Rehabilitation Through a Sensor- Based 6-DoF Robotic Exoskeleton: Clinical Validation and Ergonomic Assessment,” Sensors, vol. 26, no. 1, 2026. [58] M. Ergin and V. Patoglu, “AssistOn-SE: A Self-Aligning Shoulder- Elbow Exoskeleton,” in IEEE Int. Conf. on Robotics and Automation, 2012, p. 2479–2485. [59] M. Yalcin and V. Patoglu, “Kinematics and Design of AssistOn-SE: A Self-Adjusting Shoulder-Elbow Exoskeleton,” in IEEE Int. Conf. on Biomedical Robotics and Biomechatronics, 2012, p. 1579–1585. Harun Tolasa received his B.Sc. degree in me- chanical engineering from Bilkent University (2021) and his M.Sc. in mechatronics engineering from Sabanci University (2024). Currently, he is pursuing his Ph.D. degree at Sabanci University. His research interests include active learning, human-in-the-loop optimization, and haptic rendering. Volkan Patoglu is a full professor in mechatron- ics engineering at Sabanci University. He received his Ph.D. degree in mechanical engineering from the University of Michigan, Ann Arbor (2005) and worked as a post-doctoral researcher at Rice Univer- sity (2006). His research is in the area of physical human-machine interaction, in particular, design and control of force feedback robotic systems with ap- plications to rehabilitation. His research extends to cognitive robotics. He has served as an associate ed- itor for IEEE Transactions on Haptics (2013–2017), IEEE Transactions on Neural Systems and Rehabilitation Engineering (2018– 2023), and IEEE Robotics and Automation Letters (2019–2024).