Paper deep dive
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/16/2026, 2:39:37 AM
Summary
The paper introduces HumanTracker, a comprehensive benchmark for humanoid motion tracking, and HumanScore, a preference-aligned evaluation metric. HumanTracker provides 153 hours of optical motion data across four families (Daily, Highly Dynamic, Interaction, Ground) to address the lack of diversity in existing test suites. HumanScore is a reward model trained on 12K motion pairs (24K motions) that better predicts human preferences and identifies physical artifacts like foot skating and instability, which traditional kinematic metrics (e.g., MPJPE) often miss.
Entities (11)
Relation Signals (9)
HumanTracker → containsdatafrom → optical motion trajectories
confidence 95% · The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers
HumanScore → istrainedon → 12K motion pairs
confidence 94% · HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions.
HumanScore → outperforms → MPJPE
confidence 93% · HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
HumanTracker → usessimulator → MuJoCo
confidence 92% · execute the policy through a common MuJoCo evaluation entry point.
HumanScore → usesarchitecture → Transformer
confidence 91% · producing a scalar HumanScore via a temporal Transformer.
HumanTracker → evaluates → SONIC
confidence 90% · We comprehensively evaluate several state-of-the-art trackers, including... SONIC... under our benchmark
HumanTracker → evaluates → Humanoid-GPT
confidence 90% · We comprehensively evaluate several state-of-the-art trackers, including... Humanoid-GPT... under our benchmark
HumanTracker → evaluates → TWIST2
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Tags
Links
- Source: https://arxiv.org/abs/2608.13555v1
- Canonical: https://arxiv.org/abs/2608.13555v1
Trouble viewing inline? Open PDF directly →
Full Text
57,683 characters extracted from source content.
Expand or collapse full text
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Dairu Liu 1,3,∗ , Zekun Qi 2,3,∗ , Jiayu Zeng 3,∗ , Ruixi Yu 2,3 , Yu Guan 2,3 , Yintianrun Zhang 3 , Xuchuan Chen 2,3 , Sikai Liang 3,4 , Zekai Li 2,3 , Chenghuai Lin 3 , Xinqiang Yu 3 , Wenyao Zhang 4 , He Wang 3,5,† , Li Yi 2,3,6,† 1 Nankai University 2 Tsinghua University 3 Galbot 4 Shanghai Jiao Tong University 5 Peking University 6 Shanghai Qi Zhi Institute ∗ Equal Contribution, † Corresponding author Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss. Date: August 14, 2026 Project Page: https://dairuliu.github.io/humantracker Code: https://github.com/GalaxyGeneralRobotics/HumanTracker Daily Highly Dynamic Interactions Ground (b).EvaluationBenchmark (a).HumanScoreRewardModel ✗ ✓ TemporalTransformer . . . HumanScore Observation Sequence Humanpreferencedata Optical Mocap 150+ Hours 4Categories Figure 1 HumanTracker overview. Left: we train a Human-Aligned Reward Model on pairwise preference data from tracking rollouts, producing a scalar HumanScore via a temporal Transformer. Right: the HumanTracker Benchmark provides 153 hours of motion trajectories across four families: daily tasks, highly dynamic motions, interaction motions, and ground-level movements. 1 arXiv:2608.13555v1 [cs.RO] 13 Aug 2026 1 Introduction Motion tracking is becoming the cornerstone of hu- manoid robot control [22,4,24,42,19,30,37]. This ability drives applications in teleoperation and whole- body imitation [43,44,9,7,16,10,6]. Most systems track a reference motion using a learned feedback policy in physics simulation [22,4,24,42]. Yet, mea- suring the true quality of this tracking remains sur- prisingly difficult. Two major issues currently limit progress in this field. The first major issue is how we measure success. Mo- tion tracking looks easy to measure: one can compare the rollout pose to the reference pose and report kine- matic errors such as mean joint error and key point error [4,24,42,49,29]. However, we find a clear gap between these numbers and what people see in videos. A rollout may achieve low kinematic error but still look bad, especially under contacts [12,41,1,40]. Feet may slide on the ground, and contacts may break and reattach at the wrong time; these artifacts are exactly what contact-aware global motion reconstruc- tion and refinement methods target in video-based settings as well [32,47,45]. These artifacts are the key aspect to decide whether the motion looks stable, smooth and human-like. This gap is not a corner case but a consequence of what kinematic metrics measure. By treating tracking as a per-frame pose matching problem and averaging errors over joints and time, these metrics fail to capture contact, support, balance, and the accumulation of errors in closed-loop control. As a result, two rollouts can have similar pose errors yet differ substantially in stability. Observers reliably prefer the rollout with clean contacts and steady support, even when MPJPE is similar. Figure 1 highlights this mismatch. The second major issue is the scope of evaluation data. Despite the availability of large motion repos- itories such as PHUMA [12] and SONIC [24], the most commonly used evaluation suite for humanoid tracking is still an AMASS test set with only 140 sequences [26]. This small set lacks diversity and under-represents the long tail of human movement, including challenging contact transitions, asymmetric balancing, and complex recoveries. Moreover, results are often summarized as a single aggregate score, without a detailed breakdown by motion category, making it difficult to pinpoint exactly where and why a tracker fails. Table 1 compares HumanTracker with representative large-scale motion datasets in scale, categorization and text annotation. We introduce HumanTracker to address these prob- Table 1 Comparison of large-scale humanoid motion datasets. HumanTracker uniquely provides approximately 153 hours of motion trajectories explicitly organized into dis- tinct motion categories for comprehensive evaluation. DatasetClips Hours Categories Text Label AMASS [26]>11K >40NoNo HumanML3D [2] 14.6K 28.6NoYes PHUMA [12]76K 73NoNo HumanTracker25K 1534Yes lems. To address the metric misalignment, we de- velop a preference-aligned evaluation: we collect hu- man comparisons on synchronized tracking videos and train a reward model to predict these prefer- ences [5,27,38]. We call this metric HumanScore. Unlike joint error, HumanScore explicitly targets hu- man preferences, perceptual stability, and realistic physical contacts. It penalizes the exact physical artifacts that look wrong to human observers. Impor- tantly, we demonstrate that HumanScore captures nuanced perceptual qualities that cannot be simply reduced to rule-based diagnostics such as foot slip or root drift. Extensive experiments and visual analyses demonstrate the accuracy and zero-shot generaliza- tion of HumanScore. To address the lack of diversity in existing test suites, we provide approximately 153 hours of op- tical motion trajectories paired with category and text labels, complementing existing large-scale mo- tion sources [26,20,46,12]. All source motions are recorded in a controlled studio with a multi- camera optical system by 24 professional performers, including dance teachers and fitness coaches, yield- ing high-fidelity references for contact-rich tracking evaluation. We explicitly organize the dataset into four distinct families covering daily tasks, highly dy- namic movements, interaction motions, and ground- level motions, enabling per-family metrics and fine- grained diagnosis of failure modes. HumanTracker substantially exceeds prior mocap datasets used for humanoid tracking. We comprehensively eval- uate several state-of-the-art trackers, including GMT, TWIST2, SONIC, and Humanoid-GPT, under our benchmark [4, 44, 24, 29]. In summary, our main contributions are threefold. First, we highlight the failure of traditional kinematic metrics and introduce HumanScore to align evalu- ation with humans. Second, we provide a massive motion tracking benchmark featuring approximately 153 hours of diverse, categorized optical trajectories. Third, we establish a rigorous and standardized evalu- ation protocol to ensure future progress in humanoid tracking is both measurable and meaningful. DailyHighly Dynamic Interaction Ground Figure 2 Motion taxonomy. HumanTracker groups motions into four families according to their dominant tracking and contact regimes. The taxonomy supports category-level diagnosis rather than only a single aggregate score. 2 Related Work 2.1 Humanoid motion tracking DeepMimic [28] established reference conditioned pol- icy learning for physics-based character control. PHC and UHM [22,23] extended this paradigm to diverse humanoid motion. Recent systems broaden motion coverage and control robustness through several mech- anisms. GMT, SONIC, Humanoid-GPT and Uni- Tracker [4,24,29,42] emphasize scale and general tracking. ResMimic, MHC and iCTRL [52,6,40] ex- plore residual correction, multimodal commands and constrained control. Any2Track [49] adapts tracking policies to changes in terrain, external forces and physical properties. HumanoidPF [39] uses a hu- manoid potential field to learn collision avoidance skills for cluttered indoor scenes. Adversarial Dif- ferential Discriminators [51] provide a learned mo- tion imitation objective. OmniTrack [17] constructs physically consistent references, while LIMMT [8] studies how carefully selected training motions can improve tracking. The same capability underpins whole-body teleoperation and imitation systems, in- cluding TWIST, OmniH2O, HumanPlus, CLONE and H2O [43,9,7,16,10]. These advances have produced increasingly capable trackers, but their re- ported performance is difficult to compare because results depend not only on the learned policy, but also on the reference set, simulator, action representation, initialization, and termination rule. HumanTracker therefore treats evaluation as a controlled experi- ment. Tracker specific policy interfaces are preserved, whereas the reference representation, rollout account- ing, and reported metrics are standardized. 2.2 Motion data evaluation Large motion repositories such as AMASS, Motion-X and Motion-X++ [26,20,46] have substantially ex- panded the diversity of human motion available for learning. Recent works [2,48,14,50] further support semantic retrieval and conditional generation. For Table 2 HumanTracker dataset statistics. The benchmark contains approximately 153 hours and 25K clips across four complementary tracking regimes. FamilyHours #ClipsTypical challenges Daily899.7ksteady locomotion, mild contacts Highly Dynamic112.7kimpacts, aerial phases, fast footwork Interaction4810.9k human-like, stable, smooth hands-body coordination Ground51.6klow posture, multi-contact transitions Total15325Kdiverse humanoid tracking, however, dataset size alone does not define a useful benchmark. Human motion must also be retargeted to the robot morphology, remain physically plausible around contacts, and contain suf- ficiently difficult regimes to expose controller failures. PHUMA, OmniRetarget and GMR [12,41,1] ad- dress these requirements through physically grounded data or robot retargeting. WHAM, ProxyCap and RoHM [32,47,45] address related errors in global human motion reconstruction. Switch-JustDance [11] provides a complementary benchmark based on whole- body skills from a commercial motion game. Exist- ing tracking evaluations nevertheless remain concen- trated on comparatively small test sets and often col- lapse heterogeneous motions into one aggregate value. HumanTracker complements prior motion sources with a large, explicitly categorized test bed whose four families separate steady locomotion, rapid im- pact rich motion, interaction motions, and ground level transitions with multiple contacts. 3HumanTracker: Benchmark and Preference-Aligned Evaluation 3.1 The HumanTracker benchmark Motion collection and processing. The released Hu- manTracker benchmark contains approximately 153 hours of optical motion trajectories from 24 profes- sional performers. The performers include dance teachers, fitness coaches, tennis coaches and full-time motion-capture actors. The performers and recording plan were chosen to cover both routine movement and motions that place substantially different demands on a humanoid controller. We retarget each fitted human motion to the benchmark humanoid using General Motion Retargeting (GMR) [1]. Because a visually valid human recording is not automatically a valid robot reference, we inspect the retargeted sequences and remove segments with capture or pro- cessing artifacts such as unexplained floating, ground penetration and discontinuous contacts. Each re- leased clip contains a top-level motion-family label, a natural-language description, a fitted SMPL se- quence [21] and a robot-space reference trajectory in qposformat. These representations support seman- tic subset selection while keeping the input to every evaluated tracker identical. Figure 2 illustrates the resulting motion coverage, and Table 2 summarizes the family-level scale of the benchmark. A taxonomy for diagnostic evaluation. The four mo- tion families are defined by the failure regimes that they expose. Daily contains walking, turning and routine gestures, and therefore measures steady-state stability and residual drift under comparatively regu- lar support. Highly Dynamic contains jumps, kicks, acrobatics and fast dance footwork, for which im- pacts and rapid support switching amplify phase and timing errors. Interaction contains human body mo- tions associated with actions involving objects or the surrounding environment. It evaluates hand, arm and whole body coordination in the human refer- ence. These trajectories can also provide kinematic priors for humanoid manipulation. Ground covers kneeling, sitting, rolling and recovery, where a low centre of mass and multiple simultaneous contacts make the controller sensitive to contact geometry and friction. The distribution intentionally reflects the fre- quency of the captured activities rather than forcing equal family sizes; all benchmark results are there- fore reported by family as well as in aggregate. The complete dataset is split 9:1 into disjoint training and test partitions, with the family distribution preserved and duplicate motions kept within one partition. 3.2 Standardized tracker evaluation For every method, we convert the reference to the same 29-DoF humanoidqposrepresentation and ex- ecute the policy through a common MuJoCo evalua- tion entry point. Each tracker retains its native policy observations and action decoder, and the evaluator instead standardizes the motion list, robot model, reference indexing, rollout accounting and metric im- plementation. The control trajectory is recorded at 50 Hz. At every step, the evaluator stores the simu- lated generalized position and velocity, policy action and motor target, foot contacts and contact forces, Figure 3 Preference collection interface. Annotators view two synchronized rollouts of the same reference segment. Display order is randomized and the available responses are Left is better, Right is better, Similar and Cannot compare. foot and pelvis velocities, and 14 keypoint poses and spatial velocities. The same state history is used for conventional diagnostics and HumanScore, which prevents differences in post-processing from being mistaken for differences between trackers. We use the same tracking metric and success crite- rion with SONIC [24] for each tracker. It measures vertical position error at the pelvis, both ankles and both wrists, together with pelvis rotation error. The episode fails when any vertical error exceeds 0.25 m, the pelvis rotation error exceeds 1 rad, or the gener- alized position or velocity contains a nonfinite value. We report Succ as the fraction of completed episodes and MPJPE as the mean absolute error over the 29 actuated joint angles, in radians, over the exe- cuted portion of each rollout. Additional diagnostics include joint-velocity error, keypoint-position error, foot-contact agreement and finite-difference joint ac- celeration and jerk. All reported benchmark values use the HumanTracker test split. 3.3 Preference data construction Rollout segmentation for paired comparison. The pref- erence pool is generated exclusively from motions in the HumanTracker training split, so no bench- mark test motion is used to train HumanScore. For each source motion, GMT [4], Humanoid-GPT [29], SONIC [24] and TWIST2 [44] produce aligned roll- outs of the same robot-space reference. At 50 Hz, every rollout is divided into consecutive 250-frame windows, each spanning 5 s. If a rollout ends with a shorter window, the final short window is retained. Uniform rollout pair sampling. We concatenate all aligned windows in a deterministic order by motion family, source motion and temporal position, and as- sign each window a global index. From this catalogue we select a uniformly spaced set of unique indices over the full ordered list. This construction samples the complete HumanTracker training distribution without manually favouring visually difficult exam- ples or any particular family. Each selected window yields exactly one comparison between two of the four trackers. The six unordered tracker combinations are allocated equally, so every pairing contributes the same number of comparisons and each tracker ap- pears equally often. Candidate order is alternated within each combination and the final task order is deterministically shuffled, so no tracker is favoured by display position. The resulting preference pool is partitioned into train and test sets. Human annotation and split. The annotation panel comprised six doctoral researchers specializing in hu- manoid robotics, providing domain expert judgments of balance, contact, tracking stability and motion naturalness. Using the interface shown in Figure 3, they first decide whether either rollout fails to com- HumanScore Reward Model Chosen Trajectory Feature Extraction pose velocity contact ... Feature Projection Linear + LayerNorm Rejected Trajectory Feature Extraction pose velocity contact ... Feature Projection Linear + LayerNorm Transformer Encoder Positional Encoding Sinusoidal + L × Multi-Head Self-Attention Add & LayerNorm Feed-Forward Network Add & LayerNorm Pooling Reward Head Reward Head r chosen ∈ℝ r rejected ∈ℝ Bradley–Terry Loss ℒ=−log σ(r chosen − r rejected) σ: sigmoid r chosen r rejected Figure 4 HumanScore trajectory reward model. Current reference and simulated state features form one token per frame. A temporal Transformer processes the valid tokens, and masked mean pooling forms a trajectory representation. The diagram shows the Bradley–Terry loss on the unbounded rewards r chosen and r rejected for a strict comparison. plete the motion or loses balance; when both remain viable, they compare jitter, foot sliding, locomotion consistency and whole-body naturalness in that or- der. They choose Similar when neither candidate is meaningfully better and Cannot compare when both are unusable or the evidence is insufficient. The interface randomizes display order and records the underlying candidate indices, eliminating a fixed as- sociation between display side and tracker identity. The collected labels comprise strict preferences, simi- lar pairs and cannot-compare pairs. Cannot-compare pairs are excluded from reward-model optimization. We construct an 80/20 training/test split by group- ing records by the originalmotion_id; all clips from one source motion therefore remain in one partition. The split is balanced jointly over family, tracker pair- ing, label type, annotator and the remaining records define the training and test preference sets. 3.4 HumanScore Trajectory representation. Figure 4 summarizes how the reward model maps a rollout segmentτto an unbounded rewardr θ (τ)∈R. Although human preferences are collected from rendered videos, the model operates directly on simulator trajectories, avoiding dependence on camera viewpoint or ren- dering choices. Each frame is represented by a 539- dimensional vector. It contains 70 dimensions for the current reference state and 469 dimensions for simulated state, control, measured contact dynamics, root motion and current keypoint kinematics. The reported model does not use future reference residu- als. Conditioning on the current reference and rollout allows HumanScore to assess tracking quality rather than motion plausibility alone. The complete feature decomposition is provided in the appendix. The frame vector is linearly projected, normalized, and augmented with sinusoidal positional encoding before being processed by a Transformer encoder [35]. A padding mask is applied in every attention layer. The elementwise mean of the valid output tokens then forms the trajectory representation, which an MLP maps to the scalar reward. The same mask is used for the retained tail clips. Segments shorter than 250 frames are padded with zeros on the right, while the validity mask excludes padded positions from both attention and temporal pooling. Thus, full and truncated windows can be processed by the same model without introducing padding artifacts. Preference objective. For a strict pairi ∈ D, the model assigns rewardsr (i) chosen andr (i) rejected to the chosen and rejected trajectories. We define ∆ i = Table 3 Zero-shot evaluation on HumanTracker. All trackers are evaluated without training or fine-tuning on the test set. We report completion rate under the whole-body termination criterion (Succ), mean absolute joint-angle error over 29 actuated joints (MPJPE), and perceptual trajectory quality (HumanScore, 0 to 100). Higher Succ and HumanScore and lower MPJPE indicate better performance; bold denotes the best result within each motion family. DailyHighly DynamicInteractionGround Method Succ (%) MPJPE (rad) Human Score Succ (%) MPJPE (rad) Human Score Succ (%) MPJPE (rad) Human Score Succ (%) MPJPE (rad) Human Score GMT [4]17.0 0.250 2.4 36.2 0.196 7.0 81.4 0.205 11.7 0.0 0.456 4.0 TWIST2 [44]60.1 0.105 10.1 39.9 0.112 16.9 91.3 0.111 28.3 0.0 0.341 4.5 SONIC [24]93.8 0.102 49.5 82.1 0.118 41.0 97.6 0.128 54.6 20.1 0.23126.5 Humanoid-GPT [29] 94.4 0.04654.7 86.9 0.04749.2 97.2 0.07056.8 32.9 0.216 24.9 r (i) chosen − r (i) rejected and use the Bradley–Terry loss [3] L (i) diff =− logσ(∆ i ). The less frequent Similar pairs provide an equality constraint. For a Similar pairj ∈S, letr (j) a andr (j) b denote the rewards of its two trajectories, and define ∆ j = r (j) a − r (j) b . Their symmetric loss is L (j) similar =− 1 2 logσ(∆ j )− 1 2 logσ(−∆ j ). Every pair contributes once to the batch objective L = 1 |D| +|S| X i∈D L (i) diff + X j∈S L (j) similar . 3.5 HumanScore computation At evaluation time, a rolloutτofFframes is divided intoN=⌈F/250⌉consecutive windows. LetL i ≤ 250 be the number of actual frames in windows i , so P i L i =F. A short final window is padded on the right and evaluated with its validity mask. For reporting, the unbounded reward of each window is mapped to a bounded reward ρ θ (s i ) = σ(r θ (s i ))∈ (0, 1). We then average the window rewards according to the number of frames that they represent HumanScore(τ ) = 100 F N X i=1 L i ρ θ (s i ). Padding is used only to form the model input and contributes no weight to the average. Under the sig- moid mapping used at inference,ρ θ (s i ) ranges from 4×10 −8 to 0.99 over the training data distribution, providing sufficient dynamic range to distinguish tra- jectories according to human preferences. 4 Experiments 4.1 Experimental setup Evaluation setting. We evaluate GMT, TWIST2, SONIC and Humanoid-GPT through the standard- ized protocol of Sec. 3.2. Each policy retains its native observation and action-processing stack, while all methods receive the same retargeted reference motions and are measured by the same evaluator. We apply the whole-body metric defined in Sec. 3.2 to every tracker. None of the trackers is trained or fine-tuned on HumanTracker. Every benchmark number is computed on the HumanTracker test split and is reported separately for Daily, Highly Dynamic, Interaction and Ground. Metrics. We report three complementary measures of tracking quality. Succ captures catastrophic loss of tracking but cannot distinguish the quality of two completed rollouts. MPJPE measures local refer- ence fidelity but averages over time and joints. Hu- manScore measures trajectory-level preference and is designed to respond to contact, stability and smooth- ness as well as pose. We evaluate HumanScore in consecutive 5-second windows and retain a final shorter window using right-zero padding and a valid- ity mask. For the preference study, we additionally compare joint-velocity error, keypoint-position error, foot-contact agreement, and mean joint acceleration. 4.2 Results Table 3 shows that Humanoid-GPT is the strongest overall tracker, leading most comparisons and all three metrics on Daily and Highly Dynamic. SONIC is the closest competitor. It achieves the high- est completion rate on Interaction and the highest HumanScore on Ground. These exceptions reveal a meaningful difference between the two trackers. Humanoid-GPT generally completes motions more re- liably and follows the reference more closely, whereas Daily Highly Dynamic InteractionGround Align Rate Baseline + Future Reference − Contact Features Kinematics Only 0.8810.9290.9440.879 0.908 0.8930.9170.9310.879 0.905 0.8980.8870.9240.759 0.867 0.8920.9290.9160.776 0.878 (a) Models trained with different input features. 12345 Available temporal context (s) 0.84 0.86 0.88 0.90 0.92 Align Rate 0.857 0.876 0.906 0.900 0.908 (b) Baseline with restricted temporal context. Figure 5 HumanScore sensitivity. Align Rate is computed within each motion family and then averaged across 4 families. Panel a evaluates different input feature sets. Panel b restricts the temporal prefix available to the baseline at inference. Table 4 Alignment with human preferences. Align Rate is first computed within each motion family on strict test samples and then averaged across Daily, Highly Dy- namic, Interaction and Ground. Each family therefore contributes one quarter of the reported value. MetricAlign Rate HumanScore0.9083 MPJPE (rad)0.8049 MPJVE (rad/s)0.8404 KPT Position MAE (m)0.8405 Foot Contact Accuracy0.7882 Avg Joint Accel (rad/s 2 )0.6933 Avg Joint Jerk (rad/s 3 )0.7232 SONIC produces Ground rollouts that are perceived as more natural and stable. The family-level results therefore distinguish consistent overall tracking from strengths that emerge under a particular evaluation criterion. Moreover, it also shows that traditional metrics are not always aligned with HumanScore. 4.3 Human preference comparison Table 4 shows that HumanScore agrees with human preferences more consistently than any individual an- alytic diagnostic. Conventional metrics isolate pose, velocity or contact fidelity, whereas annotators also consider human-like, smoothness, stability and how errors develop over time. Their lower agreement in- dicates that no single diagnostic captures the full trajectory quality reflected in human judgments. The comparison also rules out several distributional shortcuts. Grouping the test set by source motion prevents adjacent clips from the same sequence from crossing partitions. Averaging agreement within each family prevents the more frequent Daily and Interac- tion samples from obscuring performance on Ground and Highly Dynamic. Tracker pairings are balanced by construction, so the result cannot be explained by a dominant family or tracker matchup. 4.4 Sensitivity Analysis Figure 5 shows that removing measured contact fea- tures degrades performance most clearly on Ground. This indicates that contact information is important for motions with complex contact transitions. Adding future reference information performs slightly worse than the baseline, suggesting that this extra signal offers limited benefit and is difficult to exploit, it is sufficient to rely on the current and historical states to evaluate the quality of the action. Longer context improves alignment by revealing sliding, jitter, drift and recovery that isolated poses cannot capture. Alignment also improves steadily as the available con- text grows from one to five seconds. Short segments capture instantaneous pose errors, whereas longer context reveals evolving artifacts such as foot sliding, repeated jitter, progressive drift, and recovery from instability. HumanScore therefore benefits from in- tegrating complementary evidence over time, rather than evaluating motion frame by frame. 5 Conclusion We introduce a large-scale benchmark and Hu- manScore, a human-aligned metric, for evaluating humanoid motion tracking. Built from optical motion recorded from 24 professional performers, the released benchmark contains approximately 153 hours across four motion families. It reveals persistent weaknesses in highly dynamic and ground-contact motions. Hu- manTracker provides a reproducible framework for tracker comparison and failure analysis, with future work extending to cross-embodiment settings, real- world hardware, and reward optimization. References [1]Joao Pedro Araujo, Yanjie Ze, Pei Xu, Jiajun Wu, and C Karen Liu. Retargeting matters: General motion retargeting for humanoid motion tracking. ArXiv preprint, abs/2510.02252, 2025.https://ar xiv.org/abs/2510.02252. [2]Léore Bensabath, Mathis Petrovich, and Gül Varol. A cross-dataset study for text-based 3d human motion retrieval. volume abs/2405.16909, 2024. https://arxiv.org/abs/2405.16909. [3] Ralph Allan Bradley and Milton E Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345, 1952. doi: 10.1093/biomet/39.3-4.324. [4] Zixuan Chen, Mazeyu Ji, Xuxin Cheng, Xuanbin Peng, Xue Bin Peng, and Xiaolong Wang. Gmt: General motion tracking for humanoid whole-body control. ArXiv preprint, abs/2506.14770, 2025.ht tps://arxiv.org/abs/2506.14770. [5]Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep re- inforcement learning from human preferences. In Isabelle Guyon, Ulrike von Luxburg, Samy Ben- gio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Ad- vances in Neural Information Processing Systems 30: Annual Conference on Neural Information Pro- cessing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 4299–4307, 2017.https: //proceedings.neurips.c/paper/2017/hash/d5e 2c0adad503c91f91df240d0cd4e49-Abstract.html. [6] Pranay Dugar, Aayam Shrestha, Fangzhou Yu, Bart van Marum, and Alan Fern. Learning multi-modal whole-body control for real-world humanoid robots. In Proceedings of the AAAI Symposium Series, vol- ume 7, pages 650–657, 2025. [7]Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wet- zstein, and Chelsea Finn. Humanplus: Humanoid shadowing and imitation from humans. 270:2828– 2844, 2024. [8] Yu Guan, Zekun Qi, Chenghuai Lin, Xuchuan Chen, Dairu Liu, Wenyao Zhang, Jilong Wang, Xinqiang Yu, He Wang, and Li Yi. LIMMT: Less is more for motion tracking, 2026.https://arxiv.org/abs/26 06.06953. [9]Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoper- ation and learning. ArXiv preprint, abs/2406.08858, 2024. https://arxiv.org/abs/2406.08858. [10]Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleopera- tion. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8944–8951. IEEE, 2024. [11]Jeonghwan Kim, Wontaek Kim, Yidan Lu, Jin Cheng, Fatemeh Zargarbashi, Zicheng Zeng, Zekun Qi, Zhiyang Dou, Nitish Sontakke, Donghoon Baek, et al. Switch-justdance: Benchmarking whole body motion tracking policies using a commercial console game. arXiv preprint arXiv:2511.17925, 2025. [12]Kyungmin Lee, Sibeen Kim, Minho Park, Hyunse- ung Kim, Dongyoon Hwang, Hojoon Lee, and Jaegul Choo. Phuma: Physically-grounded humanoid lo- comotion dataset. ArXiv preprint, abs/2510.26236, 2025. https://arxiv.org/abs/2510.26236. [13]Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, and Chelsea Finn. Robore- ward: General-purpose vision-language reward mod- els for robotics. ArXiv preprint, abs/2601.00675, 2026. https://arxiv.org/abs/2601.00675. [14]Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis. ACM Trans- actions on Graphics (TOG), 42(6):1–11, 2023. [15] Mingzhe Li, Mengyin Liu, Zekai Wu, Xincheng Lin, Junsheng Zhang, Ming Yan, Zengye Xie, Chang- wang Zhang, Chenglu Wen, Lan Xu, Siqi Shen, and Cheng Wang. Towards motion turing test: Evalu- ating human-likeness in humanoid robots. ArXiv preprint, abs/2603.06181, 2026.https://arxiv.or g/abs/2603.06181. [16]Yixuan Li, Yutang Lin, Jieming Cui, Tengyu Liu, Wei Liang, Yixin Zhu, and Siyuan Huang. Clone: Closed-loop whole-body humanoid teleoperation for long-horizon tasks. ArXiv preprint, abs/2506.08931, 2025. https://arxiv.org/abs/2506.08931. [17] Yuhan Li, Peiyuan Zhi, Yunshen Wang, Tengyu Liu, Sixu Yan, Wenyu Liu, Xinggang Wang, Baox- iong Jia, and Siyuan Huang. Omnitrack: Gen- eral motion tracking via physics-consistent refer- ence. ArXiv preprint, abs/2602.23832, 2026.https: //arxiv.org/abs/2602.23832. [18]Anthony Liang, Yigit Korkmaz, Jiahui Zhang, Miny- oung Hwang, Abrar Anwar, Sidhant Kaushik, Aditya Shah, Alex S. Huang, Luke Zettlemoyer, Dieter Fox, Yu Xiang, Anqi Li, Andreea Bobu, Abhishek Gupta, Stephen Tu, Erdem Biyik, and Jesse Zhang. Robometer: Scaling general-purpose robotic reward models via trajectory comparisons. ArXiv preprint, abs/2603.02115, 2026.https://arxiv.org/abs/26 03.02115. [19] Qiayuan Liao, Takara E Truong, Xiaoyu Huang, Guy Tevet, Koushil Sreenath, and C Karen Liu. Beyondmimic: From motion tracking to versatile hu- manoid control via guided diffusion. ArXiv preprint, abs/2508.08241, 2025.https://arxiv.org/abs/25 08.08241. [20] Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. In Alice Oh, Tristan Nau- mann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural In- formation Processing Systems 36: Annual Confer- ence on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023.http://papers.nips.c/paper _files/paper/2023/hash/4f8e27f6036c1d8b4a66b 5b3a947d7b-Abstract-Datasets_and_Benchmar ks.html. [21] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi-person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023. [22]Zhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani, and Weipeng Xu. Perpetual hu- manoid control for real-time simulated avatars. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 10861–10870. IEEE, 2023. doi: 10.1109/ICCV 51070.2023.01000.https://doi.org/10.1109/ICCV 51070.2023.01000. [23]Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris M. Kitani, and Weipeng Xu. Universal humanoid motion representations for physics-based control. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024.https://openreview.net/forum?id=OrOd8P xOO2. [24] Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Sirui Chen, Fernando Castañeda, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, et al. Sonic: Supersizing motion tracking for natural humanoid whole-body control. ArXiv preprint, abs/2511.07820, 2025. https://arxiv.org/abs/2511.07820. [25]Yecheng Jason Ma, Vikash Kumar, Amy Zhang, Os- bert Bastani, and Dinesh Jayaraman. LIV: Language- image representations and rewards for robotic con- trol. In Proceedings of the 40th International Confer- ence on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 23301–23320. PMLR, 2023.https://proceedings.mlr.press/v2 02/ma23b.html. [26] Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: archive of motion capture as surface shapes. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019, pages 5441–5450. IEEE, 2019. doi: 10.1109/ICCV.2019.00554. https://doi.org/10.1109/ICCV.2019.00554. [27]Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Pe- ter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow in- structions with human feedback. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neu- ral Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - De- cember 9, 2022, 2022.http://papers.nips.c/pap er_files/paper/2022/hash/b1efde53be364a739 14f58805a001731-Abstract-Conference.html. [28] Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel Van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based charac- ter skills. ACM Transactions On Graphics (TOG), 37(4):1–14, 2018. [29]Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin, Yunrui Lian, Sikai Liang, Zhikai Zhang, Yu Guan, Ji- long Wang, Wenyao Zhang, Xinqiang Yu, He Wang, and Li Yi. Humanoid-gpt: Scaling data and struc- ture for zero-shot motion tracking. ArXiv preprint, abs/2606.03985, 2026.https://arxiv.org/abs/26 06.03985. [30]Ilija Radosavovic, Bike Zhang, Baifeng Shi, Jathushan Rajasegaran, Sarthak Kamat, Trevor Dar- rell, Koushil Sreenath, and Jitendra Malik. Hu- manoid locomotion as next token prediction. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural In- formation Processing Systems 38: Annual Confer- ence on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024.http://papers.nips.c/paper _files/paper/2024/hash/90afd20dc776bc8849c31 d61a0763a0b-Abstract-Conference.html. [31] Jenny Sheng, Matthieu Lin, Andrew Zhao, Kevin Pruvost, Yu-Hui Wen, Yangguang Li, Gao Huang, and Yong-Jin Liu. Exploring text-to-motion gen- eration with human preference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 1888–1899, 2024.https://openaccess.thecvf.com/content/ CVPR2024W/HuMoGen/html/Sheng_Exploring_Tex t-to-Motion_Generation_with_Human_Preferen ce_CVPRW_2024_paper.html. [32]Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J. Black. WHAM: reconstructing world-grounded hu- mans with accurate 3d motion. In IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 2070–2080. IEEE, 2024. doi: 10.1109/CVPR 52733.2024.00202.https://doi.org/10.1109/CVPR 52733.2024.00202. [33] Sumedh A. Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch, Erdem Biyik, Dorsa Sadigh, Chelsea Finn, and Laurent Itti. Roboclip: One demonstration is enough to learn robot policies. In Advances in Neural Information Processing Systems, volume 36, pages 55681–55693, 2023. doi: 10.522 02/075280-2430.https://papers.nips.c/paper _files/paper/2023/hash/ae54ce310476218f26d4 8c1626d5187-Abstract-Conference.html. [34] Huajie Tan, Sixiang Chen, Yijie Xu, Zixiao Wang, Yuheng Ji, Cheng Chi, Yaoxu Lyu, Zhongxia Zhao, Xiansheng Chen, Peterson Co, Shaoxuan Xie, Guo- cai Yao, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang. Robo-dopamine: General pro- cess reward modeling for high-precision robotic ma- nipulation. ArXiv preprint, abs/2512.23703, 2025. https://arxiv.org/abs/2512.23703. [35]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Ad- vances in Neural Information Processing Systems 30: Annual Conference on Neural Information Pro- cessing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.https: //proceedings.neurips.c/paper/2017/hash/3f5 e243547dee91fbd053c1c4a845a-Abstract.html. [36]Haoru Wang, Wentao Zhu, Luyi Miao, Yishu Xu, Feng Gao, Qi Tian, and Yizhou Wang. Aligning human motion generation with human perceptions. In International Conference on Learning Represen- tations, 2025.https://proceedings.iclr.c/pape r_files/paper/2025/hash/c129741a2451e5fefe 447591e39de30e-Abstract-Conference.html. [37] Weiji Xie, Jinrui Han, Jiakun Zheng, Huanyu Li, Xinzhe Liu, Jiyuan Shi, Weinan Zhang, Chenjia Bai, and Xuelong Li. Kungfubot: Physics-based humanoid whole-body control for learning highly- dynamic skills. ArXiv preprint, abs/2506.12851, 2025. https://arxiv.org/abs/2506.12851. [38]Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. Iterative preference learning from human feedback: Bridging theory and practice for RLHF under kl- constraint. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024.https: //openreview.net/forum?id=c1AKcA6ry1. [39] Han Xue, Sikai Liang, Zhikai Zhang, Zicheng Zeng, Yun Liu, Yunrui Lian, Jilong Wang, Qingtao Liu, Xuesong Shi, and Li Yi. Collision-free humanoid traversal in cluttered indoor scenes, 2026.https: //arxiv.org/abs/2601.16035. [40] Yashuai Yan, Esteve Valls Mascaro, Tobias Egle, and Dongheui Lee. I-ctrl: Imitation to control humanoid robots through constrained reinforcement learning. ArXiv preprint, abs/2405.08726, 2024.https://ar xiv.org/abs/2405.08726. [41] Lujie Yang, Xiaoyu Huang, Zhen Wu, Angjoo Kanazawa, Pieter Abbeel, Carmelo Sferrazza, C Karen Liu, Rocky Duan, and Guanya Shi. Omnire- target: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction. ArXiv preprint, abs/2509.26633, 2025. https://arxiv.org/abs/2509.26633. [42] Kangning Yin, Weishuai Zeng, Ke Fan, Minyue Dai, Zirui Wang, Qiang Zhang, Zheng Tian, Jingbo Wang, Jiangmiao Pang, and Weinan Zhang. Unitracker: Learning universal whole-body motion tracker for humanoid robots. ArXiv preprint, abs/2507.07356, 2025. https://arxiv.org/abs/2507.07356. [43] Yanjie Ze, Zixuan Chen, JoÃG , o Pedro AraÚjo, Zi- ang Cao, Xue Bin Peng, Jiajun Wu, and C Karen Liu. Twist: Teleoperated whole-body imitation system. ArXiv preprint, abs/2505.02833, 2025.https://ar xiv.org/abs/2505.02833. [44]Yanjie Ze, Siheng Zhao, Weizhuo Wang, Angjoo Kanazawa, Rocky Duan, Pieter Abbeel, Guanya Shi, Jiajun Wu, and C Karen Liu. TWIST2: Scal- able, portable, and holistic humanoid data collec- tion system. ArXiv preprint, abs/2511.02832, 2025. https://arxiv.org/abs/2511.02832. [45] Siwei Zhang, Bharat Lal Bhatnagar, Yuanlu Xu, Alexander Winkler, Petr Kadlecek, Siyu Tang, and Federica Bogo. Rohm: Robust human motion re- construction via diffusion. In IEEE/CVF Con- ference on Computer Vision and Pattern Recog- nition, CVPR 2024, Seattle, WA, USA, June 16- 22, 2024, pages 14606–14617. IEEE, 2024. doi: 10.1109/CVPR52733.2024.01384.https://doi. org/10.1109/CVPR52733.2024.01384. [46] Yuhong Zhang, Jing Lin, Ailing Zeng, Guanlin Wu, Shunlin Lu, Yurong Fu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x++: A large-scale multimodal 3d whole-body human motion dataset. ArXiv preprint, abs/2501.05098, 2025.ht tps://arxiv.org/abs/2501.05098. [47]Yuxiang Zhang, Hongwen Zhang, Liangxiao Hu, Ji- ajun Zhang, Hongwei Yi, Shengping Zhang, and Yebin Liu. Proxycap: Real-time monocular full-body capture in world space via human-centric proxy-to- motion learning. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 1954– 1964. IEEE, 2024. doi: 10.1109/CVPR52733.2024.0 0191.https://doi.org/10.1109/CVPR52733.2024 .00191. [48]Zhikai Zhang, Yitang Li, Haofeng Huang, Mingx- ian Lin, and Li Yi. Freemotion: Mocap-free human motion synthesis with multimodal large language models. In European Conference on Computer Vi- sion, pages 403–421. Springer, 2024. [49]Zhikai Zhang, Jun Guo, Chao Chen, Jilong Wang, Chenghuai Lin, Yunrui Lian, Han Xue, Zhenrong Wang, Maoqi Liu, Jiangran Lyu, et al. Track any motions under any disturbances. ArXiv preprint, abs/2509.13833, 2025.https://arxiv.org/abs/25 09.13833. [50] Zhikai Zhang, Haofei Lu, Yunrui Lian, Ziqing Chen, Yun Liu, Chenghuai Lin, Han Xue, Zicheng Zeng, Zekun Qi, Shaolin Zheng, et al. Learning athletic humanoid tennis skills from imperfect human motion data. arXiv preprint arXiv:2603.12686, 2026. [51]Ziyu Zhang, Sergey Bashkirov, Dun Yang, Yi Shi, Michael Taylor, and Xue Bin Peng. Physics-based motion imitation with adversarial differential dis- criminators. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers, pages 1–12. ACM, 2025. doi: 10.1145/3757377.3763819.https: //doi.org/10.1145/3757377.3763819. [52]Siheng Zhao, Yanjie Ze, Yue Wang, C Karen Liu, Pieter Abbeel, Guanya Shi, and Rocky Duan. Resmimic: From general motion tracking to hu- manoid whole-body loco-manipulation via residual learning. ArXiv preprint, abs/2510.05070, 2025. https://arxiv.org/abs/2510.05070. A Implementation details A.1 Segment construction and padding Following Sec. 3.3, a segment withL <250 frames is padded with zeros on the right to form a 250×539 input. A Boolean validity mask marks the firstL frames. The Transformer uses this mask to exclude padding from attention, and temporal mean pooling uses the same mask to exclude padding from the trajectory representation. For encoded tokenh t and feature index j, the pooled feature is ̄ h j = P 250 t=1 m t h t,j P 250 t=1 m t , wherem t ∈ 0,1. Short terminal segments can therefore contribute without padded frames affecting attention or pooling. When window rewards are combined across a rollout, each window is weighted by its number of actual frames. Padded positions contribute neither to the window representation nor to the trajectory average. A.2 Frame features Table 5 decomposes the reported 539-dimensional frame token into 70 current reference dimensions and 469 rollout dimensions. The rollout features are com- puted from the simulated robot. Foot contact and force are obtained from contacts between the robot and floor, while foot acceleration is the temporal derivative of measured foot velocity. Keypoint fea- tures describe 14 bodies in a navigation frame aligned with gravity. Table 5 HumanScore input features. Dimensions are re- ported for one frame. GroupContentsDim. Current refer- ence root pose and navigation ve- locity; joint position and ve- locity; foot contact 70 Robot state and action root and IMU pose; action; motor target; joint position and velocity 126 Measured con- tact dynamics foot contact, force, velocity and acceleration 20 Root motionpelvis and root velocities in local and navigation frames 15 Current key- points 14 4×4 poses and six dimen- sional spatial velocities 308 Totalconcatenated frame token539 The model without measured contact features re- moves the 20-dimensional rollout contact block and has 519 input dimensions. The Kinematics Only vari- ant also removes the three-dimensional gravity frame angular velocity and has 516 dimensions. Both vari- ants retain the current reference foot contact target. A.3 Reward model optimization The reported model uses a bidirectional Transformer with normalization before each sublayer and a reward head comprising three linear layers. Table 6 lists the architecture and optimization settings recovered from the reported checkpoint. At inference, a sigmoid maps each unbounded window reward to a bounded reward between zero and one. HumanScore is 100 times their mean weighted by the number of actual frames in each window. Table 6 Settings of the reported HumanScore checkpoint. HyperparameterValue Model dimension256 Transformer layers4 Attention heads8 FFN dimension1024 Poolingmasked mean Maximum sequence length 250 frames Temporal downsampling none Batch size8 OptimizerAdamW Learning rate 1× 10 −4 Training epochs20 Learning rate schedulecosine with 10% warmup Maximum gradient norm 1.0 Dropout0.1 Weight decay 1× 10 −5 Preference temperature1.0 Similar pair weight1.0 Training precisionfloat32 Random seed42 Padding right zero padding with a validity mask B Additional related work B.1 Evaluation from human preferences Pairwise judgments are useful when people can recog- nize the quality of a structured output more readily than researchers can express it as a fixed analytic objective. Learning from human preferences [5] es- tablished this approach for reward learning, while InstructGPT [27] and iterative RLHF [38] scaled it for language model alignment. In human motion, InstructMotion [31] uses comparisons to refine mo- tion generation from text. MotionCritic [36] learns a quality metric from MotionPercept, and the Mo- tion Turing Test [15] measures whether humanoid motion appears human from kinematic observations. These methods evaluate generated motion or general human likeness. Our setting instead compares the physical quality of robot rollouts that follow the same reference. Robotic reward learning provides a second line of related work. LIV [25] learns dense rewards from videos without actions, and RoboCLIP [33] derives rewards from video or text demonstrations. RoboRe- ward [13] trains vision and language reward models on large robot datasets. Robo-Dopamine [34] mod- els manipulation progress from multiple views, while Robometer [18] combines progress supervision with trajectory comparisons. These methods primarily support task completion or policy optimization. Hu- manScore instead evaluates humanoid motion track- ing by integrating pose, contact, force and velocity evidence over synchronized rollouts of the same ref- erence. C Preference data statistics The preference catalogue uses only motions from the HumanTracker training set. We divide the aligned tracker rollouts for every source motion into consecu- tive clips and retain the final shorter clip. After sort- ing the complete catalogue by family, source motion and temporal position, we select uniformly spaced clips. A balanced schedule assigns one of the six tracker combinations to each selected clip, and the corresponding two rollouts form a comparison. Six doctoral researchers specializing in humanoid robotics annotated 6,000 original trajectory pairs. We bilat- erally mirrored each pair, yielding 12,000 preference records for model development. Mirrored variants remain in the same source-motion partition, and pref- erence alignment is evaluated only on the original unmirrored test comparisons. The label set contains strict preferences, Similar judg- ments and Cannot compare judgments. We exclude Cannot compare from reward model optimization. Records are split 80/20 by sourcemotion_idusing seed 42, so clips from one motion cannot cross par- titions. The split balances motion family, tracker pair, label, annotator, clip length and the number of sampled clips from each source motion. Alignment analysis uses strict test samples. Agreement is com- puted within each family before the four family rates are averaged with equal weight. Samples too short to define acceleration or jerk are omitted for those metrics. D Dataset scale and release format All released robot trajectories are stored at 50 Hz. The training manifest lists 22,495 trajectories and 24,793,129 frames, while the test manifest lists 2,500 trajectories and 2,687,461 frames. This gives a 9:1 split by trajectory count. Table 7 reports the com- bined scale of each motion family. Table 7 HumanTracker statistics by motion family. FamilyTrajectories Frames Duration (h) Daily9,739 16,072,017 89.29 Highly Dynamic 2,6761,981,84311.01 Ground1,640825,7924.59 Interaction10,940 8,600,93847.78 Total24,995 27,480,590 152.67 Thetrain.jsonandtest.jsonmanifests define the released partitions. All trajectories derived from the same source motion remain in one partition, prevent- ing related entries from crossing between training and test data. Each manifest entry records its path, motion family and frame count. Each NPZ archive contains seven arrays with a com- mon temporal lengthF. The robot state arrays are qposof shapeF×36 andqvelof shapeF×35, both infloat32. Keypoint motion is stored askpt2gv_- poseof shapeF ×14×4×4 infloat32andkpt_- cvel_in_gv of shapeF ×14×6 infloat64. Global motion usesgv_velof shapeF×3 andgv2wrd_pose of shapeF ×4×4, both infloat32. The Boolean arrayfoot_contacthas shapeF ×2. A complete scan confirmed this schema across all 24,995 released archives. E Analysis of motion coverage We characterize each motion family using four descrip- tors computed from one source reference trajectory per motion. Horizontal root path is the accumulated root translation in the ground plane. Root height range is the difference between the maximum and minimum root height. For joint speed, we take finite differences of the 29 actuated joint coordinates at 50 Hz and then compute the 95th percentile of abso- lute speed across frames and joints. Table 8 reports the median, first quartile and third quartile of each descriptor across source motions. Daily motions are longer and cover more horizontal distance, whereas Interaction motions are spatially localized. Ground spans a broader upper half of the Table 8 Kinematic coverage of HumanTracker reference motions. Each entry reports median [first quartile, third quartile] across source motions. FamilyDuration (s)Horizontal root path (m) Root height range (m) Joint speed P95 (rad/s) Daily30.81 [27.62, 33.68] 15.25 [10.03, 20.99]0.273 [0.108, 0.428] 2.316 [1.865, 2.770] Highly Dynamic 8.38 [6.70, 11.17]1.40 [0.63, 2.69]0.081 [0.045, 0.114] 1.657 [1.221, 2.051] Interaction10.24 [8.58, 11.70]0.42 [0.29, 1.16]0.008 [0.005, 0.045] 0.962 [0.626, 1.157] Ground9.97 [9.14, 11.14]0.66 [0.14, 1.20]0.102 [0.087, 0.620] 1.079 [0.843, 1.255] root height distribution. Highly Dynamic contains many short skills, so its label does not imply that every kinematic descriptor exceeds those of other families. These descriptors do not measure contact timing, support or recovery and therefore cannot replace Succ, MPJPE or HumanScore. F Analysis of preference alignment Table 9 supplements Table 4 with uncertainty for Align Rate. We use 20,000 bootstrap draws with seed 20260813. Within each family, source motions are sampled with replacement and all associated com- parisons are retained. We then recompute the four family rates and their equal-weight mean. This pro- cedure preserves the reported family weighting while accounting for repeated comparisons from the same source motion. Table 9 Uncertainty of preference alignment. We report 95% intervals from a bootstrap stratified by motion family and clustered by source motion. MetricAlign Rate (%)95% CI (%) HumanScore90.83 [87.36, 93.83] MPJPE80.49 [75.95, 84.76] MPJVE84.04 [79.80, 87.87] KPT Position MAE84.05 [79.67, 88.04] Foot Contact Accuracy78.82 [73.73, 83.59] Avg Joint Accel69.33 [64.23, 74.12] Avg Joint Jerk72.32 [67.52, 76.93] G Discussion and Limitations The benchmark and preference results point to the same conclusion from different directions. Category- level evaluation shows that contact regime changes the relative behaviour of trackers: a method that is reliable on upright daily motion can still fail almost completely during ground-level transitions. The pref- erence experiment explains why this distinction is not fully represented by joint error. HumanScore gains most of its advantage when it can observe contact- related state over several seconds, suggesting that the perceptual unit of failure is often an event such as a slide, impact, support switch or recovery, rather than an isolated pose. HumanScore should therefore be read alongside success and kinematic error, not as a replacement for all analytic diagnostics: the former summarizes perceived trajectory quality, whereas the latter remain valuable for locating a specific source of error. Several boundaries remain. HumanScore is trained from HumanTracker training motions and rollouts produced by four trackers, so its motion-disjoint test measures generalization to unseen motions, not to an entirely unseen robot, simulator or controller family. Its 539-dimensional input also includes privileged sim- ulator state and contact quantities that are available for benchmarking but may not be observable on hard- ware; applying the metric to real-world trajectories will require an observable feature set or a separately validated state estimator. The preference pool assigns one primary judgment to each pair, which provides broad coverage but does not quantify uncertainty through repeated independent labels. Finally, the present metric comparison covers representative indi- vidual diagnostics; broader tests against fitted linear and nonlinear diagnostic composites would further delimit what must be learned from trajectory-level preference data. HumanTracker also inherits the limits of its data distribution. The four families are deliberately diag- nostic but imbalanced, with far fewer Ground clips than Daily or Interaction clips, and the evaluation uses one 29-DoF humanoid embodiment in MuJoCo. Results should not be interpreted as a population estimate over all human activity or as evidence of hardware robustness. Extending the benchmark to additional embodiments, real-robot rollouts and rare contact regimes is therefore a natural next step. We also restrict HumanScore to evaluation in this work. Directly optimizing a learned score can create be- haviours that exploit model imperfections; using it as a reinforcement-learning reward will require explicit regularization and an independent human evalua- tion rather than assuming that benchmark alignment transfers automatically to policy optimization.