Paper deep dive
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction
Ayda Eghbalian, Kevin Desai
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 8:15:18 AM
Summary
The paper introduces BioModule, a lightweight plug-in temporal transformer that bridges 3D human pose estimation and biomechanical attribute prediction. It takes standard 17-joint 3D skeletons from any pose estimator and predicts 17 biomechanical attributes across kinematic, kinetic, and neuromuscular tiers. The authors construct Human3.6Mplus, a dataset aligning Human3.6M video/keypoints with Human3.6Mplus musculoskeletal simulation labels, enabling frame-accurate cross-modal supervision. BioModule is evaluated across seven state-of-the-art pose estimators to analyze how upstream pose quality affects downstream biomechanical fidelity.
Entities (8)
Relation Signals (6)
BioModule → predicts → Biomechanical Attributes
confidence 95% · BioModule... predicts biomechanical attributes from standard 17-joint 3D skeletons.
BioModule → uses → Temporal Transformer
confidence 95% · BioModule is a temporal transformer module that maps a sequence of root centred 3D skeletons to biomechanical attributes.
Human3.6M → alignedwith → Human3.6Mplus
confidence 90% · We establish and verify anatomical correspondence between coordinate systems of the two datasets... enabling frame-accurate cross-modal supervision.
BioModule → trainson → Human3.6Mplus
confidence 90% · To train and evaluate BioModule, we construct a large-scale aligned dataset pairing Human3.6M video and 3D keypoints with the biomechanical label space of Human3.6Mplus.
BioModule → benchmarkedon → 3D Human Pose Estimation
confidence 85% · We further benchmark BioModule across seven state-of-the-art 3D pose estimators...
OpenCap → uses → 3D Human Pose Estimation
confidence 80% · OpenCap uses videos from smartphones to estimate human movement kinematics and dynamics through a pipeline that combines pose estimation...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable. However, most pose estimators remain optimized for geometric keypoint accuracy, while many real-world applications in rehabilitation, sports science, ergonomics, and clinical movement analysis require biomechanical quantities that describe how the body moves, loads, and activates. In this work, we propose BioModule, a lightweight plug-in temporal transformer that attaches downstream of any 3D pose estimator and predicts biomechanical attributes from standard 17-joint 3D skeletons. BioModule is estimator-agnostic and requires no modification of the upstream pose model, enabling existing pose estimators to be extended toward physically interpretable motion analysis. To train and evaluate BioModule, we construct a large-scale aligned dataset pairing Human3.6M video and 3D keypoints with the biomechanical label space of Human3.6Mplus. We establish and verify anatomical correspondence between coordinate systems of the two datasets, enabling frame-accurate cross-modal supervision. Using this aligned supervision, BioModule predicts biomechanical quantities. We further benchmark BioModule across seven state-of-the-art 3D pose estimators, providing the first systematic analysis of how upstream pose estimation quality propagates to downstream biomechanical prediction fidelity. The results position BioModule as a compact, modular bridge between vision-based pose estimation and biomechanically meaningful human motion analysis.
Tags
Links
- Source: https://arxiv.org/abs/2607.08725v1
- Canonical: https://arxiv.org/abs/2607.08725v1
Trouble viewing inline? Open PDF directly →
Full Text
60,563 characters extracted from source content.
Expand or collapse full text
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction Ayda Eghbalian and Kevin Desai Department of Computer Science, University of Texas at San Antonio, One UTSA Circle, San Antonio, 78249, Texas, United States. Contributing authors: ayda.eghbalian@utsa.edu; kevin.desai@utsa.edu; Abstract Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable. However, most pose estima- tors remain optimized for geometric keypoint accuracy, while many real-world applications in rehabilitation, sports science, ergonomics, and clinical move- ment analysis require biomechanical quantities that describe how the body moves, loads, and activates. In this work, we propose BioModule, a lightweight plug-in temporal transformer that attaches downstream of any 3D pose estima- tor and predicts biomechanical attributes from standard 17-joint 3D skeletons. BioModule is estimator-agnostic and requires no modification of the upstream pose model, enabling existing pose estimators to be extended toward physically interpretable motion analysis. To train and evaluate BioModule, we construct a large-scale aligned dataset pair- ing Human3.6M video and 3D keypoints with the biomechanical label space of Human3.6Mplus. We establish and verify anatomical correspondence between coordinate systems of the two datasets, enabling frame-accurate cross-modal supervision. Using this aligned supervision, BioModule predicts biomechanical quantities. We further benchmark BioModule across seven state-of-the-art 3D pose estimators, providing the first systematic analysis of how upstream pose estimation quality propagates to downstream biomechanical prediction fidelity. The results position BioModule as a compact, modular bridge between vision- based pose estimation and biomechanically meaningful human motion analysis. The complete source code and additional qualitative results are available at:https: //utsa-virlab.github.io/BioModule/ Keywords: Human Pose Estimation, Vision-based biomechanics, Markerless biomechanics, Musculoskeletal model 1 arXiv:2607.08725v1 [cs.CV] 9 Jul 2026 1 Introduction 3D human pose estimation models have become increasingly effective at recovering geometric skeletons from images and videos, yet a gap remains between kinematic pose and the biomechanical quantities required for physically meaningful motion analysis. Existing pose estimation benchmarks commonly optimize for Mean Per Joint Position Error (MPJPE) and related geometric metrics without quantifying physiological cor- rectness, including torques, ground reaction forces, or muscle activation signals, which are essential assets in rehabilitation, and clinical movement assessment.[1–6]. Bridging this gap has traditionally required marker based motion capture lab- oratories, force plates, and electromyography, instrumentation that is expensive, environment constrained, and difficult to scale to in the wild video. Recent markerless pipelines such as OpenCap [7] and BioPose [8] have demonstrated that video derived kinematics can seed inverse dynamics solvers, but they either require multi view cal- ibrated capture or per subject optimization at inference, limiting their scalability. Meanwhile, body model approaches improve anatomical realism but do not expose the full musculoskeletal output space needed for biomechanical analysis [9, 10]. The result is that no existing method enables conversion or integration of the 3D skeleton output of standard pose estimations with a comprehensive biomechanical attributes without additional instrumentation. In this paper, we propose BioModule, a plug-in temporal transformer that attaches downstream of any 3D pose estimator and simultaneously predicts 17 biome- chanical criteria across three tiers: kinematic, kinetic, and neuromuscular. These criteria are inferred from a short temporal receptive field of root centered 3D joint positions. The key insight is that a shared temporal encoder, trained on biomechani- cally labeled skeleton sequences, can learn the implicit mapping from pose dynamics to musculoskeletal state without physics simulation at inference. BioModule introduces no changes to the upstream estimator and requires only a temporally ordered 17-joint 3D skeleton as input, making it compatible with the full spectrum of contemporary lifting models. To support training and evaluation at scale, we construct Human3.6Mplus, a large- scale dataset that aligns Human3.6M [1] keypoints with musculoskeletal marker-based simulation labels from Human3.6Mplus [11]. A core contribution of this dataset is the establishment and geometric verification of anatomical correspondence between the coordinate systems of the two modalities, anchored at the shared pelvis root, enabling frame accurate cross-modal supervision across all frames of the seven subjects and thirty activities included in Human3.6M. Finally, we conduct a systematic cross-estimator benchmarking of downstream biomechanical prediction quality. Training BioModule once on ground-truth poses and evaluating it, both frozen and after lightweight fine-tuning, across seven state- of-the-art 3D pose estimators, we quantify precisely how upstream pose accuracy propagates to each biomechanical tier. This analysis reveals which architectural fami- lies of pose estimator are most biomechanically faithful, providing actionable guidance for practitioners deploying markerless biomechanical pipelines. The contributions of this work are: 2 • BioModule: We have trained a plug-in temporal transformer with a tiered multi- head prediction architecture that regresses 17 biomechanical criteria, spanning joint coordinates, torques, ground reaction forces, and muscle activation signals, from any 3D pose estimator’s output. • Aligned pose and biomechanics dataset: We have curated a biomechanically annotated dataset pairing Human3.6M video and skeleton sequences with full mus- culoskeletal simulation labels, including geometric verification of the cross-modal anatomical correspondence required for frame accurate supervision. • Cross-estimator biomechanical benchmark: We have set up an evaluation pro- tocol across seven contemporary 3D pose estimators, providing the first quantitative analysis of how pose estimation architecture and accuracy determine downstream biomechanical prediction fidelity. 2 Related Work 2.1 Deep Learning-Based Human Pose Estimation Human pose estimation has progressed from direct coordinate regression in single images to increasingly structured models that exploit spatial, temporal, and kinematic priors. Early deep learning approaches such as DeepPose formulated pose estima- tion as direct regression from image evidence to body-joint locations [12]. Subsequent methods improved localization by using convolutional heatmap representations and multi-stage refinement, including convolutional pose machines [13], stacked hourglass networks [14], and high-resolution representations such as HRNet [15]. These methods established strong 2D pose detectors that later became the input backbone for many monocular 3D human pose estimation systems. A major line of 3D pose estimation first estimates 2D joints and then lifts them into 3D space. Martinez et al. showed that a simple fully connected residual network can be highly effective for 2D-to-3D pose lifting when accurate 2D detections are available [2]. VideoPose3D extended this formulation with temporal convolutional networks, using motion context across frame windows to improve robustness against depth ambiguity and 2D detection noise [3]. Large-scale motion-capture datasets such as Human3.6M have provided the standard benchmark for evaluating these models in controlled indoor settings [1]. More recently, transformer-based models have become prominent for video-based 3D pose estimation. PoseFormer introduced a spatial-temporal transformer for mod- eling joint relations within frames and temporal dependencies across frames [16], building on the general attention mechanism introduced in Transformer architec- tures [17]. Subsequent methods improved efficiency, ambiguity handling, and motion representation. MHFormer introduced a multi-hypothesis transformer to address monocular depth ambiguity [18], PoseFormerV2 used frequency-domain representa- tions for efficient temporal modeling [19], and MotionAGFormer combined graph- based skeletal structure with attention mechanisms [20]. More recent models further explore diffusion-based pose aggregation [21], kinematic and trajectory priors [22], implicit temporal pose proxies [23], state-space sequence modeling [24], and general 3 human motion representations [25]. Pre-layer normalization has also been shown to improve transformer training stability [26]. These methods show that temporal and structural modeling are essential for accu- rate 3D pose recovery. However, their objective remains primarily geometric: the output is a 3D skeleton or pose representation optimized by joint-position error. Even when such models implicitly encode motion dynamics, they do not directly supervise or evaluate biomechanical quantities such as torques, reaction forces, contact, acti- vation, or excitation. BioModule uses the 3D skeletons produced by such estimators as input, but evaluates them through biomechanical prediction fidelity rather than geometric accuracy alone. 2.2 Biomechanically Accurate Body Models and Motion Representations Beyond sparse keypoints, parametric body models provide richer representations of human shape and pose. SMPL introduced a learned skinned body model that rep- resents human body shape and articulation with a compact parameterization [27]. Image- and video-based mesh recovery methods such as HMR and VIBE estimate body pose and shape from monocular visual input, enabling temporally coherent reconstruc- tion of human motion in mesh space [28, 29]. These methods provide a more complete geometric representation than a sparse 3D skeleton and have become widely used in human motion analysis. Recent work has further sought to make body models more anatomically and biomechanically meaningful. Keller et al. introduced SKEL and BioAMASS to connect surface body models with a biomechanically grounded skeleton [9]. Xia et al. recon- structed humans with biomechanically accurate skeletons, further emphasizing the importance of anatomical structure in human reconstruction [10]. SKEL-CF extends this direction through coarse-to-fine recovery of biomechanical skeleton and surface mesh representations [30]. These works are important because they move beyond visual surface reconstruction toward body representations that are more consistent with human anatomy. However, body models and mesh recovery methods primarily address how the body is represented or reconstructed. They do not directly provide a general estimator- agnostic mechanism for converting the standard 17-joint outputs of existing 3D pose estimators into a broad biomechanical state space. BioModule is complementary to these approaches: rather than proposing a new body model, it learns to infer kinematic, kinetic, contact, and neuromuscular attributes from the sparse skeleton representation already produced by contemporary 3D pose estimators. 2.3 Markerless Biomechanics and Video-to-Biomechanics Pipelines Biomechanical movement analysis traditionally relies on marker-based motion cap- ture, force plates, and musculoskeletal modeling to estimate kinematics, kinetics, and muscle-related quantities. Recent markerless systems have attempted to reduce this dependence on laboratory instrumentation. OpenCap uses videos from smartphones to 4 estimate human movement kinematics and dynamics through a pipeline that combines pose estimation, musculoskeletal modeling, and simulation [7]. OpenCap Monocu- lar further extends this direction toward single-video biomechanical analysis [31]. OpenCapBench explicitly frames the gap between pose estimation and biomechan- ics by evaluating whether pose-estimation outputs preserve biomechanically relevant correctness, not only geometric accuracy [4]. Other video-to-biomechanics and markerless motion-capture studies examine dif- ferent parts of this pipeline. Cotton et al. studied trajectory optimization and inverse kinematics for biomechanical analysis of markerless motion-capture data [32]. Ruescas- Nicolau et al. investigated keypoint augmentation for markerless motion capture in biomechanical applications [33]. Auer et al. evaluated markerless motion capture com- bined with musculoskeletal models for kinematic analysis [34], while Barzyk et al. studied smartphone-based markerless capture of lower-limb joint angles during coun- termovement jumps [35]. Rode et al. assessed monocular human pose estimation models for clinical movement analysis [36]. Together, these studies show that mark- erless motion capture is increasingly relevant for biomechanics, rehabilitation, sports science, and clinical assessment. Recent methods also integrate vision models more directly with biomechanical con- straints. BioPose estimates biomechanically accurate 3D pose from monocular video by combining mesh recovery with biomechanical constraints [8]. Lin et al. used biome- chanical models and synthetic training data to estimate 3D kinematics from video [37]. Miller et al. showed that integrating machine learning with musculoskeletal simula- tion can improve OpenCap video-based dynamics estimation [38]. These approaches demonstrate the value of combining learned visual representations with biomechanical modeling. Most existing markerless biomechanics systems are designed as complete pipelines involving video processing, pose recovery, trajectory refinement, inverse kinematics, musculoskeletal modeling, or simulation. BioModule addresses a different setting: given the 17-joint 3D skeleton output of an existing pose estimator, it learns a compact temporal mapping to multiple biomechanical attributes. This makes the proposed framework suitable for comparing different upstream pose estimators under a shared biomechanical prediction interface. 2.4 Musculoskeletal Simulation, Physics-Informed Modeling, and Biomechanical Datasets Musculoskeletal simulation provides the physical foundation for estimating biome- chanical variables that are not directly visible from video. OpenSim is a widely used framework for creating and analyzing dynamic simulations of movement [5], and later extensions support musculoskeletal dynamics and neuromuscular control modeling for human and animal movement [6]. AddBiomechanics automates model scaling, inverse kinematics, and inverse dynamics from motion-capture data and musculoskeletal mod- els [39]. Its associated dataset captures the physics of human motion at scale, providing a broader source of motion and biomechanical supervision [40]. 5 Human3.6Mplus provides another important form of biomechanical supervision by pairing Human3.6M motion with physically consistent musculoskeletal labels, includ- ing kinematic, dynamic, and muscle-related quantities [11]. This type of dataset is critical for learning mappings from pose sequences to biomechanical attributes because quantities such as joint torques, reaction forces, activation, and excitation are not directly annotated in conventional computer-vision pose datasets. Related work has also explored physics-informed or physics-based learning for pose and dynamics estimation. IPMAN incorporates intuitive physics constraints, includ- ing floor contact and center-of-pressure or center-of-mass consistency, into 3D human pose estimation [41]. SSPINNpose uses a self-supervised physics-informed neural net- work for inertial pose and dynamics estimation [42], and OrientationNN provides a lightweight physics-informed approach for real-time joint kinematics estimation from IMU data [43]. These methods show that physical constraints can improve motion estimation, although they often target pose, inertial kinematics, or specific dynam- ics settings rather than a broad simulation-derived biomechanical label space from video-based 3D skeletons. Overall, musculoskeletal simulation and physics-informed learning provide the label sources and physical constraints needed for biomechanical inference. However, these tools are typically used either as explicit simulation pipelines or as constraints inside task-specific models. BioModule uses simulation-derived supervision differently: it learns a reusable temporal mapping from standard 17-joint 3D pose sequences to mul- tiple biomechanical criteria, enabling biomechanical prediction without running full musculoskeletal simulation at inference time. 3 Methodology 3.1 Overview We propose a pipeline that maps monocular RGB video to a biomechanical state repre- sentation of the human body through two decoupled stages, as shown in Figure 1. First, a 3D human pose estimator reconstructs a temporally ordered sequence of 3D skeletal joint positions from the input video. Second, BioModule receives the root-centred 3D pose sequence and predicts biomechanical attributes derived from the H3.6Mplus musculoskeletal simulation labels [11]. This separation allows BioModule to operate downstream of different 3D pose estimators without architectural changes. BioMod- ule is trained from scratch using ground-truth 3D poses and H3.6Mplus biomechanical labels. Let P t ∈R J×3 denote the 3D skeleton at frame t, where J =17 joints follow the Human3.6M joint convention. The input to BioModule is the described sequence of Euclidean 3D joint coordinates. To remove global translation, each frame is centered at the pelvis joint p t,0 ∈R 3 : ̄ P t = P t − 1p ⊤ t,0 (1) The centred pose is then flattened as x t = vec( ̄ P t )∈R 51 . The resulting input window is X = [x 1 ,...,x W ]∈R W×51 (2) 6 where W is the temporal receptive field. In this work, W =81, which corresponds to approximately 1.62 seconds at 50 fps. BioModule receives raw meter scale, root centered 3D joint coordinates. BioModule predicts C=17 biomechanical attributes ˆ Y a C a=1 , where ˆ Y a ∈ R B×W×d a and d a denotes the output dimension of attribute a. These attributes are organized into three groups according to their interpretation in biomechanical muscu- loskeletal analysis and their relationship to vision based pose estimation, as defined in Eq. 3. The kinematic attributes describe motion geometry and its temporal deriva- tives. In biomechanics, kinematics refers to motion quantities without directly considering the forces that caused the motion. In human pose estimation, this group is closest to the information explicitly represented by pose sequences, because 3D joint locations encode body configuration over time. In our output set, the kine- matic attributes are coordinates, speed, and acceleration. Here, there are 2 sets of coordinates: the output pose of H36M Euclidean joint positions, and the OpenSim generalized marker coordinates used in H3.6Mplus based on which the degrees of free- dom for each joint are defined. Therefore, the model must learn the mapping from Human3.6M joint positions to OpenSim generalized marker coordinates as well. This mapping depends on the OpenSim skeleton, joint degrees of freedom constraints, coordinate conventions, and the anatomical model used to generate the H3.6Mplus labels. The kinetic attributes describe force and load related quantities. In biome- chanics, kinetics refers to the quantities associated with producing or constraining movement. In pose estimation research, these quantities are usually not predicted directly, because standard pose benchmarks evaluate geometric joint accuracy rather than physical loading. In our output set, the kinetic attributes are active torque, passive torque, ideal torque, instantaneous power, instantaneous power raw, ground reaction, seat reaction, and touch. The touch attribute is binary and represents whether or not the right and left foot have contact with floor. The neuromuscular attributes describe quantities associated with excitation, activation, actuator scaling, and torque generation capacity which are derived from biomechanical simulations. The term neuromuscular in the sense is used in neuromus- culoskeletal modeling, where neural excitation and activation dynamics are linked to muscle force and joint torque generation [5, 6]. In our output set, the neuromuscu- lar attributes are activation signal, excitation signal, normalized active torque, angle scaling, velocity scaling, and maximum joint torque. This grouping is used both to interpret the predicted outputs and to define the weighted multitask objective and is further used in the equations 15, 16, 17, and 18. 7 Fig. 1 End to end vision to biomechanics pipeline and BioModule architecture. A monocular video sequence is first converted into a temporally ordered 3D kinematic pose sequence by a 3D human pose estimator. BioModule then receives the root centred pose sequences, embeds the frame vectors into a latent space, and processes the temporal window with multiple layers of transformer encoders. Inde- pendent prediction heads estimate kinematic, kinetic, and neuromuscular biomechanical attributes, which are optimized using a tiered multi task loss with weights 1.0, 0.5, and 0.3, respectively. A kin =coordinate, speed, acceleration, A knt =activetorque, passivetorque, idealtorque, instantaneouspower, instantaneouspowerraw, groundreaction, seatreaction, touch, A nmsc =activationsignal, excitationsignal, normalizedactivetorque, anglescaling, velocityscaling, maximumjointtorque. (3) 3.2 BioModule Architecture BioModule is a temporal transformer module that maps a sequence of root centred 3D skeletons to biomechanical attributes. The model has three main components: a framewise pose embedding, a temporal transformer encoder, and a set of independent attribute prediction heads. Figure 1 gives an overview of the architecture. 8 3.2.1 Pose Embedding and Temporal Encoding Each input frame vector x t ∈R 51 is projected into a hidden representation using a shared linear embedding: e t = W e x t + b e e t ∈R d (4) where d=256. Since self attention does not encode temporal order by itself, we add a fixed sinusoidal positional encoding [17]: PE(t, 2i) = sin t 10000 2i/d PE(t, 2i+1) = cos t 10000 2i/d (5) The transformer input is then z (0) t = e t + PE(t) Z (0) ∈R B×W×d (6) The positional encoding is fixed and introduces no additional trainable parameters. 3.2.2 Temporal Transformer Encoder The temporal encoder contains L=4 transformer layers with pre layer normaliza- tion [26]. Each layer applies multi head self attention across the full temporal window, followed by a feed forward network. For layer ℓ, the update is: ̃ Z (ℓ) = Z (ℓ−1) + MHA LN Z (ℓ−1) (7) Z (ℓ) = ̃ Z (ℓ) + FFN LN ̃ Z (ℓ) (8) The attention block uses h=8 heads with per head dimension d k =32. Attention is bidirectional over the full window. No causal mask is applied because BioModule pre- dicts the biomechanical state associated with the centre frame rather than forecasting an unseen future frame. This allows the representation to use both preceding and following motion context. The feed forward network uses a four times hidden expansion with GELU activation: FFN(u) = W 2 GELU ( W 1 u + b 1 ) + b 2 (9) where W 1 ∈R 4d×d and W 2 ∈R d×4d . Dropout with probability p=0.1 is used in the encoder. After the final transformer layer, a layer normalization operation produces the encoded sequence: H = LN Z (L) H∈R B×W×d (10) 9 3.2.3 Biomechanical Attribute Prediction Heads The encoded sequence H is passed to C=17 independent attribute prediction heads. Each head g a is a two layer MLP applied framewise: ˆ Y a = g a (H) ˆ Y a ∈R B×W×d a (11) For each attribute, the head has the form Linear(d→ d/2)→ GELU→ Dropout(0.1)→ Linear(d/2→ d a ) The shared encoder learns a temporal representation of skeletal motion, while the separate heads allow each biomechanical attribute to learn its own mapping from that representation. Continuous outputs are trained and predicted in normalized space. At inference, they are converted back to physical units using ˆ y = ̃ y σ a + μ a (12) where μ a and σ a are the training set mean and standard deviation for the corre- sponding attribute dimension. The binary touch head outputs logits, and a sigmoid is applied at inference to obtain contact probabilities. 3.3 Weighted Multi Task Loss BioModule is trained in a per-joint manner over 17 biomechanical attributes with different physical meanings, dimensionalities, and levels of uncertainty. A direct sum over all outputs would make the objective sensitive to the number of dimensions and noise level of each attribute. To reduce this effect, we use a tiered weighted multi task loss. Each attribute belongs to exactly one of the three biomechanical groups defined in Eq. 3. For each continuous attribute a, the loss is the mean squared error over the full output tensor: L a = 1 B W d a B X b=1 W X t=1 d a X j=1 ˆ y a b,t,j − y a b,t,j 2 (13) For the binary foot contact attribute, we use binary cross entropy with logits: L touch = BCEWithLogits ˆ Y touch Y touch (14) The three group losses are computed by averaging the individual attribute losses within each group: ̄ L kin = 1 |A kin | X a∈A kin L a (15) ̄ L knt = 1 |A knt | X a∈A knt L a (16) 10 ̄ L nmsc = 1 |A nmsc | X a∈A nmsc L a (17) The total training objective is: L total = 1.0 ̄ L kin + 0.5 ̄ L knt + 0.3 ̄ L nmsc (18) The weights reflect the relevance of each group to the observed pose sequence. Kinematic attributes receive the highest weight because they are most directly con- strained by skeletal motion. Kinetic attributes receive an intermediate weight because torques, powers, and reaction forces depend on inverse dynamics, contact assump- tions, and pose quality. Neuromuscular attributes receive the lowest weight because excitation, activation, and actuator level quantities are the most indirect and model dependent. Averaging within each group before applying the group weight prevents high dimensional outputs, such as the 68 dimensional neuromuscular attributes, from dominating the objective simply because they contain more output dimensions. The loss is computed over all W frames in the input window during training. During evaluation, metrics are computed only at the center frame t ⋆ = ⌊W/2⌋. This avoids boundary effects and gives each reported prediction symmetric temporal context. 4 Experimental Setup 4.1 Dataset Base data. H3.6Mplus [11] extends Human3.6M [1] with dense, per-frame biomechan- ical annotations derived from subject-specific OpenSim musculoskeletal simulations. The underlying motion capture corpus provides multi-camera video, 2D keypoint anno- tations, and 3D pose data for 7 subjects (S1, S5, S6, S7, S8, S9, S11) performing 30 standardized activities in a controlled laboratory environment, recorded at 50 fps across 4 synchronized camera views; the 3D joint set spans 32 body landmarks in the H36M convention. Layered on top of this skeleton data, H36Mplus pairs every frame with the outputs of subject-specific inverse kinematics and inverse dynamics pipelines run in OpenSim using a sex-matched musculoskeletal model (female: S1, S5, S7; male: S6, S8, S9, S11). These simulations provide per-frame estimates of general- ized joint coordinates, velocities, and accelerations, active and passive joint torques, ideal torques, instantaneous mechanical power both filtered and raw, bilateral ground reaction forces and torques, vertical seat reaction force, binary foot contact labels, and neuromuscular signals comprising muscle activation, neural excitation, force-length scaling, force-velocity scaling, and maximum isometric torque, all for 68 actuators spanning 34 actuated degrees of freedom (J7-J40). The degrees of freedom represent anatomical movements of their respective joint such as abduction, adduction, flexion, extension, etc. Cross-modal alignment. H36M world coordinates and OpenSim body-frame coor- dinates occupy entirely different reference systems: H36M joints live in a camera-rig world frame, while OpenSim generalized coordinates use a musculoskeletal body frame, making direct algebraic alignment impossible. We resolve this by using the H36M 11 camera calibration matrices namely intrinsics, rotation, translation as a shared projec- tion target, independently projecting both H36M 3D joints and OpenSim K-markers into the same four camera image planes. Co-registration is verified geometrically: the 17-joint H36M subset reproduced from projected 3D coordinates matches the native 2D keypoint annotations to sub-pixel accuracy (< 0.28 px), and the OpenSim pelvis marker (K1) coincides with H36M joint 0 to machine precision across all frames and cameras. This shared pelvis anchor ties the two coordinate systems together frame by frame, enabling reliable cross-modal supervision. This alignment produces frame-level paired pose and biomechanical labels for each subject-activity sequence. Data split. The dataset comprises 520,509 frames across 210 subject–activity clips. Subjects S1, S5, S6, S7, and S8 (157 clips, ≈108 min) form the training split; subjects S9 and S11 (53 clips,≈44 min) are held out as the test split and are never used during any training phase. Per-dimension Z-score statistics (μ, σ) are computed exclusively from training frames and applied identically across all experimental phases. 4.2 Implementation Details Hardware and framework. All experiments are implemented in PyTorch and trained on a single NVIDIA GeForce RTX 4080 GPU. The DataLoader uses 4 worker processes with pinned memory for accelerated host-to-device transfer. Input preprocessing. Given a clip of raw 3D joint positions, we extract the 17-joint H36M subset, apply root-centring (Eq. 1), flatten each frame to x t ∈R 51 , and tile into receptive-field sequences of W frames. During training, sequences are sampled with stride 1, yielding dense overlapping samples; during validation and testing, non- overlapping stride-W sampling is used. Edge frames at clip boundaries are handled by edge-replication padding. Continuous biomechanical attributes are Z-score normalized per output dimension using statistics computed only from the training subjects, and the same statistics are reused during frozen evaluation and pose-estimator adaptation. Input pose coordinates are not normalized, so BioModule learns a mapping from raw, root-centered 3D skeleton geometry to normalized OpenSim-derived biomechanical attributes. Training setup. The base model is trained from random initialization for 50 epochs using AdamW [44] (β 1 =0.9, β 2 =0.999, weight decay 10 −4 ) with an initial learning rate of 10 −4 annealed via cosine scheduling [45] (T max =50, η min =10 −6 ), batch size 64, dense stride-1 sampling, and global gradient norm clipping at 1.0. Fine-tuning from the base checkpoint runs for 10 epochs at a fixed learning rate of 10 −5 with all other hyperparameters unchanged and no validation set; the best checkpoint per estimator is selected by minimum training loss. 4.3 Evaluation Metrics Predictions are extracted at the center frame t ⋆ =⌊W/2⌋ of each sequence. Continuous outputs are de-normalized via ˆ y = ̃ y·σ +μ; touch logits are converted to contact prob- abilities via sigmoid. We report mean absolute error (MAE), root mean squared error (RMSE), and normalized MAE (nMAE) as a percentage of each criterion’s ground- truth range over the test set for continuous attributes s, and per-foot classification 12 accuracy for touch. Tier-level MAE is the unweighted mean of per-attribute MAEs within each tier. 4.4 Benchmarked Pose Estimators Seven contemporary 3D kinematic pose estimators spanning dilated convolution, transformer, diffusion, state-space, and graph architectures serve as upstream sup- pliers: VideoPose3D [3] (2019), MHFormer [18] (2022), D3DP [21] (2023), Pose- Mamba [24] (2024), MotionAGFormer [20] (2024), KTPFormer [22] (2024), and TCPFormer [23] (2025). For each estimator, pre-extracted 3D pose predictions are used as drop-in replacements for ground-truth poses without post-processing, isolating BioModule’s contribution from estimator-specific design choices. 4.5 Evaluation Protocol After training on ground-truth 3D poses, the learned BioModule weights are evalu- ated in a frozen setting by replacing the ground-truth pose input with the 3D output of each pose estimator. In this setting, no weight updates are performed. This mea- sures how biomechanical prediction quality changes when the input skeletons come from estimated 3D poses rather than ground-truth poses. In the adaptation set- ting, a separate BioModule checkpoint is fine-tuned for each pose estimator using that estimator’s predicted poses on the training subjects. Fine-tuning starts from the ground-truth-trained BioModule weights and runs for 10 epochs at a learning rate of 10 −5 . Continuous metrics are computed after de-normalization, so errors are expressed in the original physical units of each attribute. For touch, logits are con- verted to probabilities using a sigmoid and evaluated as binary contact predictions. Although BioModule predicts outputs for all frames in a window, evaluation uses only the center-frame prediction. The implementation of the whole pipeline can be found at: https://github.com/UTSA-VIRLab/BioModule 5 Results 5.1 Quantitative Results The quantitative results indicates the influence of the quality and temporal consis- tency of the upstream 3D pose sequence on biomechanical prediction accuracy. Across the evaluated models, BioModule generally produces more reliable estimates when the input skeletons preserve anatomically plausible joint relationships and stable temporal motion patterns. This trend is expected because the predicted biomechanical vari- ables are not independent frame-level labels. They are consequences of coordinated motion over time. Therefore, even when two pose estimators have similar average joint-position errors, their downstream biomechanical predictions may differ if one estimator produces smoother trajectories, more consistent limb orientations, or fewer local joint distortions. The results also indicate that errors do not propagate uniformly across all biome- chanical targets. Kinematic-related outputs are generally more directly tied to the observed skeletal geometry and therefore tend to be more stable across pose-estimator 13 Table 1 Biomechanical attributes’ MAE with the frozen weights protocol (RF = 81), evaluated on subjects S9 and S11 from Human3.6M. The testing on the ground-truth 3D poses obviously plays as the upper bound to those of 3D poses elicited from pose estimation models. KinematicKinetic 3D Pose / Model Coord. Speed Accel. Act.T. Pass.T. Ideal T. Inst.P. Inst.P. r GRF Seat R. H36M GT0.228 0.193 1.240 9.529 5.452 4.420 3.516 7.235 28.900 21.700 MHFormer0.658 0.403 2.377 15.400 6.927 10.600 4.502 7.52439.000 68.400 TCPFormer0.5340.587 4.290 16.900 6.53612.300 7.002 8.080 37.500 42.900 PoseMamba0.655 0.328 1.938 16.0007.008 10.8004.120 8.047 38.600105.400 VideoPose3D0.661 0.397 2.344 15.400 6.953 10.600 4.4767.704 38.900 68.000 MotionAGFormer 0.530 0.594 4.318 16.900 6.515 12.300 7.037 8.062 37.500 43.000 KTPFormer0.658 0.399 2.34315.400 6.939 10.600 4.484 7.537 39.100 67.100 D3DP0.656 0.400 2.346 15.400 6.937 10.600 4.500 7.512 39.100 66.600 NeuromuscularBinary 3D Pose / Model Act.Sig. Exc.Sig. N.Act.T. Ang.Sc. Vel.Sc. MaxJT. Touch H36M GT0.058 0.0560.057 0.054 0.117 2.049 67.500% MHFormer0.189 0.1240.101 0.217 0.187 16.10043.700% TCPFormer0.174 0.1680.0900.150 0.225 14.300 41.800% PoseMamba0.159 0.089 0.103 0.212 0.175 16.200 42.400% VideoPose3D0.190 0.1240.101 0.218 0.186 16.200 43.600% MotionAGFormer0.1740.167 0.090 0.150 0.226 14.300 42.100% KTPFormer0.190 0.1240.101 0.218 0.186 16.10043.700% D3DP0.190 0.1240.100 0.218 0.186 16.10043.600% Table 2 Biomechanical attributes’ MAE with the fine-tuned protocol (RF = 81), evaluated on S9 and S11 from Human3.6M. Fine-tuning done through 10 epochs on the estimator’s training subject poses (S1–S8) at lr = 10 −5 . KinematicKinetic ModelCoord. Speed Accel. Act.T. Pass.T. Ideal T. Inst.P. Inst.P. r GRF Seat R. MHFormer0.274 0.282 1.724 10.6005.556 5.523 3.892 6.872 30.800 27.000 TCPFormer0.260 0.282 1.689 10.500 5.5325.498 3.8259.516 29.700 25.300 PoseMamba0.269 0.276 1.703 10.700 5.545 5.560 3.842 6.779 31.000 27.300 VideoPose3D0.278 0.283 1.730 10.600 5.546 5.538 3.903 6.999 30.900 25.900 MotionAGFormer 0.2610.282 1.68810.500 5.524 5.475 3.816 9.590 29.90024.800 KTPFormer0.266 0.2711.690 10.500 5.537 5.4053.857 6.921 31.100 25.900 D3DP0.266 0.270 1.679 10.500 5.540 5.355 3.840 6.86230.600 24.200 NeuromuscularBinary ModelAct.Sig. Exc.Sig. N.Act.T. Ang.Sc. Vel.Sc. MaxJT. Touch MHFormer0.069 0.0620.0640.070 0.159 4.635 67.600% TCPFormer0.071 0.0670.065 0.070 0.167 5.809 68.600% PoseMamba0.069 0.0620.064 0.070 0.158 4.791 67.600% VideoPose3D0.069 0.0630.064 0.070 0.160 6.136 68.300% MotionAGFormer 0.071 0.0670.065 0.070 0.167 5.026 68.200% KTPFormer0.068 0.0610.063 0.067 0.1563.874 67.500% D3DP0.067 0.061 0.063 0.0670.155 4.27868.600% inputs. In contrast, kinetic and neuromuscular quantities are more sensitive to subtle changes in joint angle, velocity, and temporal coordination. Thus, the performance gap among upstream pose estimators becomes more visible as the target variable moves 14 from geometric motion description toward physically and physiologically interpretable quantities. The comparison across the seven pose estimators further demonstrates that pose- estimation accuracy alone is not sufficient to fully explain biomechanical prediction quality. A pose estimator that performs well in terms of joint localization may still introduce errors that are biomechanically meaningful, such as inconsistent knee flex- ion, unstable hip orientation, or unnatural ankle positioning during walking. These errors may have limited impact on conventional pose metrics but can strongly affect torque and muscle-related predictions. This finding supports the central motivation of BioModule: biomechanical evaluation requires attention not only to where the joints are located, but also to whether the estimated motion preserves physically meaningful relationships among body segments. Another important observation is that BioModule remains functional across all evaluated upstream models, which supports its estimator-agnostic design. Since BioModule operates on standard 17-joint 3D skeletons, it does not require retrain- ing or redesigning the original pose-estimation models. This makes the framework practical for comparing different pose estimators. At the same time, the variation in performance across models shows that modularity does not eliminate the influence of upstream error. Instead, the downstream structure of this research makes that influence measurable. The results therefore provide not only a benchmark of BioMod- ule performance, but also an analysis of how pose-estimation quality propagates into biomechanical inference. 5.2 Qualitative Results Skeletal visualizations of four biomechanical attributes from test subject S9 during a walking sequence are presented in Figure 2. The figure compares the BioModule outputs obtained from ground-truth 3D poses and from seven upstream 3D pose estimators for the same representative frame. The red spectrum indicates the value magnitude of the corresponding attribute across body segments. The walking frame provides an informative case because gait involves coordinated loading across the hip, knee, and ankle, with different muscles and torques contributing at different phases of the movement.In the ground-truth visualization, higher biome- chanical responses are concentrated around the lower-limb joints, particularly the hip, knee, and ankle, reflecting their dominant role in generating and controlling walk- ing motion. This spatial distribution serves as a reference for assessing whether the predicted biomechanical patterns remain anatomically plausible. For active torque, the qualitative results are expected to reveal how well each pose- estimator input preserves the joint-level loading pattern of the walking frame. Active torque reflects the net muscular effort required to generate or control motion. Models that produce less accurate pose input with local distortions around the lower limbs may shift the predicted torque magnitude or produce unnatural concentration at the wrong joint. Unlike active torque, passive torque is strongly related to joint configuration and soft-tissue resistance and hence more sensitive to joint-angle errors. If an upstream pose estimator produces excessive or insufficient knee flexion in the walking frame, 15 GTD3DP KTPFormer MHFormer MotionAG Former PoseMamba TCPFormer VideoPose3D Active Torque 7.2110.249.239.185.204.971.609.91 Passive Torque 0.651.000.961.240.941.440.921.07 Muscle Activation 0.060.140.140.130.170.090.150.13 Neural Excitation 0.090.030.030.040.110.050.110.03 Fig. 2 Qualitative comparison of BioModule predictions for a sample walking frame from test subject S9. Rows show four biomechanical attributes. Columns compare predictions obtained from ground- truth 3D poses and from seven upstream 3D pose estimators. Bone coloring encodes the predicted attribute magnitude and the number above each panel reports the mean predicted value across all body segments. 16 the predicted passive torque may become exaggerated or suppressed relative to the ground truth. This makes passive torque a useful qualitative indicator of whether the estimated skeleton remains within plausible biomechanical ranges. Muscle activation is not directly visible from the skeleton, but it is inferred from the relationship between posture, motion, and the learned biomechanical supervision. Scattered or misplaced activation patterns may indicate that the input skeleton lacks the temporal coherence needed to support reliable muscle-level inference. Neural excitation is expected to be among the more challenging outputs because it represents a control signal rather than a directly observable geometric quantity. Qualitative differences in neural excitation therefore provide insight into the limits of downstream inference from 3D skeletons alone. If the predicted excitation pat- terns remain spatially and functionally consistent with the ground truth, this supports the ability of BioModule to infer higher-level biomechanical attributes from pose sequences. If the predictions become noisy or anatomically inconsistent, this suggests that some neuromuscular quantities may require richer input representations, stronger temporal modeling, or additional physical constraints. The qualitative results are expected to show that the best-performing upstream models do not simply produce cleaner skeletons; they preserve biomechanically mean- ingful structure which also depends highly on the action scenario. When the upstream pose contains local errors, the downstream biomechanical maps may amplify those errors in different manners. The visual patterns in Figure 2 support the quantitative findings. They show that the downstream formulation makes biomechanical consequences of pose-estimation error visible. Rather than treating all pose errors as equally important, the visual- izations reveal which errors matter more for physical interpretation. This distinction is central to the purpose of BioModule: to connect pose-estimation outputs with biomechanical meaning. 6 Discussion This study evaluates BioModule as a modular temporal transformer for predicting biomechanical attributes from 3D human pose sequences. In this work, biomechan- ical prediction is treated as a downstream task as a methodological design choice which allows BioModule to attach to existing 3D pose estimators without modify- ing their architectures. biomechanical attributes can be learned from pose sequences alone. In this sense, BioModule provides an initial bridge between the conventional 3D pose-estimation setting and biomechanical analysis of human motion. As a byprod- uct,this setting makes it possible to compare how different upstream pose models affect downstream biomechanical prediction. Benchmarking of BioModule on the outputs of SOTA 3D pose estimation model sets an important milestone because most video-based human pose estimation pipelines produce sparse skeletal representations rather than full musculoskeletal states. Therefore, by learning the mapping from 3D pose sequences to simulation- derived biomechanical attributes, BioModule shows that biomechanical interpretation can be approached directly from video-compatible pose representations. 17 In qualitative walking visualizations, the skeletal maps of active torque, passive torque, muscle activation provide insight beyond numerical error values by showing where biomechanical demand is concentrated and whether the predicted distribution remains consistent with expected movement behavior. BioModule provides a baseline formulation for connecting standard 3D pose esti- mation with biomechanical motion analysis. By predicting biomechanical attributes from pose sequences, it treats the skeleton as an intermediate representation rather than only a geometric output. This provides an initial step toward video-based biome- chanical assessment without motion-capture systems, wearable sensors, force plates, or separate inverse-kinematics pipelines. Limitations and Future Work: The current study is intended as a baseline rather than a complete solution for in-the-wild biomechanical analysis. BioModule relies on aligned biomechanical supervision and therefore inherits limitations from the underlying simulation-derived labels, musculoskeletal assumptions, and dataset alignment process. In addition, the use of a reduced 17-joint skeleton improves com- patibility with standard 3D pose estimators but limits anatomical detail compared with full-body marker sets or subject-specific musculoskeletal models. The evaluation is also limited to Human3.6M-based motion data, which is captured in a controlled environment and does not fully represent outdoor videos, occlusion, camera motion, clothing variation, clinical movement patterns, or complex sports activities. Therefore, the present results should be interpreted as evidence that biome- chanical attributes can be learned from pose sequences under controlled conditions, and more data including pose and the corresponding biomechanical attributes is needed to bring the field to the point where validation for unconstrained real-world deployment is feasible. Future work should extend this framework to richer skeletal representations, subject-specific biomechanical modeling, uncertainty-aware prediction, and broader activity domains. Further validation is needed on real-world video, sports move- ments, rehabilitation tasks, and clinical populations. A longer-term direction is to combine video-based pose estimation with BioModule-like biomechanical prediction so that human motion analysis can be performed without intrusive sensors, expensive laboratory equipment, or separate inverse-kinematics pipelines. 7 Conclusion This work presented BioModule, a lightweight temporal transformer for predicting biomechanical attributes from 3D human pose sequences. In this study, biomechanical prediction was formulated as a downstream modular extension of 3D pose estima- tion. This formulation was chosen to allow BioModule to operate on the outputs of existing pose estimators without modifying their architectures. It should therefore be understood as the design adopted in this research, not as a claim that biomechanical estimation must always be performed downstream of pose estimation. Using aligned Human3.6M and Human3.6Mplus supervision, BioModule was trained to infer biomechanical quantities from standard 17-joint skeletons. The exper- iments across seven upstream pose estimators showed that biomechanical prediction is 18 strongly affected by the anatomical plausibility and temporal consistency of the input pose sequence. The results also showed that different biomechanical outputs have dif- ferent sensitivity to upstream error. Kinematic quantities are more directly related to skeletal geometry, while kinetic and neuromuscular quantities, including active torque, passive torque, muscle activation, and neural excitation, are more sensitive to subtle pose and motion artifacts. The qualitative walking visualizations further demonstrated that biomechani- cal evaluation provides insight beyond conventional pose-estimation metrics. Errors that may appear small geometrically can become important when they occur near mechanically active joints such as the knee, hip, or ankle. By visualizing predicted biomechanical quantities across the skeleton, BioModule helps reveal whether an estimated pose sequence preserves physically meaningful movement structure. Overall, the findings indicate that BioModule can serve as a compact bridge between computer vision-based 3D pose estimation and biomechanical motion analy- sis. Its estimator-agnostic design makes it useful for extending existing pose-estimation pipelines, while its output space provides a more functionally meaningful way to evaluate human motion. Future work will explore richer skeletal representa- tions, uncertainty-aware biomechanical prediction, stronger physical constraints, and broader validation across activities, subjects, and real-world movement conditions. 8 Acknowledgments This material is partially based upon work supported by the National Science Foundation under Grant No. 2153249. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. References [1] Ionescu, C., Papava, D., Olaru, V., Sminchisescu, C.: Human3.6M: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence 36(7), 1325– 1339 (2014) https://doi.org/10.1109/TPAMI.2013.248 [2] Martinez, J., Hossain, R., Romero, J., Little, J.J.: A simple yet effective base- line for 3d human pose estimation. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV), p. 2640–2649 (2017) [3] Pavllo, D., Feichtenhofer, C., Grangier, D., Auli, M.: 3d human pose estimation in video with temporal convolutions and semi-supervised training. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7753–7762 (2019) [4] Gozlan, Y., et al.: OpenCapBench: A benchmark to bridge pose estimation and biomechanics. In: Proceedings of the IEEE/CVF Winter Conference on Applica- tions of Computer Vision (WACV) (2025). Also introduces SynthPose for denser 19 keypoint prediction [5] Delp, S.L., Anderson, F.C., Arnold, A.S., Loan, P., Habib, A., John, C.T., Guen- delman, E., Thelen, D.G.: OpenSim: Open-source software to create and analyze dynamic simulations of movement. IEEE Transactions on Biomedical Engineering 54(11), 1940–1950 (2007) https://doi.org/10.1109/TBME.2007.901024 [6] Seth, A., Hicks, J.L., Uchida, T.K., Habib, A., Dembia, C.L., Dunne, J.J., Ong, C.F., DeMers, M.S., Rajagopal, A., Millard, M., Hamner, S.R., Arnold, E.M., Yong, J.R., Lakshmikanth, S.K., Sherman, M.A., Ku, J.P., Delp, S.L.: Open- Sim: Simulating musculoskeletal dynamics and neuromuscular control to study human and animal movement. PLOS Computational Biology 14(7), 1006223 (2018) https://doi.org/10.1371/journal.pcbi.1006223 [7] Uhlrich, S.D., et al.: OpenCap: Human movement dynamics from smartphone videos. PLOS Computational Biology (2023). Open-source platform; uses at least 2 smartphone videos for kinematics and dynamics [8] Koleini, M., et al.: BioPose: Biomechanically-accurate 3d pose estimation from monocular videos. In: Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision (WACV) (2025). Combines MQ-HMR mesh recovery with NeurIK for biomechanical constraints [9] Keller, M., Werling, K., Shin, S., Delp, S., Pujades, S., Liu, C.K., Black, M.J.: From skin to skeleton: Towards biomechanically accurate 3d digital humans. ACM Transactions on Graphics (TOG) 42(6), 1–12 (2023) [10] Xia, Y., Zhou, X., Vouga, E., Huang, Q., Pavlakos, G.: Reconstructing humans with a biomechanically accurate skeleton. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 5355–5365 (2025) [11] Nasr, A., Zhu, K., McPhee, J.: Using musculoskeletal models to generate physically-consistent data for 3d human pose, kinematic, dynamic, and muscle estimation. Multibody System Dynamics (2024) [12] Toshev, A., Szegedy, C.: DeepPose: Human pose estimation via deep neural net- works. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 1653–1660 (2014) [13] Wei, S.-E., Ramakrishna, V., Kanade, T., Sheikh, Y.: Convolutional pose machines. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 4724–4732 (2016) [14] Newell, A., Yang, K., Deng, J.: Stacked hourglass networks for human pose estimation. In: Proceedings of the European Conference on Computer Vision (ECCV), p. 483–499 (2016) 20 [15] Sun, K., Xiao, B., Liu, D., Wang, J.: Deep high-resolution representation learning for human pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 5693–5703 (2019) [16] Zheng, C., Zhu, S., Mendieta, M., Yang, T., Chen, C., Ding, Z.: PoseFormer: 3d human pose estimation with spatial and temporal transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 11656–11665 (2021) [17] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems, vol. 30 (2017) [18] Li, W., Liu, H., Tang, H., Wang, P., Van Gool, L.: MHFormer: Multi-hypothesis transformer for 3d human pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 13147– 13156 (2022) [19] Zhao, Q., Zheng, C., Liu, M., Wang, P., Chen, C.: PoseFormerV2: Exploring frequency domain for efficient and robust 3d human pose estimation. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 8877–8886 (2023) [20] Mehraban, S., Adeli, V., Taati, B.: MotionAGFormer: Enhancing 3d human pose estimation with a transformer-GCNformer network. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 6920–6930 (2024) [21] Shan, W., Liu, Z., Zhang, X., Wang, Z., Han, K., Wang, S., Ma, S., Gao, W.: Diffusion-based 3D human pose estimation with multi-hypothesis aggregation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 14761–14771 (2023) [22] Peng, J., Zhou, Y., Mok, P.Y.: KTPFormer: Kinematics and trajectory prior knowledge-enhanced transformer for 3D human pose estimation. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 1123–1132 (2024) [23] Liu, J., Liu, M., Liu, H., Li, W.: TCPFormer: Learning temporal correlation with implicit pose proxy for 3D human pose estimation. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, p. 5478–5486 (2025) [24] Huang, Y., Liu, J., Xian, K., Qiu, R.C.: PoseMamba: Monocular 3D human pose estimation with bidirectional global-local spatio-temporal state space model. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, p. 3842– 3850 (2025) 21 [25] Zhu, W., Ma, X., Liu, Z., Liu, L., Wu, W., Wang, Y.: MotionBERT: A unified perspective on learning human motion representations. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 15085– 15099 (2023) [26] Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., Liu, T.: On layer normalization in the transformer architecture. In: International Conference on Machine Learning, p. 10524–10533 (2020). PMLR [27] Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., Black, M.J.: SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (Proc. SIGGRAPH Asia) 34(6), 248–124816 (2015) [28] Kanazawa, A., Black, M.J., Jacobs, D.W., Malik, J.: End-to-end recovery of human shape and pose. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 7122–7131 (2018) [29] Kocabas, M., Athanasiou, N., Black, M.J.: VIBE: Video inference for human body pose and shape estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 5253–5263 (2020) [30] Li, D., Jin, J., Yu, X., Liu, W., Cun, X., Chen, K., Fan, R., Kong, J., Shen, X.: Skel-cf: Coarse-to-fine biomechanical skeleton and surface mesh recovery. arXiv preprint arXiv:2511.20157 (2025) [31] Gilon, O., et al.: OpenCap monocular: Biomechanical analysis from a single smartphone video. arXiv preprint (2026). Extends OpenCap to monocular input; refines via WHAM and physics-based simulation [32] Cotton, R.J., DeLillo, A., Cimorelli, A., Shah, K., Peiffer, J.D., Anarwala, S., Abdou, K., Karakostas, T.: Optimizing trajectories and inverse kinematics for biomechanical analysis of markerless motion capture data. In: Proceedings of the IEEE International Conference on Rehabilitation Robotics (ICORR), p. 1–6 (2023) [33] Ruescas-Nicolau, A.-V., et al.: Markerless motion capture with keypoint augmen- tation for biomechanical applications. Sensors (2024). Investigates error factors in deep learning markerless pose estimation for biomechanics [34] Auer, S., et al.: Using markerless motion capture and musculoskeletal models for kinematic analysis. Proceedings of the Institution of Mechanical Engineers (2024). Lower-limb joint angle MAEs 4–6°; GRF errors 6% body weight for video-based MoCap [35] Barzyk, T., et al.: AI-smartphone markerless capture of hip, knee, and ankle angles during countermovement jumps. European Journal of Sport Science (2024). Validated against Vicon; uses multilevel CNN for full-skeleton keypoint 22 localization [36] Rode, C., et al.: Assessment of monocular human pose estimation models for clin- ical movement analysis. Scientific Reports (2025). Benchmarks markerless pose estimation as an alternative to clinical MoCap [37] Lin, K., et al.: 3d kinematics from video with a biomechanical model and synthetic training data. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (2024). Bridges vision pipelines with musculoskeletal constraints using synthetic data [38] Miller, E.Y., Tan, T., Falisse, A., Uhlrich, S.D.: Integrating machine learning with musculoskeletal simulation improves opencap video-based dynamics estimation. bioRxiv, 2025–12 (2025) [39] Werling, K., et al.: AddBiomechanics: Automating model scaling, inverse kine- matics, and inverse dynamics from optical motion capture data and musculoskele- tal models. PLOS ONE (2023) [40] Werling, K., et al.: AddBiomechanics dataset: Capturing the physics of human motion at scale. arXiv preprint (2024). Built on the Rajagopal Full Body Model; bridges CV and biomechanics via SMPL compatibility [41] Tripathi, S., et al.: IPMAN: 3d human pose estimation via intuitive physics. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2023). Differentiable physics terms encouraging floor contact and CoP/CoM alignment [42] Gambietz, M., Dorschky, E., Akat, A., Sch ̈ockel, M., Miehling, J., Koelewijn, A.D.: Sspinnpose: A self-supervised pinn for inertial pose and dynamics estimation. ACM Transactions on Intelligent Systems and Technology (2025) [43] Bian, Q., Wang, H., Alsayed, K., Ding, Z.: Orientationnn: a physics-informed lightweight neural network for real-time joint kinematics estimation from imu data. Frontiers in Bioengineering and Biotechnology 13, 1737916 (2025) [44] Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: Proceedings of the International Conference on Learning Representations (ICLR) (2019) [45] Loshchilov, I., Hutter, F.: SGDR: Stochastic gradient descent with warm restarts. In: Proceedings of the International Conference on Learning Representations (ICLR) (2017) 23