Paper deep dive
MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/29/2026, 3:57:03 AM
Summary
The paper introduces MyoMechanix, a multimodal ecosystem for biomechanically-grounded action quality assessment (AQA) in weight-loaded exercises. It presents a large-scale dataset with synchronized RGB video, 3D pose, sEMG, and physiological signals from 38 subjects. The work also introduces the Fitness Knowledge Graph (FKG) for structured reasoning and CUBIST, a compositional ontological reasoning engine for fine-grained error attribution. Key tasks include AQA, VideoQA, and a novel Video2EMG task.
Entities (7)
Relation Signals (6)
MyoMechanix → includesmodality → sEMG
confidence 98% · MyoMechanix... contains... synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals
MyoMechanix → supportstask → Action Quality Assessment
confidence 96% · MyoMechanix... forming the largest multimodal AQA benchmark to date.
CUBIST → utilizes → Fitness Knowledge Graph
confidence 95% · Building on these representations [FKG], we develop CUBIST... which performs decomposition-analysis-recomposition
MyoMechanix → supportstask → Video2EMG
confidence 94% · We further establish... a novel MyoMechanix-Video2EMG task.
CUBIST → achievesstateoftheart → Action Quality Assessment
confidence 93% · CUBIST achieving state-of-the-art results
Fitness Knowledge Graph → structures → Action Quality Assessment
confidence 92% · FKG... enables compositional scoring and interpretable assessment.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic patterns. These limitations hinder fine-grained, biomechanically grounded feedback. We introduce MyoMechanix, a multimodal ecosystem for weight-loaded actions that aligns motion with muscle activity. Expert-annotated, it contains 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, and additional physiological signals, forming the largest multimodal AQA benchmark to date. We further construct the Fitness Knowledge Graph (FKG), which organizes expert annotations into structured relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment. Building on these representations, we develop CUBIST (Compositional Ontological Reasoning Engine), which performs decomposition-analysis-recomposition for fine-grained error attribution and feedback generation. We also establish MyoMechanix-AQA, MyoMechanix-VideoQA, and a novel MyoMechanix-Video2EMG task. Experiments show that multimodal sensing and structured representations improve performance, interpretability, and error attribution, with CUBIST achieving state-of-the-art results; VideoQA enhances language-grounded action understanding; and Video2EMG suggests video-based alternatives to costly EMG sensing. MyoMechanix advances skilled activity understanding toward biomechanically grounded, multimodal, and compositional reasoning for Physical AI applications in fitness, rehabilitation, healthcare, and machine learning. Project page: this https URL
Tags
Links
- Source: https://arxiv.org/abs/2608.26094v1
- Canonical: https://arxiv.org/abs/2608.26094v1
Trouble viewing inline? Open PDF directly →
Full Text
175,956 characters extracted from source content.
Expand or collapse full text
MYOMechanix| Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching Hao Yin 1, 2* · Paritosh Parmar 3* Lijun Gu 1, 2 · Lin Xu 4 · Tianxiao Guo 5 · Xiujin Liu 6 · Tianyou Zheng 2 Yang Zhang 2 · Weiwei Fu 1, 2 AbstractExisting skilled activity understanding/action quality assessment (AQA) datasets & methods face two key limitations: they rely primarily on explicit visual inputs (e.g., RGB & Pose data), overlooking implicit physiological dynamics such as muscle mechanics, & they model actions as monolithic, non-compositional patterns, limiting semantic richness. These constraints hinder fine-grained, biomechanically grounded feedback. To address this, we pioneer a paradigm that integrates multimodal sensing, structured representations, & com- positional reasoning. We introduce MyoMechanix, the first of its kind biomechanically grounded multi- modal ecosystem for weight-loaded actions with strong internal–external alignment, enabling faithful coupling between motion & muscle activity. Expert-annotated, it comprises 7,500+ samples of 20 actions from 38 subjects, with synchronized multiview RGB video, 3D pose, sEMG, & additional physiological signals, form- ing the largest multimodal AQA benchmark to date. Building upon this sensing foundation, we construct the Fitness Knowledge Graph (FKG), which or- *These authors contributed equally to this work. Weiwei Fu E-mail: fuww@sibet.ac.cn Yang Zhang E-mail: zhangyang@sibet.ac.cn 1 School of Biomedical Engineering (Suzhou), Division of Life Sciences and Medicine, University of Science and Technology of China, Hefei, Anhui, China 2 Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou, Jiangsu, China 3 Arexeni Research and Technologies Inc. 4 School of Psychology, Beijing Sports University, Beijing, China 5 School of Competitive Sports, Beijing Sports University, Beijing, China 6 Department of Robotics, University of Michigan, Ann Arbor, Michigan, USA ganizes annotations from field experts into structured relationships linking actions, phases, key steps, error types, & corrective feedback. This representation sup- ports a compositional scoring & reasoning framework that enables interpretable & fine-grained quality as- sessment. To capitalize on the multimodal richness & structured representations, we develop a new model- ing paradigm, CUBIST (Compositional Ontological Reasoning Engine), a framework that performs decom- position–analysis–recomposition over structured actions to enable fine-grained error attribution and feedback generation. We further establish the MyoMechanix- AQA benchmark & MyoMechanix-VideoQA, sup- porting tasks ranging from action quality assessment to error diagnosis & cross-modal reasoning, including a novel MyoMechanix-Video2EMG task. Three key findings emerge from extensive experiments: 1) multi- modal sensing with structured representations improves performance, interpretability, & error attribution, with CUBIST achieving state-of-the-art results; 2) VideoQA enhances language-grounded, fine-grained action under- standing in vision–language models; 3) Video2EMG re- sults suggest cheap video-based alternatives to expensive EMG sensors. Nonetheless, MyoMechanix remains chal- lenging, motivating further research. This work advances skilled activity understanding from visual recognition to biomechanically grounded, multimodal, & compositional reasoning, providing a foundation for next-generation Physical AI with applications in fitness, rehabilitation, healthcare, & core machine learning research, particu- larly, representation learning. Project page: MyoMechanix Ecosystem. Keywords Interpretable Action Quality Assessment, Physical Skills Assessment, Video Understanding, AI-driven Coaching, Fitness, sEMG, Physical AI arXiv:2608.26094v1 [cs.CV] 26 Aug 2026 2 CUBIST: Compositional Ontological Reasoning Engine Action Quality Assessment VideoQA/ Video Caption Video2EMG Cross-modal Coach Agent Self-supervised Federated Learning ... More Compositional Score Engine Stage I Ontology-driven Error Reasoning Stage I Action Perception- driven Expert Routing Stage I Compositional Score = [1-(10+9)/134]×100=85.82 Knowledge Graph A1 6 AK 56 AK 43 AK 44 AK 45 ET 061 ET 003 ET 059 ET 002 ET 007 AK 46 ET 032 No error.ET032: Excessive knee flexion. (9) ET055: Excessive forward movement of the knees. (10) Eccentric Phase AK44 Concentric Phase AK45 Preparation AK56+AK43 5 Viewpoints 0246810121416 01000020000300004000050000 Biomechanical Information Heart Rate Breath Rate sEMG 20 Actions 38 Subjects 7512 Samples Fig. 1 An overview of our MyoMechanix ecosystem. We introduce the MyoMechanix ecosystem, the first biomechanically grounded multimodal framework for skilled activity understanding and coaching. MyoMechanix advances fitness activity understanding along two complementary dimensions: physical fidelity and the semantic richness of actions. For data collection, we use high-precision motion capture to record 7,512 samples from 38 subjects spanning three ability levels and performing 20 fitness actions, making it the largest multimodal fitness AQA benchmark. The dataset includes synchronized recordings from five RGB camera views, 3D poses, sEMG, and physiological signals, and is annotated by multiple fitness experts. To support structured reasoning, MyoMechanix incorporates a fitness knowledge graph (FKG) that formalizes relationships among action key steps (AK), error types (ET), and action feedback. FKG recasts fitness actions as structured procedures, enabling fine-grained error reasoning, feedback generation, and interpretable scoring. This framework shifts the paradigm from subjective overall scoring to objective, fine-grained error annotation and establishes a computational scoring system. Building on these resources, we develop benchmark datasets for AQA, VideoQA, and Video2EMG, providing a versatile foundation for future multidisciplinary research. Furthermore, we propose CUBIST, a Compositional Ontological Reasoning Engine. Through a three-stage progression from action perception to scoring, CUBIST enables precise quantitative evaluation with high interpretability at both the intermediate reasoning and outcome levels. 1 Introduction Last spring, Christina, a fitness enthusiast attempt- ing a personal-best barbell squat, appeared to exhibit correct form when assessed visually—both through self- observation and by an AI-based coaching system. How- ever, subtle misalignments in load distribution and im- proper muscle activation ultimately led to a lower-back injury. Crucially, neither the system nor the athlete was able to identify which specific phase of the movement caused the failure, nor to uncover its underlying physi- ological origin. This example highlights a fundamental limitation of current action analysis approaches: their reliance on external visual cues, lack of access to in- ternal physiological signals, and inability to decompose actions into interpretable components for fine-grained reasoning. Human motion is fundamentally governed by the neuromuscular system, in which observable kinematics emerge from complex coordination patterns of muscle activation. Due to the redundancy of the musculoskele- tal system, visually similar movements can arise from substantially different internal control strategies—some of which may compromise safety and performance. Con- sequently, errors in execution often manifest not as obvious deviations in motion trajectories, but as latent inconsistencies in muscle coordination and force gen- eration. Effective assessment, therefore, requires both physiological grounding and structured representations that capture the compositional nature of human actions. Existing research on Action Quality Assessment (AQA) 1 has achieved significant progress in modeling 1 The task of computational assessment of how well an action is performed has been described using various terms, 3 During the squat, the individual demonstrates a tendency to ... Enhance performance, and minimize the risk of injury. I will award X points. AQA Model Referee Teach Mimic Fitness Action Action Rule Description Feedback Score More ... Fig. 2 Illustration of AQA task. Given an action video, the AQA model plays the role of a referee to evaluate perfor- mance and score skill levels. spatiotemporal dynamics using convolutional neural net- works [1], Transformers [2], and graph-based architec- tures [3, 4] (see Fig. 2). However, these approaches re- main in several critical aspects. Firstly, they primarily concentrate on exteroceptive visual inputs and explicit kinematic features such as joint positions, limb config- urations, and temporal patterns [5]. As a result, they struggle to capture crucial information regarding im- plicit physiological states—including muscle activation, fatigue, and coordination strategies [6]. Secondly, they tend to treat actions as monolithic, non-compositional patterns rather than structured compositions. In addi- tion, widely used AQA datasets are often derived from web-crawled videos, which suffer from inconsistent qual- ity, limited diversity, and coarse-grained annotations lacking semantic structure and causal interpretability. Collectively, these limitations hinder the development of models capable of fine-grained, physiologically informed reasoning about complex human actions. To address these challenges, we introduce a new paradigm for action understanding that integrates: 1) sensorimotor integration, 2) structured representation, and 3) compositional reasoning. Concretely, we propose to model human actions as biomechanically grounded, structured processes, in which observable motion is cou- pled with underlying physiological dynamics and orga- nized into semantically meaningful components. Our main contributions can be summarized as follows. First, we introduce MyoMechanix, a multimodal data acquisition and analysis ecosystem for weight- loaded exercises. In contrast to prior datasets, My- oMechanix combines synchronized multi-view video (in- cluding ubiquitous phone camera), high-precision mo- tion capture, and surface electromyography (sEMG) [7, 8], enabling direct measurement of muscle mecha- including Skilled Activity Understanding, AQA, Skills As- sessment, & Proficiency Estimation. Consequently, this study treats these terms as conceptually equivalent & uses them interchangeably. nisms/activity alongside external motion. The inclusion of sEMG provides access to internal neuromuscular sig- nals, offering a crucial complementary perspective to visual observation and enabling a fundamentally more comprehensive characterization of action quality. The resulting dataset comprises over 7,500 samples across 20 exercise types performed by 38 subjects, totaling more than 40 hours of multimodal recordings (see Fig. 1). Second, we propose the Fitness Knowledge Graph (FKG), which formalizes human actions as structured action representations. Specifically, each action is decom- posed into hierarchical phases (e.g., preparation, con- centric, eccentric) and further subdivided into ordered action key steps (AKs) associated with domain-specific error types (ETs) and corrective feedback (FBs). This representation introduces an explicit semantic structure that captures both the procedural organization of ac- tions and the causal relationships between execution errors and their consequences. Building upon this ontol- ogy, we design a compositional, penalty-based annota- tion protocol that assigns weights to errors according to their biomechanical impact, enabling interpretable and fine-grained evaluation. Third, we introduce a new modeling paradigm— CUBIST (Compositional Ontological Reasoning En- gine), a framework for compositional reasoning over hu- man actions that leverages the multimodal richness and structured representations of FKG. Inspired by the prin- ciple of decomposition–analysis–recomposition, CUBIST operates by (i) decomposing actions into structured components, (i) analyzing each component through multimodal inputs and specialized processing modules, and (i) recomposing the results to produce coherent assessments, including error attribution and corrective feedback. This design enables the model to move beyond holistic scoring toward interpretable, step-wise reasoning that more closely reflects expert analysis. Based on multimodal data and structured annota- tions of this ecosystem, we establish the MyoMechanix- AQA benchmark for structured action quality assess- ment, as well as MyoMechanix-VideoQA, a dataset sup- porting tasks ranging from action recognition to fine- grained error diagnosis and feedback generation. Impor- tantly, we introduce Video2EMG, a novel cross-modal task that aims to infer muscle activation patterns from visual input. Extensive experiments demonstrate that firstly, the integration of multimodal sensing and structured rep- resentations significantly improves performance over existing SOTA baselines, while also enabling richer inter- pretability and more precise error attribution. Our CU- BIST achieves SOTA performance, significantly outper- forming existing AQA models. Secondly, the VideoQA 4 dataset helps instill deeper action understanding in vision-language models, supporting the development of foundation models for language-grounded, query- conditioned, fine-grained action understanding. Thirdly, we demonstrate promising results on the cross-modal task—Video2EMG, paving the way for video-based al- ternatives to expensive, specialized EMG sensors. Over- all, despite these advances, MyoMechanix remains a challenging benchmark with considerable room for im- provement, motivating future research in physiologically grounded and compositional action understanding. In summary, this work advances action understand- ing from purely visual pattern recognition toward biome- chanically grounded, multimodal, and compositional reasoning. We believe this paradigm provides a founda- tion for future research in physically grounded AI, with broad implications for applications in fitness training, rehabilitation, and smart healthcare systems. The remainder of our paper is organized as follows: –Section 2 provides a comprehensive literature review of relevant datasets and research methodologies. –Section 3 details the construction process and proto- cols of the MyoMechanix ecosystem. –Section 4 elaborates on the architectural design and algorithms of the proposed CUBIST model. – Section 5 presents comprehensive experimental re- sults and in-depth analyses of various baseline models on the proposed benchmarks. –Finally, Section 6 concludes the paper and discusses potential directions for future research. 2 Related Work We comprehensively compare our work with prior art in related areas, especially action quality assessment datasets and models, video question answering datasets, electromyography datasets, and muscle state estimation. 2.1 Action Quality Assessment Datasets The AQA field has advanced significantly over the past decade, with high-quality datasets (e.g., [9, 16, 26–30]) playing a central role in this progress. However, despite this growth, existing datasets remain fundamentally limited in capturing the full complexity of human move- ment, particularly with respect to internal physiological processes, skill diversity, and structured reasoning. Our dataset aims to address this as follows. Lack of physiological signals such as muscle mechanics. A wide variety of AQA datasets are available, contain- ing various modalities such as Video [9, 19, 25, 26, 31], Pose [12, 16, 24, 32], Audio [15], and Language [11]. How- ever, to the best of our knowledge, incorporating physi- ological signals such as electromyography (EMG), heart rate, etc., remains to be accomplished. Signals such as EMG can be very useful in AQA because they are probes of muscles—actuators of human movement—that tell how the muscles are being fired/activated and whether they are performing optimally. These are important in- sights because, often, a person’s posture on the surface may look normal, but underneath, muscles are not being activated properly, are producing imbalanced forces, or are overcompensating for other muscles, etc., resulting in injury or the development of bad form in the long term. Importantly, a lack of muscle-level information hinders the ability to provide precise coaching recom- mendations. MyoMechanix bridges this crucial gap by providing multimodal information, including synchro- nized multiview RGB video, 3D pose, sEMG signals, heart signals, and breathing signals, across over 7.5K samples for 20 different exercise procedural actions that involve multiple joints and are highly prone to injury Table. 2. Lack of diversity in ability levels. A majority of the datasets in the AQA field have been collected by web crawling (see Table. 1). This collection process, although fast and scalable, suffers from an imbalance in terms of skill level as summarized in Table. 3. Recent works such as EgoExo-4D [33], EgoExo-Fitness [23], LucidAc- tion [24], and Struggle [31] mark a shift in data collection from rough web crawling to physical acquisition in con- trolled environments. While this paradigm effectively mitigates limitations such as poor data quality and ac- tion homogenization, these emerging datasets still fail to fully accommodate subjects’ diverse skill levels and overlook the research gap in modeling implicit physio- logical states. Specifically, existing works often restrict their focus to self-loaded fitness actions, completely ig- noring the highly challenging and complex domain of weight training. To overcome this limitation, we make a major contribution by meticulously collecting the full, large-scale MyoMechanix dataset in-house. This is a significantly more effortful option, but it results in a much more desirable, balanced representation of skill level, allowing for the mitigation of skill-level bias. Lack of intermediate structured reasoning. Another lim- itation stemming from web-crawled datasets is that they lack an intermediate structured reasoning process be- hind the scores predicted by the models [1, 9, 19]. This lack of intermediate reasoning naturally hinders the de- velopment of causal and interpretable models. To bridge this gap, we construct the Fitness Knowledge Graph 5 Table 1 Comparison of the MyoMechanix dataset with representative AQA datasets in terms of sample size, action types, views, modalities, annotations, skill level coverage & sources. V: Video, Sk: Skeleton, sE: sEMG, P: Physiological Info., S: Score, G: Grade, A: Action, F: Formation, D: Description, E: Error, Fb: Feedback. N: Novice, Am: Amateur, Ex: Expert, Nor: Normal, Abn: Abnormal, Rehab: Rehabilitation. DatasetSampleTypeViewModalityAnnotationSkill LevelsDomainSource MIT-Dive [9]15911VSExSportWeb MIT-Skate [9]15011VSExSportWeb UNLV-Dive [1]37011VSExSportWeb UNLV-Vault [1]17611VSExSportWeb AQA-7 [10]118971VS, AExSportWeb MTL-AQA [11]1412521VS, A, DExSportWeb Waseda-Squat [12]200111V, SkENFitnessCamera Fis-V [13]50011VSExSportWeb TASD-2 [14]60621VS, AExSportWeb Rhy.Gym. [15]100041VS, AExSportWeb QMAR [16]30626V, SkS, ANor, AbnRehabCamera FR-FS [17]41711VG, AExSportWeb Fitness-AQA [18]1304931VG, AAmFitnessWeb FineDiving [19]3000521VS, AExSportWeb LOGO [20]200121VS, A, FExSportWeb FineFS [21]116711VS, AExSportWeb PaSk [22]101811VSExSportWeb EgoExo-Fitness [23]6131126VS, A, DN, AmFitnessCamera LucidAction [24]670288V, SkS, AAm, ExSportMoCap FitAQA [25]5512301VS, A, DN, AmFitnessWeb MyoMechanix7512205V, Sk, sE, PS, A, D, E, FbN, Am, ExFitnessMoCap (FKG), which organizes annotations from field experts into structured relationships linking actions, phases, key steps, error types, and corrective feedback [34–37]. This representation supports a compositional reasoning and scoring framework that enables interpretable and fine- grained quality assessment. Scientific blind spot in load-based movement model- ing and analysis. AQA datasets cover a wide range of domains, including fitness. However, current fitness AQA datasets have two limitations: 1) they largely fo- cus on self-loaded or low-load exercises/actions (e.g., [12, 23, 33])—failing to capture critical factors unique to weightlifting, such as load magnitude, joint torque distribution, and deep muscle activation; and 2) they are limited in terms of exercises they cover [12, 18]. Weight- loaded training introduces significantly greater biome- chanical complexity and injury risk than self-loaded movements, requiring precise coordination among mul- tiple joints and muscle groups under increased external forces. Weight-loaded training also has unique movement and error patterns that are not covered in self-loaded exercises. This creates a scientific blind spot in modeling and understanding human movement under load. As a result, existing approaches are insufficient to accurately assess performance quality or capture the true physi- ological and neuromuscular demands of weight-loaded exercises. Our MyoMechanix overcomes this inherent limitation by covering 20 diverse weight-loaded exer- cises. These weight-loaded exercises are also compound exercises that involve multiple body parts and joints and carry a higher risk of injury. 2.2 Action Quality Assessment Models Action Quality Assessment (AQA) focuses on automati- cally evaluating the quality of human actions from video data, with important applications in sports analytics and intelligent coaching. Recent research has achieved substantial performance gains through advances in repre- sentation learning, evaluation strategies, and multimodal modeling. In this section, we organize existing methods into four key directions: fine-grained spatiotemporal rep- resentation, objective evaluation under subjective uncer- tainty, multimodal and vision-language modeling, and interpretability. This taxonomy provides a structured view of current progress while highlighting remaining challenges, which our proposed MyoMechanix dataset and CUBIST modeling paradigm aim to address. Fine-grained action representation. Recent AQA re- search has shifted from coarse global feature learning to- 6 ward fine-grained spatiotemporal modeling. To capture the internal evolution of complex actions, mainstream methods employ temporal segmentation mechanisms or phase-aware attention modules to decompose continu- ous action sequences into semantically meaningful sub- phases. This decomposition enables precise local align- ment and fine-grained feature aggregation. Representa- tive methods include PSL [38], GDLT [39], HGCN [4], MCoRe [40], T2CR [41], and FineParser [42]. In parallel, spatial modeling has evolved toward actor-object-centric representations, which highlight the action subject while suppressing irrelevant background noise, thereby improv- ing sensitivity to subtle motion variations. Representa- tive methods include JR-GCN [43], ACTION-NET [15], Sport-Cap [44], and NS-AQA [45]. Some works also explore self-supervised learning to acquire robust pose- motion representations [18, 46, 47]. Objective evaluation strategies. To address the inher- ent subjectivity and label ambiguity in manual scoring, one line of work introduces uncertainty modeling to explicitly capture cognitive ambiguity by learning the distribution of human scores rather than predicting a sin- gle deterministic value. Representative methods include USDL [48], UD-AQA [49], and DAE [50]. Another line of work adopts contrastive regression, which leverages ex- pert demonstrations as reference anchors and estimates relative quality by measuring spatiotemporal differences between the evaluated sample and benchmark exem- plars. Representative methods include CoRe [51] and TPT [2]. In addition, NS-AQA [45] proposes a causal framework that decomposes scores into base-level com- ponents, thereby reducing subjectivity. Notably, such approaches can operate without relying on potentially biased ground-truth labels, which are influenced by hu- man judgment (e.g., bias toward splash size in Olympic diving). Multimodal and vision-language modeling. Beyond rep- resentation and evaluation, recent advances enhance information richness and interaction by incorporating multi-source modalities and vision-language alignment. Given the limitations of a single visual modality in capturing complex kinematic patterns, multimodal ap- proaches integrate heterogeneous signals such as RGB, skeleton, audio, and physiological data to improve ro- bustness. Representative methods include PISA [52], MS-GCN [53], PAMFN [54], and SkillSight [55]. Fur- thermore, vision-language frameworks introduce prior knowledge from large language models, enabling AQA systems to move beyond scalar score prediction toward natural language feedback with fine-grained error local- ization and corrective suggestions. Representative meth- ods include MTL-AQA [11], NAE [56], VATP-Net [57], and VidDiff [58]. Interpretability. Despite these advances in representa- tion, evaluation, and multimodal understanding, most existing approaches rely on end-to-end deep architec- tures and thus remain inherently black-box, limiting the interpretability of their predictions, particularly in fitness-related domains. Interpretability, however, is crit- ical for bridging the gap between theoretical research and real-world deployment. Models equipped with rea- soning and explanatory capabilities can significantly enhance transparency, reliability, and debuggability in intelligent coaching systems. Moreover, they facilitate the generation of detailed performance analysis reports, providing actionable and practical corrective feedback for practitioners [45, 59, 60]. Nevertheless, interpretable approaches for procedural fitness activity understanding remain underexplored. Lack of detailed, fine-grained mechanism-level reason- ing. Despite the growing emphasis on interpretability, interpretable modeling remains largely underexplored in existing AQA research. A major reason is that main- stream AQA datasets are predominantly constructed from web-scraped sports competition videos [1, 9, 10, 19], and their annotation schemes typically provide only a final single score or multiple referee scores. As a result, these datasets largely lack fine-grained and composi- tional breakdowns of performance scores, as well as intermediate reasoning steps that explain how different aspects of the action execution contribute to the final assessment. Although some recent studies attempt to incorporate textual priors or perform sub-action decom- position, native, authoritative, and structured process- level annotations remain largely absent. This data-level limitation further constrains progress on the modeling side, as models have limited supervision for learning interpretable, mechanism-aware assessment processes rather than relying primarily on outcome-level score prediction. To address this critical data and knowledge bottle- neck, we introduce the MyoMechanix dataset, which leverages the Fitness Knowledge Graph (FKG) as a domain-specific ontology to reformulate traditional subjective scoring into objective, rule-based, fine-grained error annotations meticulously provided by multiple do- main experts throughout the action execution procedure (see Sec. 3.4). Building upon this design, the dataset es- tablishes a compositional scoring system based on error weights. As a result, MyoMechanix is the first benchmark in AQA to simultaneously provide final macroscopic 7 scores, microscopic execution parameters, and highly interpretable intermediate reasoning step annotations. Furthermore, we introduce a new modeling paradigm – CUBIST (Compositional Ontological Reasoning Engine)— a framework for compositional reasoning over human actions that leverages the multimodal richness and struc- tured representations of FKG. Inspired by the princi- ple of decomposition–analysis–recomposition, CUBIST operates by (i) decomposing actions into structured components, (i) analyzing each component through multimodal inputs and specialized processing modules, and (i) recomposing the results to produce coherent assessments, including error attribution and corrective feedback. This design enables the model to move beyond holistic scoring toward interpretable, step-wise reasoning that more closely reflects expert analysis. 2.3 Video Question Answering Datasets The Video Question Answering (VideoQA) task involves answering questions related to corresponding videos. VideoQA datasets have evolved from simple short-video understanding to more advanced multimodal and spa- tiotemporal reasoning. Early datasets focused on high- level semantic questions about short clips [61, 62], while later datasets [63, 64] introduced longer videos and im- proved grounding for cross-modal alignment. Recent work [65] further explores complex reasoning tasks such as causal inference and human behavior analysis. How- ever, despite these advances, existing datasets still repre- sent human actions at a coarse level, lacking fine-grained annotations of posture, movement quality, and execution details. There is a lack of datasets that provide corrective feedback, limiting the ability to answer questions about how to improve movement. Most existing AQA datasets only include simple supervision signals, such as numeri- cal scores or phase labels [1, 10, 19, 20]. Although some efforts have introduced text annotations, these are often noisy, unstructured, or incomplete—particularly lacking coverage of weight-based fitness actions [23, 33]. More- over, their reliability is uncertain due to the lack of expert validation. To address these limitations, we introduce MyoMech- anix-VideoQA, a fine-grained and feedback-oriented Vid- eoQA benchmark for weight-based fitness actions. Un- like existing datasets that represent human actions at a coarse level, lacking fine-grained annotations of posture, movement quality, and execution details, MyoMechanix- Video-QA enables language-conditioned reasoning over posture, movement phases, key steps, execution errors, muscle involvement, causal relations, and corrective strategies. Built from the expert-grounded Fitness Knowl- edge Graph (FKG), it converts structured action knowl- edge into graph-grounded question–answer pairs, provid- ing reliable supervision for both diagnostic understand- ing and actionable feedback generation. The dataset supports training vision-language models to reason not only about what action is performed, but also how well it is executed, why specific errors occur, and how they can be corrected, thereby addressing the lack of fine- grained annotations, instructional feedback, and expert validation in existing VideoQA datasets. More broadly, MyoMechanix-VideoQA provides a flexible framework for developing models capable of interactive and inter- pretable reasoning, with applications ranging from AI coaching systems to general-purpose foundation models for fine-grained human action understanding. 2.4 Electromyography Datasets Electromyography (EMG) captures the electrophysio- logical signals generated during skeletal muscle contrac- tion. It effectively quantifies muscle activation levels, evaluates motor coordination across multiple muscle groups, and provides a physiological basis for detect- ing anomalous action. Regarding acquisition techniques, EMG is primarily categorized into surface electromyog- raphy (sEMG) using non-invasive skin patch electrodes and needle electromyography (nEMG) relying on micro- probes implanted into muscle tissues [6]. Due to the non-invasive nature and high portability of the former, existing EMG datasets widely adopt sEMG technology. This technology sees extensive application in advanced computational fields, including biomechanical analysis and gesture intention estimation. However, a review of current research indicates that the vast majority of datasets treat sEMG signals merely as auxiliary features for coarse-grained action classifica- tion [66–75]. They overlook the deep underlying relation- ship between electromyography signals and fine-grained action quality. In complex fitness scenarios, precise neu- romuscular control and the functional state of deep muscles provide insight into underlying neuromuscu- lar dynamics, including muscle activation patterns and coordination efficiency. These factors are some of the key determinants of whether an exercise is performed correctly. Since visual modalities alone often fail to cap- ture such internal activation patterns, incorporating EMG signals can enhance the physiological fidelity of human models and support more objective assessments of movement quality. To address this gap, MyoMechanix introduces the first benchmark dataset that systematically incorporates 8 real electromyography signals into action quality assess- ment. The dataset is collected using a sixteen-channel sEMG system operating at a high sampling rate of 2 KHz [7], which enables continuous recording of muscle activation patterns during weight-loaded fitness exer- cises. Unlike vision-only datasets, MyoMechanix cap- tures both external movement trajectories and internal neuromuscular responses, providing a more physiologi- cally grounded representation of exercise performance. In addition, strict hardware-level synchronization be- tween the sEMG and MoCap systems ensures precise temporal alignment between physiological and kinematic signals. This design supports fine-grained cross-modal analysis and enables more objective assessment of move- ment quality with improved physiological fidelity. Beyond action quality assessment, the synchronized multimodal design of MyoMechanix enables emerging cross-modal tasks such as Video-to-EMG prediction, where models learn to estimate neuromuscular activa- tion patterns from visual observations. This capability opens a pathway toward video-based surrogate sensing of muscle activity, reducing dependence on expensive and specialized EMG hardware. With promising vali- dation, MyoMechanix further supports translation be- yond fitness scenarios, offering a foundation for low-cost healthcare sensing applications such as rehabilitation monitoring, motor function assessment, remote neuro- muscular analysis, and clinically informed evaluation of movement disorders. 2.5 Muscle Activity Estimation In muscle state estimation and cross-modal analysis, some existing works infer active muscles from visual fea- tures [76]. While these methods establish useful associa- tions between visual actions and muscular involvement, they often formulate muscle estimation as a categorical prediction problem, similar to action recognition. In this setting, the model predicts muscle names rather than measured activation states. Consequently, the con- tinuous and subject-specific neuromuscular dynamics reflected in real EMG signals remain insufficiently ex- plored. To address this limitation, this work introduces the Video2EMG task, which aims to estimate continuous muscle activation patterns from visual input. Com- pared with traditional active muscle detection methods, Video2EMG represents a fundamentally different formu- lation in terms of task objective, supervision signal, and evaluation granularity. Instead of predicting discrete muscle names or binary activation states, it uses real sEMG signals as physiological supervision to model fine- grained neuromuscular activity across consecutive video frames. By aligning external visual motion with inter- nal muscle activation, Video2EMG encourages models to learn physiologically grounded representations and provides a pathway from coarse action-level recognition toward fine-grained assessment of movement quality. More broadly, Video2EMG offers a promising di- rection toward video-based surrogate sensing of muscle activity, which may reduce reliance on expensive and specialized EMG hardware in scalable sensing scenar- ios. This capability further extends the relevance of MyoMechanix beyond fitness assessment, providing a foundation for clinically relevant healthcare applications such as rehabilitation monitoring, motor function as- sessment, remote neuromuscular analysis, and clinically informed evaluation of movement disorders. 3 MyoMechanix Dataset 3.1 Motivation Motivated by the limitations identified in existing bench- marks related to AQA, VideoQA, and EMG, we intro- duce MyoMechanix, the largest and first-of-its-kind multimodal dataset for weight-loaded action assessment and coaching. As discussed in the previous section, cur- rent datasets remain limited in four key aspects: 1) they lack physiological signals that reflect implicit muscle mechanics, 2) they provide insufficient diversity across performer ability levels, 3) they offer limited support for intermediate structured reasoning, and 4) they largely overlook load-dependent movement modeling and anal- ysis. MyoMechanix is constructed to address these gaps through in-house data acquisition, synchronized multi- view and multimodal recordings, and physiological sens- ing, including sEMG, heart rate, and respiratory rate. By bridging explicit visual movement features with im- plicit muscle mechanics, MyoMechanix provides a foun- dational testbed for biomechanically grounded Physical AI. Each action is meticulously annotated by multi- ple fitness experts with comprehensive error labels and detailed actionable feedback. To support structured and interpretable reasoning, we construct a comprehensive Fitness Knowledge Graph (FKG) that provides an ontology-like representation of fitness actions. The FKG captures the procedural struc- ture of actions and the relationships between execution errors and corrective guidance by encoding actions, ac- tion key steps (AKs), fine-grained error types (ETs), and prescriptive feedback messages (FBs). Unlike existing fitness datasets that primarily focus on self-loaded or low-load actions, MyoMechanix cen- ters on weight-training actions, where external loads 9 introduce greater biomechanical complexity, novel move- ment and error patterns, neuromuscular coordination demands, and injury-related assessment challenges. This section presents the design of MyoMechanix, including the fitness action scope, data collection and annotation procedures, quality-control mechanisms, Fit- ness Knowledge Graph construction, compositional scor- ing, and benchmark design. 3.2 MyoMechanix Design Fitness actions are complex, skilled procedural activi- ties consisting of temporally ordered movement steps that require precise posture, coordination, force control, and joint alignment. Consequently, they provide a struc- tured setting for linking explicit visual motion patterns with implicit physiological and biomechanical mecha- nisms. While existing fitness-related benchmarks—such as Waseda-Squat [12], Fitness-AQA [18], and EgoExo- Fitness [23]—have established a foundation for action assessment, they focus almost exclusively on self-loaded or low-load exercises. Although these movements are accessible and safe for data collection, they fail to cap- ture the unique biomechanical complexity introduced by external resistance. Weight-loaded exercises differ fundamentally from bodyweight movements, as external loads (such as bar- bells or dumbbells) significantly alter the human move- ment profile. The introduction of resistance increases joint torque, demands deeper muscle activation, and requires more intricate neuromuscular coordination. Fur- thermore, the heightened injury risk associated with im- proper technique under load necessitates a more granular level of postural assessment than bodyweight datasets currently provide. This discrepancy creates a signifi- cant scientific gap: existing models cannot sufficiently account for the specific movement patterns and error modes unique to strength training. Current action quality assessment (AQA) datasets remain insufficient for bridging this gap due to two primary limitations. First, by emphasizing low-load ex- ercises, they overlook critical factors such as load mag- nitude and joint torque distribution under resistance. Second, the diversity of these datasets is often restricted; for instance, the Waseda-Squat dataset [12] is confined to a single exercise performed by a single subject with simulated errors rather than realistic weight-loaded con- ditions. As a result, there is a clear need for benchmarks that represent the physiological and biomechanical real- ities of professional and recreational strength training. To address these limitations, MyoMechanix focuses on weight-loaded compound exercises. These exercises Table 2 Comparison between MyoMechanix and ex- isting SOTA fitness datasets. EWL: equipment (barbell, dumbbell, etc.)-based weight loading; RI: level of risk of injury. DatasetExercisesEWLRI Fitness-AQA [18]3✓high EgoExo-Fitness [23]12✗low FitAQA [25]30✗low MyoMechanix20✓high involve multiple joints and muscle groups, impose higher biomechanical demands, and require stricter form con- trol for safe and effective execution. Given the practical constraints of dataset collection, we select 20 representa- tive weight-loaded actions according to four principles: 1) covering both upper- and lower-body muscle groups; 2) involving complex and injury-prone joints, such as the shoulders, knees, and hips; 3) including free-weight actions that require greater stabilization and carry ele- vated injury risk; and 4) focusing on widely practiced exercises to ensure broad relevance to common strength- training regimens. Through this design, MyoMechanix complements ex- isting fitness benchmarks by capturing diverse, realistic, and biomechanically challenging weight-loaded actions. This makes the benchmark intrinsically more challeng- ing, as weight-loaded action assessment requires reason- ing beyond visible pose geometry. External resistance changes movement dynamics, introduces equipment oc- clusions, amplifies subtle but safety-critical error modes, and requires models to account for implicit biomechani- cal factors such as joint torque, load distribution, sta- bilization, fatigue-induced tremors, and neuromuscular control. By focusing on these challenging movements, MyoMechanix advances fitness AQA beyond coarse ge- ometric evaluation toward fine-grained assessment of biomechanical integrity, performance quality, and in- jury risk. A detailed comparison between the proposed dataset and existing benchmarks is provided in Table. 2. Multiview Capture. The biomechanical and visual complexity of weight-loaded actions directly motivates our multiview capture design. Since action quality in strength training depends on fine-grained joint align- ment, limb trajectories, and equipment-body interac- tions, single-view observations are often insufficient for reliable visual assessment. This limitation is further exacerbated by frequent self-occlusions and equipment occlusions caused by barbells, dumbbells, and compact lifting postures. Multiview capture is therefore essen- tial for robust visual human modeling and sensing, as complementary viewpoints reduce occlusion, improve 3D pose reconstruction, and enable more reliable assess- 10 ment of subtle postural and movement errors. To obtain comprehensive ground-truth motion representations, in- cluding multi-view videos and 3D skeletal points, we use a high-precision markerless motion capture system [8]. This setup uses four synchronized ZCAM E2M4 cinema cameras [77] with Olympus 14–150 m lenses fixed at 14 m [78], capturing at 4K resolution (3840×2160 pixels) and 120 FPS. We also include a standard smartphone camera, OnePlus 7 [79], recording at 1080P resolution (1920×1080 pixels) and 60 FPS. Setup visualized in Fig. 3. Adding this smartphone viewpoint replicates typical user-recording scenarios, improving the dataset’s practical utility for real-world fitness applications while maintaining precise multimodal synchronization. Multimodal Information. Beyond visual complex- ity, weight-loaded action assessment also involves latent physiological and biomechanical factors that cannot be fully observed from videos or skeletal points alone. While multiview capture improves visual human modeling and 3D motion reconstruction, it does not directly measure muscular activation, exertion, stabilization effort, or breathing control, all of which are closely tied to action quality and safety in strength training. This motivates a multimodal sensing design that complements visual observations with physiological signals, enabling more comprehensive assessment of both external movement patterns and internal bodily responses. Existing AQA datasets primarily focus on videos [9], text [11], skeletal points [80], or audio [52], with limited attention to phys- iological data. EMG, in particular, captures muscle acti- vation patterns that are closely tied to force production, stabilization, and movement quality, providing critical information for fitness AQA and fine-grained feedback. Accordingly, the MyoMechanix dataset records three key physiological signals: 1) EMG: we use a 16-channel FastMove sEMG device [7] operating at a sampling rate of 2000 Hz to capture the target muscle groups’ activity; 2) heart rate: collected with a customized wearable vest to provide insight into exercise intensity and physiologi- cal exertion; and 3) respiratory rate: also recorded via the vest, since proper breathing is an important pro- cedural component of safe and effective weight-loaded exercise. Accordingly, our action procedures and rules incorporate breathing as an AQA criterion. Skill-Level Diversity for Robust Modeling. In ad- dition to multimodal sensing, a diverse subject pool is essential for capturing the variability of weight-loaded action quality. Since strength-training performance de- pends strongly on proficiency, strength capacity, motor control, and exercise experience, datasets constructed from narrow subject populations may fail to represent Data Collection Environment - Five Viewpoints Setup Fig. 3 The Data Acquisition Environment. We equipped four cinema cameras fixed at the four corners of the collection area, along with one smartphone camera included to simulate typical user-recording scenarios and improve in-the-wild gen- eralizability. Video, sEMG, heart rate, and respiratory rate signals were recorded synchronously during exercise perfor- mance by subjects. the full range of movement patterns and error modes encountered in real-world fitness scenarios. To address the uneven proficiency distributions prevalent in existing AQA benchmarks, we recruit participants across three proficiency levels: novice, amateur, and expert. This proficiency-balanced design reduces skill-level bias and supports computational modeling of a wide spectrum of movement patterns, error modes, and performance qualities across novice, amateur, and expert subjects. Participants were recruited through two channels: re- search institutions and local gymnasiums. During the preliminary selection phase, over 60 applications were re- ceived from students and professional coaches. To ensure reliable proficiency stratification, all applicants under- went a standardized screening protocol consisting of online questionnaires, on-site athletic performance tests, and 3D body-composition analysis. This procedure en- abled a quantitative evaluation of their strength capacity, weight-bearing ability, and motor skill level. Following previous work [23], we selected 38 subjects in total: 10 experts (IDs P01 to P10), 8 amateurs (IDs A01 to A10, excluding A06 and A09), and 20 novices (IDs N01 to N20). Unlike many existing datasets with limited profi- ciency coverage, MyoMechanix spans a broad spectrum of skill levels, as illustrated in Table. 3. This proficiency- balanced design helps reduce skill-level bias and sup- ports more robust evaluation across diverse movement patterns, error modes, and performance qualities. 11 3.3 Data Acquisition Protocol Human subject safety. To ensure participant safety during the collection of risk-prone weight-loaded ac- tions, subjects were instructed to perform each action at 80% of their one-repetition maximum. In addition, all subjects were provided with sports injury insurance, and on-site safety supervisors and emergency medical kits were available throughout the experiments. The data collection protocol was reviewed and approved by the Institutional Review Board (IRB) of Beijing Sport University (Approval No.: 2023313H). All participants provided written informed consent prior to participation. Execution Procedure. At the beginning of data acqui- sition, each subject was informed of the specific actions to be performed. Staff assisted subjects in wearing the customized vest correctly and securely attaching the sEMG sensors to the target muscles, ensuring stable skin contact to reduce signal noise and acquire high- quality physiological signals. During data collection, physiological signals were monitored to capture internal exertion and fatigue-related responses, providing impor- tant references for subsequent fine-grained annotation. Unlike prior in-house datasets such as FLAG3D [81] and EgoExo-Fitness [23], subjects entered the collection area individually to avoid observing others or receiving detailed instructions, thereby minimizing external influ- ence. No textual descriptions were provided. Instead, a brief demonstration was given before each session to en- sure that subjects performed the actions based primarily on their own experience. After sensor attachment, each subject faced the curtain between Viewpoints 1 and 2 and performed the action upon the staff’s signal. The staff then recorded the corresponding weight data, and the entire process was supervised by the authors. Furthermore, to support the modeling of the tempo- ral evolution of action quality, each subject was required to complete 10 continuous repetitions for each action. Such continuous repetition sequences are largely absent from existing datasets. 3.4Expert-Guided Data Annotation Framework Accurate annotation is essential for transforming raw multiview and multimodal recordings into a benchmark for fine-grained fitness action quality assessment. For weight-loaded exercises, annotation is particularly chal- lenging because action quality depends on phase-specific movement execution, localized postural errors, biome- chanical safety, and physiological factors such as breath- ing and fatigue. To ensure consistent and interpretable Table 3 Comparison of ability/skill level diversity of the subjects in AQA datasets. N: Novice, A: Amateur, P: Professional. DatasetNAP Sport Datasets [1, 9, 19]✓ Waseda-Squat [12]✓ Fitness-AQA [18]✓ EgoExo-Fitness [23]✓ LucidaAction [24]✓ FitAQA [25]✓ MyoMechanix✓ labels, we develop an expert-guided annotation frame- work that integrates standardized fitness guidelines, biomechanical phase decomposition, structured error taxonomy, a Fitness Knowledge Graph (FKG), and a compositional scoring system. This section first de- scribes how the annotation guidelines are constructed, then introduces the structured annotation framework, and finally details the expert annotation workflow. 3.4.1 Annotation Guideline Construction Unlike competitive sports with strictly defined action objectives and scoring criteria, such as diving or gymnas- tics, fitness actions target diverse goals such as muscle growth, fat loss, rehabilitation, and general physical conditioning. Consequently, evaluation standards for a single action may vary across individuals, training con- texts, and coaching practices. To ensure high-quality and consistent annotations for MyoMechanix, we collected existing fitness-action evaluation standards from multi- ple sources, including national standards, fitness associ- ations, professional literature, fitness applications, and domain experts. By synthesizing these sources, we estab- lished an expert-guided annotation system for weight- loaded fitness actions, which can also be adopted by future work. Rule formulation for Fitness Skill Evaluation via multisource knowledge integration. Specifically, our annotation rules drew primarily on the ”National Occupational Skill Standards—Social Sports Instruc- tor (Occupation Code: 4-13-04-01)”, jointly formulated by the Ministry of Human Resources and Social Secu- rity of the People’s Republic of China and the General Administration of Sport of China [34]. Additionally, we also reference the ”Fitness and Bodybuilding Tuto- rial” from Beijing Sport University [35], ”Occupational Competency Training Textbook for Social Sports In- structors—Fitness Coaches” published by the Human Resources Development Center of the General Adminis- 12 A K 1 9 A 0 1 A K 5 6 A 0 2 A 0 3 A 0 4 A 0 5 A 0 6 A 0 7 A K 1 0 A 0 8 A K 1 1 A 0 9 A K 3 5 A K 0 8 A 1 0 A K 3 4 A K 0 9 A 1 1 A K 3 7 A 1 2 A K 3 6 A 1 3 A K 4 0 A K 3 1 A 1 4 A K 4 1 A K 3 0 A 1 5 A K 4 2 A K 3 3 A 1 6 A K 4 3 A K 3 2 A 1 7 A 1 8 A K 0 1 A 1 9 A 2 0 A K 0 2 A K 0 3 A K 3 9 A K 0 4 A K 3 8 A K 0 5 A K 0 6 A K 0 7 A K 1 2 A K 1 3 A K 1 4 A K 1 5 A K 1 6 A K 1 7 A K 1 8 A K 2 0 A K 2 1 A K 2 2 A K 2 3 A K 2 4 A K 2 5 A K 2 6 A K 2 7 A K 2 8 A K 2 9 A K 4 4 A K 4 5 A K 4 6 A K 4 7 A K 4 8 A K 4 9 A K 5 0 A K 5 1 A K 5 2 A K 5 3 A K 5 4 A K 5 5 E T 0 1 1 A K 0 1 E T 0 1 2 A K 0 2 E T 0 1 3 A K 0 3 E T 0 5 8 A K 4 8 E T 0 2 9 A K 3 9 E T 0 1 4 A K 0 4 E T 0 5 9 A K 4 9 E T 0 2 8 A K 3 8 E T 0 1 5 A K 0 5 E T 0 1 6 A K 0 6 E T 0 1 7 A K 0 7 E T 0 5 4 A K 4 4 E T 0 2 5 A K 3 5 E T 0 1 8 A K 0 8 E T 0 5 5 A K 4 5 E T 0 2 4 A K 3 4 E T 0 1 9 A K 0 9 A K 1 0 A K 1 1 E T 0 0 1 A K 1 2 E T 0 0 2 A K 1 3 E T 0 0 3 A K 1 4 E T 0 0 4 E T 0 0 5 A K 1 5 A K 1 6 E T 0 0 6 A K 1 7 E T 0 0 7 A K 1 8 E T 0 0 8 E T 0 4 4 A K 5 4 A K 1 9 E T 0 0 9 E T 0 4 5 A K 5 5 E T 0 3 0 A K 2 0 E T 0 3 1 A K 2 1 E T 0 3 2 A K 2 2 E T 0 3 3 A K 2 3 E T 0 3 4 A K 2 4 E T 0 3 5 A K 2 5 E T 0 3 6 A K 2 6 E T 0 3 7 A K 2 7 E T 0 3 8 A K 2 8 E T 0 3 9 A K 2 9 E T 0 5 1 A K 4 1 E T 0 2 0 A K 3 0 E T 0 5 0 A K 4 0 E T 0 2 1 A K 3 1 E T 0 5 3 A K 4 3 E T 0 2 2 A K 3 2 E T 0 5 2 A K 4 2 E T 0 2 3 A K 3 3 E T 0 5 7 A K 4 7 E T 0 2 6 A K 3 6 E T 0 5 6 A K 4 6 E T 0 2 7 A K 3 7 E T 0 4 0 A K 5 0 E T 0 4 1 A K 5 1 E T 0 4 2 A K 5 2 E T 0 4 3 A K 5 3 E T 0 4 6 A K 5 6 E T 0 1 0 E T 0 4 7 E T 0 4 8 E T 0 4 9 E T 0 6 2 E T 0 6 3 E T 0 6 0 E T 0 6 1 (a)(b) Fig. 4 Visualization of the entity relationships within the Fitness Knowledge Graph. (a) Mapping between actions and action keysteps. (b) Mapping between action keysteps and error types. tration of Sport of China [36], and ”Joe Weider’s Body- building System” [37]. While several of these primary texts are localized national standards, they adhere to global training prin- ciples and biomechanics. These domestic guidelines are rigorously benchmarked against frameworks from the American College of Sports Medicine (ACSM) and the National Strength and Conditioning Association (NSCA), as well as advanced European and Australian sports science practices. Consequently, these sources provide comprehensive and broadly applicable guidance on ac- tion mechanics, posture, and safety considerations. We thoroughly reviewed all sources and meticulously documented essential information relevant to the actions collected, emphasizing target muscles, action descrip- tions, and potential error types. By carefully synthesiz- ing these details, we developed annotation guidelines that adhere to existing standards while accommodating practical training scenarios. Furthermore, these rules underwent rigorous verification by experts at Beijing Sport University, thereby ensuring theoretical soundness, consistency, and practicality in the annotation process. 3.4.2 Structured Annotation Framework Building on the expert-guided rules, we develop a struc- tured annotation framework that converts fitness knowl- edge into fine-grained, machine-readable annotations for model training, evaluation, and interpretable reasoning. Each action is decomposed into biomechanical phases and action keysteps, with corresponding phase-specific error types and corrective feedback. These components are organized into a Fitness Knowledge Graph (FKG), enabling explicit modeling of the relationships among actions, execution steps, errors, and feedback. Finally, a compositional scoring system aggregates the annotated errors into an interpretable action quality score. Biomechanical phase decomposition. To guide the annotation process and ensure consistency, a biomechani- cally-grounded framework was developed for structur- ing each action. In this framework, each collected ac- tion is decomposed into three distinct phases, based on dominant-limb movement patterns: Preparation, Con- centric, and Eccentric. Action keysteps. Within each phase, the primary limb actions are further subdivided into finer-grained action keysteps (AKs), each paired with integrated tex- tual descriptions. By systematically combining these AKs, we generate rich textual representations of diverse fitness actions, supporting fine-grained action analy- sis and structured representation learning. Building on this framework, we next introduce extensions for error identification, structured knowledge representation, and annotation efficiency. Error Taxonomy paired with Corrective Feed- back. We cataloged fine-grained error types (ETs) that 13 Preparation AK01 AK02 AK03 AK04AK05 ConcentricEccentric ET003: Excessive forward flexion of the spine. (9) ET018: Lowering too quickly. (5) Compositional Score = 1 - (6+6+5+6+9+5)/161 × 100% = 77.02 ET001: Feet positioned too narrowly. (6) ET006: Open grip. (6) ET005: Trunk swaying forward and backward, indicating instability. (5) ET008: Excessive elbow hyperextension. (6) ET006 ET001 ET005 ET003 ET018 ET008 Fig. 5 Visualization of exemplary errors and compositional scoring process during one of the actions—BarBell Overhead Press. Several key errors were observed that could compromise exercise form and effectiveness, as mentioned in the following. (a) During the Preparation phase, 1) the stance was too narrow, and 2) the grip was open rather than closed. (b) During the Concentric phase, 1) trunk swaying and 2) excessive elbow hyperextension were observed. (c) During the Eccentric phase, 1) The barbell was also lowered too quickly, 2) accompanied by excessive forward spinal curvature. The corresponding score deductions are shown in the figure, resulting in a final action score of 77.02. (Zoom in for the best view.) characterize localized joint deviations within each major limb action, pairing them with specific prescriptive cor- rective feedback. To systematically quantify the severity of these deviations, we established a hierarchical crite- rion for assigning penalty weights to each ET. These weights are determined based on their potential impact on physical safety, core stability, and movement efficacy, structured as follows: •Errors that directly compromise joint safety (e.g., spine, knee, hip, shoulder) receive the highest weight (9–10). •Errors that undermine core stability, disrupt move- ment trajectory, or alter force generation (e.g., rely- ing on momentum) receive high weight (7–9). •Errors in foundational posture (e.g., initial stance, foot spacing) and insufficient movement amplitude receive moderate weight (6–7). •Auxiliary errors related to fine control receive a lower weight (5). • Breathing errors (incorrect inhalation or exhalation) receive the minimal weight (2). Fitness Knowledge Graph Construction. To sup- port structured and interpretable reasoning, we intro- duce a novel Fitness Knowledge Graph (FKG) that formalizes weight-training exercises as biomechanically phased action representations. Each exercise is decom- posed into constituent actions organized by biomechan- ical phases, such as preparation, eccentric, transition, and concentric phases. Within each phase, the FKG encodes ordered action key steps, phase-conditioned er- ror types, and corresponding corrective feedback. This explicit semantic structure captures both the proce- dural organization of weight-training actions and the relationships between execution errors, their biomechan- ical consequences, and appropriate corrective guidance. Building on this ontology, we design a compositional, penalty-based annotation protocol that weights phase- specific errors according to their biomechanical impact, enabling interpretable and fine-grained action evalua- tion. FKG is illustrated in Fig. 4. Compositional Scoring System. Based on our error taxonomy, we designed a compositional penalty-based scoring scheme to evaluate overall action quality. Our 14 compositional scoring is a crucial contribution, provid- ing a way to account for all errors in a transparent and traceable manner. Since our compositional AQA score is built in a bottom-up approach from individual errors, it enjoys full interpretability and offers full ac- countability/verifiability. This is in stark contrast to existing approaches and practices, which only produce a final single score that lumps together all the errors in a completely opaque and untraceable manner – losing all interpretability and accountability/verifiability. This system separates the mathematical complexity from the annotation process, ensuring both rapid and consistent data labeling. Annotators are required to identify and select the observed ETs within a video clip, and the overall execution score is automatically computed by composing the respective penalties. Specifically, the overall score is built up bottom-up by calculating the ratio of the sample’s cumulative error weight to the maximum possible error weight meticu- lously predefined for that action by experts as discussed in the previous section. For each sample, the compo- sitional action quality score,S, is computed using the following compositional formula: S = [1− P N i=1 W i P M j=1 W j ]× 100(1) whereNrepresents the number of errors the sample contains,W i represents the weight of errors that the sample contains,Mrepresents the number of errors that the action type contains, andW j represents the weight of errors that the action type contains. The theoretical range ofSis [0,100]. We have visualized a detailed example of action quality assessment along with exemplary errors and the associated full compositional scoring process in Fig. 5. 3.4.3 Expert Annotation Workflow We annotate the following information for each sample: 1) action segmentation; 2) action keysteps; 3) action errors; and 4) action quality assessment scores using the following systematic procedure. The complete anno- tation process is visualized in Fig. 6. Annotator Recruitment. Due to the substantial volume of the MyoMechanix dataset and the stringent require- ments for annotation quality, 16 professionals, including fitness trainers and sports science graduate students, were recruited for annotation. To ensure consistency and scientific rigor throughout the annotation, we estab- lished in-depth collaboration with all annotators before project commencement, clarifying the annotation re- quirements, personnel qualifications, and the underlying logic of the guidelines. We then conducted a small-scale pilot annotation to validate these rules. Based on the pilot results, we refined the guidelines and finalized our annotator pool. Action Segmentation. During action segmentation, con- tinuous videos containing repeated actions were tem- porally segmented into individual samples. Annotators subsequently verified each discrete sample to ensure strict alignment with our established guidelines. Annotation Task Organization. During the formal an- notation phase, we organized the workload into distinct periods based on target muscle groups. At the onset of each period, we conducted centralized training for all 16 annotators to ensure strict adherence to the guidelines and high inter-annotator consistency. Annotator Grouping. Annotators were organized into 2 distinct groups based on the level of fitness experience: 1) Group 1 – trainers with<3 years of experience and master’s students – handled the initial annotation (Stage 1); 2) Group 2 – trainers with>3 years of experience and doctoral students – was responsible for annotation verification (Stage 2). Two-Stage Annotation. Stage 1: Each sample was in- dependently annotated for error types by 3 annotators from Group 1, utilizing the RGB video, 3D pose, and respiratory rate modalities. Stage 2: The initial sets of la- bels underwent a rigorous review by 3 senior annotators from Group 2. Verification Workflow. To effectively mitigate subjec- tive bias and label noise, we implemented a progressive conflict-resolution protocol for any identified discrep- ancies. Initially, inconsistent annotations were resolved through a majority vote among the Stage 2 annota- tors. For highly ambiguous cases, the reviewers engaged in detailed discussions until a mutual agreement was reached. In exceptionally rare and complex instances, we invited collaborating domain experts to participate in the review meetings, facilitating a final consensus to rigorously determine the ground-truth annotation. Authors periodically monitored the quality of the anno- tated data to ensure that the dataset maintained a high degree of accuracy and reliability. 3.5 Benchmark Design Building on the multimodal recordings and structured annotations in MyoMechanix, we establish a suite of 15 Segmentation Annotation Start Frame 150 Raw Videos End Frame 2780 2520 10 Repetitions of Squat 10th Repetition 2780 ... 2780Eccentric AK44 Stage1: Initial Anotation Stage2: Verification Concentric AK45 No Error Source National Standard Textbook/ Tutorial Lifestyle Book Train Annotators Learn & Cultivate Integration & Refine Fig. 6 Our Annotation Framework and Process. (Left) We integrate multiple sources of knowledge, and annotators are trained on the provided guidelines, and they receive centralized instruction to ensure full understanding of the rules. (Right Top) Raw video data was segmented following predetermined criteria. (Right Bottom) We implement a two-stage expert annotation and verification process to reduce annotation errors and mitigate subjective bias. benchmarks that extend beyond standard action quality assessment to support broader applications in fitness understanding, coaching, and physiological signal pre- diction. Specifically, we define three benchmark tasks: 1) Action Quality Assessment and Coaching, with Vanilla, Cross-Subject, Cross-View, and Mix-View set- tings; 2) Fine-Grained Video Question Answering; and 3) Video-to-EMG Prediction. The following sec- tions describe each benchmark in detail. 3.5.1 MyoMechanix-AQA & Coaching Dataset The MyoMechanix-AQA & Coaching benchmark focuses on the core structured exercise assessment, including estimating execution quality, identifying form errors, explaining score deductions, and generating corrective feedback grounded in expert-defined criteria. It is de- signed to support both training and evaluation of models that learn our rubric-grounded knowledge about how weight-loaded actions should be assessed, how devia- tions from ideal execution should be penalized, and how corrective recommendations should be derived. To support action quality assessment under varying levels of generalization difficulty, we design split proto- cols that evaluate both in-distribution performance and out-of-distribution generalization. Specifically, we estab- lish a hierarchical evaluation framework with four splits: Vanilla, Cross-Subject, Cross-View, and Mix-View. For each split, samples are partitioned into training, valida- tion, and test sets with a ratio of 60:20:20. •Vanilla. This standard protocol follows common prac- tice in AQA datasets and provides an in-distribution setting for estimating upper-bound model perfor- mance under minimal distribution shift. Within each action class, samples from different subjects and rep- etitions are aggregated and randomly partitioned into training, validation, and test sets. •Cross-Subject. This protocol partitions subjects into disjoint training, validation, and test groups to eval- uate model generalization across individuals. •Cross-View. This protocol trains models on one cam- era view and tests them on another unseen view, simulating the distribution shift caused by viewpoint variation. •Mix-View. This protocol trains models on three cam- era views and tests them on the remaining unseen view, assessing generalization to unseen viewpoints under partial multiview supervision. 16 We recommend a two-stage evaluation protocol: first, establishing an in-distribution baseline on the Vanilla split, then evaluating robustness on the Cross-Subject, Cross-View, and Mix-View splits. This progression en- ables systematic assessment of both standard AQA per- formance and generalization under real-world subject and viewpoint shifts. 3.5.2 MyoMechanix-VideoQA Dataset To develop foundational multimodal AI models capable of deep fine-grained understanding of complex actions, we introduce MyoMechanix-VideoQA, the first of its kind fine-grained video question–answering benchmark built directly on top of the dataset’s multimodal record- ings & the Fitness Knowledge Graph (FKG). Existing VideoQA datasets primarily focus on macroscopic scene understanding or coarse action recognition (e.g., identi- fying an individual’s actions or forecasting subsequent movements). In contrast, MyoMechanix-VideoQA eval- uates a model’s capacity for structured and multimodal reasoning by targeting fine-grained biomechanical evalu- ation. By addressing localized postural deviations, move- ment mechanics, and prescriptive corrective feedback, the proposed benchmark shifts the paradigm from sim- ply recognizing what is occurring to providing actionable guidance/coaching on how execution can be improved. Dataset Construction. Our FKG provides a struc- tured ontology of actions, keysteps, error types, mus- cle groups, and corrective feedback. Leveraging this graph, we design a reasoning-oriented conversational QA pipeline that supports interactive user queries about exercise execution. The pipeline progresses through four sequential stages: action recognition, action standards, action evaluation, and action scoring, as shown in Fig. 7. Machine-Generated, Human-Verified QA Con- struction. Guided by this pipeline, a large-scale set of natural-language QA pairs was generated utilizing a mix- ture of rule-based templates and Large Language Mod- els (LLMs). Specifically, complex responses for action evaluation were synthesized by DeepSeek-V3-Reasoning, which integrated video context, keysteps, errors, and feedback suggestions to ensure semantic coherence. Since every question is strictly grounded in explicit nodes of the FKG, the generated QA pairs maintain high biome- chanical precision. Following the generation phase, all QA sets underwent rigorous manual verification and correction, resulting in 30,048 high-quality QA pairs. The questions span four core reasoning categories: •Descriptive: Inquire about observable attributes such as the action being performed or the current keystep (e.g., “Which action is shown in the video?”, “What keystep is the performer currently executing?”). •Relational reasoning : Require linking multiple graph entities (e.g., “Which muscle group is primarily ac- tivated during the descent phase of a squat?” or “Which error in the setup step leads to a lower-back penalty?”). • Temporal : Probe the understanding of sequential dependencies (e.g., “Which key step follows the liftoff phase?”). • Causal feedback : Demand corrective advice based on observed biomechanical errors (e.g., “What feedback should be given if the knees cave inward during the ascent?”). Comprehensive multimodal reasoning platform. By transforming the structured annotations of the Fit- ness Knowledge Graph (FKG) into a large set of graph- grounded question–answer pairs, MyoMechanix-VideoQA evolves beyond a conventional AQA dataset into a comprehensive multimodal reasoning platform. It is de- signed to support the development of foundation vision- language models (VLMs) for language-grounded, query- conditioned, fine-grained action understanding. Rather than requiring fixed assessment outputs, VideoQA ex- poses models to diverse natural-language queries span- ning action identity, movement phases, key steps, pos- ture errors, muscle activation, temporal dependencies, causal relationships, and corrective strategies. This for- mulation enables supervised fine-tuning of VLMs that must not only recognize human actions but also rea- son over the hierarchical and causal structure linking actions, execution errors, underlying physiology, and feedback. Experimental results highlight the value of such structured graph supervision in enabling VLMs to learn representations of action quality that extend well beyond surface-level visual understanding. More broadly, MyoMechanix-VideoQA provides a flexible framework for training models capable of interactive and inter- pretable reasoning, with applications ranging from AI coaching systems to general-purpose foundation models for fine-grained human action understanding. 3.5.3 MyoMechanix-Video2EMG Dataset Although visual modalities, such as RGB video and 3D poses, effectively capture the macroscopic kinematic features of human action, they inherently lack informa- tion regarding the internal activation of muscles, the exertion of force, and physiological cost. To bridge this information gap and expand the boundaries of research in multimodal action analysis, we propose Video2EMG, a novel cross-modal translation benchmark. 17 Q1: Tell me what fitness action the person in the image is performing. Q4: Based on your evaluation, assign a score to the person's action in the image on a scale of 0–100. Answer1 Answer2 Answer3 Answer4 Sample Annotator Annotation Q2: Based on the identified fitness action, please explain the standard for this action according to your understanding. Reasoning Process Deepseek V3 Q3: Using your standard, please evaluate the person's performance in the image according to the action standards you deem appropriate. Action Type Error Type Feedback Score Task 1 "from": "human", "value": "<video> What exercise move is being performed by the individual in the image?", "from": "gpt", "value": "The person is performing a Squat." Task 4 "from": "human", "value": "Based on your assessment, what score out of 100 would you give the person‛s performance in the image?", "from": "gpt", "value": "85.82" Task 3 "from": "human", "value": "<video> Please assess the person‛s performance in the image using the criteria you consider appropriate for that action.", "from": "gpt", "value": "During the squat, the individual demonstrates a tendency to ... enhance performance, and minimize the risk of injury." Task 2 "from": "human", "value": "Given the fitness move shown, describe the correct execution standards for it.", "from": "gpt", "value": "The rule of Squat is: Stand with feet slightly wider than shoulder-width, toes slightly turned outward. Keep the chest lifted, engage the core, maintain a neutral spine, ... hips and knees, moving the hips forward until the knees are straight or slightly bent." Integration & Refine Fig. 7 The construction of the MyoMechanix-VideoQA dataset. 100 150 200 250 300 350 400 100 150 200 250 300 350 400 A 0 1 A 0 2 A 0 3 A 0 4 A 0 5 A 0 6 A 0 7 A 0 8 A 0 9 A 1 0 A 1 1 A 1 2 A 1 3 A 1 4 A 1 5 A 1 6 A 1 7 A 1 8 A 1 9 A 2 0 40 50 60 70 80 90 S a m p l e N u m b e r (a) D u r a t i o n S c o r e 0102030405060708090100 0 200 400 600 800 1000 1200 1400 C o u n t Score (b) E T 0 0 4 E T 0 1 8 E T 0 0 5 E T 0 0 8 E T 0 4 2 E T 0 0 3 E T 0 0 2 E T 0 1 1 E T 0 3 6 E T 0 3 8 E T 0 2 2 E T 0 2 1 E T 0 0 1 E T 0 2 0 E T 0 5 7 E T 0 0 6 E T 0 4 7 E T 0 3 3 E T 0 4 6 E T 0 1 3 0 1000 2000 3000 4000 5000 6000 7000 C o u n t Fig. 8 Data diversity and distribution in My- oMechanix. (a) Number of samples, average duration, and scores for the 20 actions. (b) Overall score distribution of the dataset and the 20 most frequent error types. Focusing on the direct prediction of continuous sEMG signals from video, this benchmark leverages the strictly paired visual and physiological signals within the dataset. This approach facilitates the modeling of the complex, nonlinear mapping between external representations of motion and internal biomechanical states. This task of- fers significant practical value in scenarios where the deployment of wearable sensors is constrained, such as competitive sports and rehabilitation therapy, as it enables unobtrusive assessment of target muscle engage- ment and fatigue. 4040303540404040403540404040304040404040806050407060504050506050506050505050 3030303040304040303030303030303030403030804050406060403050506050404050405040 3040404040404040403030404040403040404040706050508060505050806070406060606060 3040404040404060403030406040402040404040706050508080505050906070406060606060 2020203030302030202030304030302030303020603040305040303030406050304040605060 2020303030302030303030304030302030303020603040406040303040506050505050605060 2030303040304040303040304040304040303030705050507060604050606050505050805060 3020302040303020402030202030303030203020403040304040303030404040304030203040 2020202020202020302020202020202020202020302030204030303030303030303030303040 2020302020202020202030302030203020202020402040304030303030304030303030503040 555 7.5555 7.555 7.55555555 7.55 107.5105 1010105 107.55 107.57.57.5107.5 12.5 12.512.57.57.51515 12.510107.5 12.5101510101010107.57.525151015201515101520202015151520 12.515 12.512.5105 12.51515 12.5101015151515 12.512.512.51010102515201525 12.5155 20152015151015301525 5555 7.57.55 7.555 7.57.57.57.5555 7.555 107.5107.510107.57.57.51025107.510 12.512.520 12.5 7.510 12.57.5 12.512.510107.51015 12.5151510 12.510107.57.52515201530202010 12.5202525152015252020 7.510 12.57.5 12.512.510107.51015 12.5151510 12.510107.57.5252020203020201020252525202015302025 7.5555 7.555 7.555 7.57.57.555555 7.55 105 7.57.5 12.57.5105 7.555 105 107.57.57.57.5 5 7.555 7.555555 7.510105555555 12.55 7.57.5107.57.57.57.55 7.5107.57.57.57.57.57.5 5 7.57.57.510107.51055 7.5107.5107.55 7.55 7.55 12.57.5107.520 12.512.57.510 12.512.5107.57.57.5107.5 12.5 5 7.555 7.555 7.555 7.55 7.55555555 12.55 7.57.5107.57.57.57.55 7.5107.57.57.57.57.57.5 N01 N02 N03 N04 N05 N06 N07 N08 N09 N10 N11 N12 N13 N14 N15 N16 N17 N18 N19 N20 A01 A02 A03 A04 A05 A07 A08 A10 P01 P02 P03 P04 P05 P06 P07 P08 P09 P10 A18 A17 A16 A15 A10 A09 A08 A03 A02 A01 Subject Barbell Action 20.00 30.00 40.00 50.00 60.00 70.00 80.00 90.00 Weight (KG) N01 N02 N03 N04 N05 N06 N07 N08 N09 N10 N11 N12 N13 N14 N15 N16 N17 N18 N19 N20 A01 A02 A03 A04 A05 A07 A08 A10 P01 P02 P03 P04 P05 P06 P07 P08 P09 P10 A20 A19 A14 A13 A12 A11 A07 A06 A05 A04 Dumbbell Action 5.000 10.00 15.00 20.00 25.00 30.00 Weight (KG) Fig. 9 Distribution of weight used by subjects across actions. MyoMechanix includes 20 weight-loaded fitness ac- tions, evenly split between barbell and dumbbell exercises. For barbell actions, the intrinsic weight of the barbell bar, 20 kg, is included in the calculation. For dumbbell actions, the reported value corresponds to the weight of a single dumbbell, with the total bilateral load being twice that value. In the figure, the X-axis represents subjects and the Y-axis represents fitness actions. Darker colors indicate heavier weights. 3.6 Dataset Statistics 3.6.1 Fundamental Attributes After data cleaning and annotation verification, the MyoMechanix dataset contains 7,512 action samples collected from 38 subjects across three ability levels, covering 20 distinct weight-loaded fitness actions. We 18 Table 4 Comparison of textual features between CoT-AFA and MyoMechanix-VideoQA (MyoMech-VQA). Metric CategoryMetricStatisticCoT-AFA [82]MyoMech-VQA Corpus Volume Average Words per SampleWord Count102.19192.27 Std of Words per SampleStd Dev–29.09 Average Sentences per SampleSentence Count5.2510.05 Lexical & Syntactic Complexity Total Vocabulary SizeToken Count31432008 Type-Token RatioRatio–0.0014 Avg Sentence LengthWords–19.14 Flesch-Kincaid Grade LevelGrade–13.78 Flesch Reading EaseScore–30.35 Reasoning & Actionability Avg Reasoning StepsPer Sample0.911.85 Avg Actionable SuggestionsPer Sample0.758.41 05101520253035404550 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 Distribution of Word Counts per Sentence 4681012141618 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 0.16 0.18 0.20 0.22 0.24 Distribution of Sentence Counts per Sample A01A02A03A04A05A06A07A08A09A10A11A12A13A14A15A16A17A18A19A20 150 155 160 165 170 175 180 185 190 195 200 205 210 215 220 A v g W o r d s Action Type 05101520 0 200 400 600 800 1000 1200 1400 F r e q u e n c y Actionable Suggestions 0123456 0 500 1000 1500 2000 2500 3000 F r e q u e n c y Reasoning Steps controlstabilitytrunk/coreshoulderknee/legalignmentelbowhead/neckgripwrist 0.0 0.5 1.0 1.5 2.0 2.5 3.0 A v e r a g e M e n t i o n s p e r S a m p l e Error Type Fig. 10 Linguistic complexity and fine-grained reason- ing in MyoMechanix-VideoQA. This figure characterizes the linguistic richness of the dataset through six distributions: (1) words per sentence, (2) sentences per sample, (3) average word counts across the 20 action categories, (4) frequency of actionable suggestions, (5) number of reasoning steps, and (6) average frequency of error-type keywords per sample. conducted statistical analyses of the per-action sample counts, sample durations, and action quality scores, as shown in Fig. 8. The sample counts are well balanced across actions, ranging from 370 to 380. The average clip duration is 233.76 frames, with individual clips spanning 100 to 1000 frames. Action quality scores have a mean of 68.15 and exhibit an approximately normal distribution, with values ranging from 20 to 100. This broad score range, together with the moderate mean, indicates substantial diversity in skill levels within the dataset. We also report the 20 most frequent error types. Finally, Fig. 9 presents the specific weight loads used by each subject. 3.6.2 Linguistic, Semantic and Reasoning Complexity In addition to quantitative statistics on dataset scale, video duration, and action scores, we conduct a sys- tematic analysis of the structural and distributional characteristics of the MyoMechanix dataset across tex- tual features, as summarized in Table. 4 and Fig. 10. This analysis helps characterize the linguistic complexity of the dataset and the challenges it poses for multimodal reasoning and fine-grained action understanding. In terms of corpus volume, we examine the word count and sentence count of the textual annotations. All text is converted to lowercase, and standard punc- tuation marks are removed using regular expressions before tokenization based on whitespace. The results show that the average word count across all 7,512 sam- ples is 192.27, with a standard deviation of 29.09. In addition, segmenting the text using terminal punctu- ation marks, including periods, question marks, and exclamation points, shows that each sample contains an average of 10.05 sentences. In terms of lexical & syntactic complexity, we evaluate vocabulary richness and sentence-level struc- tural characteristics. Extracting unique tokens across all samples yields a total vocabulary size of 2,008 for the entire dataset. The resulting Type-Token Ratio (TTR) is 0.0014. Since TTR is highly sensitive to corpus size, this low value should be interpreted in conjunction with the dataset’s large total word count. Together, these statis- tics suggest a thematically concentrated vocabulary cen- tered on the specialized domain of fitness and exercise. Furthermore, the average sentence length across samples is 19.14 words. This relatively long sentence length re- flects the dataset’s tendency toward detailed descriptive expressions and may indicate increased syntactic and informational complexity. To further characterize reading difficulty, we compute the Flesch–Kincaid Grade Level (FK Grade) and the Flesch Reading Ease (FRE) score. Both metrics are derived from weighted counts of words, sentences, and syllables. The dataset obtains an FK Grade of 13.78, corresponding approximately to the reading level of a first-year college student in the United States. Its FRE 19 score is 30.35 on a standard scale from 0 to 100, where lower values indicate greater reading difficulty, placing the text in the “difficult” category. Taken together, these two metrics suggest that the textual annotations in MyoMechanix are linguistically demanding and require a relatively high level of reading comprehension. For reasoning & actionability, we adopt a keyword- matching approach to characterize the presence of causal reasoning and actionable guidance within the Chain- of-Thought (CoT) annotations. This strategy is used instead of more complex semantic analysis because the evaluation texts exhibit highly consistent expression pat- terns. Causal explanations frequently rely on explicit conjunctions, while actionable suggestions commonly employ imperative constructions or modal verbs. Ac- cordingly, this keyword-based methodology provides a simple, efficient, and reproducible proxy for analyzing reasoning- and action-oriented textual patterns. For the statistical analysis of reasoning-related ex- pressions, we retrieve and count causal, inferential, and explanatory keywords, including “because,” “due to,” “thus,” “hence,” “indicating,” and “suggesting.” The results show that each sample contains an average of 1.85 explicit reasoning steps. These lexical cues suggest that the annotations often extend beyond surface-level description to include deeper causal reasoning. For ex- ample, the phrase “which reduces” in the sentence “grip was open rather than closed, which reduces control” explicitly links an observed movement pattern to its functional consequence, forming a causal explanatory relation. For the statistical analysis of actionability, we cal- culate the frequency of directive, corrective, and effect- oriented keywords, including “adjust,” “aim for,” “to correct,” “avoid,” “will enhance,” and “will reduce.” The results show that each sample contains an average of 8.41 actionable recommendations grounded in the specific erroneous movement patterns observed in that sample. These cues indicate that the annotations provide con- crete advice for performance improvement. An example of such executable guidance is the instruction to “Keep the head firmly against the bench throughout...”. Collectively, the dataset’s causal reasoning content and actionable recommendations highlight the value of MyoMechanix for developing and evaluating models that can identify erroneous movement patterns, reason about their consequences, and generate precise, sample- grounded feedback for improvement. Overall, the dataset exhibits notable strengths across multiple dimensions, including textual scale, structural complexity, and se- mantic depth. These characteristics make MyoMechanix a more demanding benchmark for assessing the fine- grained multimodal understanding and reasoning capa- bilities of existing models. 4 Our CUBIST Modeling Paradigm Existing AQA paradigms are largely monolithic black- box models: they produce a single overall score, offering limited fine-grained error attribution or explanation. In this sense, they resemble Impressionism in art history, which emphasizes an overall impression over detailed structure. By contrast, Cubism offers a more analyt- ical perspective: it decomposes complex objects into smaller components, examines each part from multiple viewpoints, and recombines them to reveal underlying structure and relationships. Motivated by these principles, we propose CUBIST (Compositional Ontological Reasoning Engine), a novel modeling paradigm for compositional reasoning over human actions that leverages rich MyoMechanix-FKG and multimodal data. CUBIST moves beyond the cur- rently prevalent single overall score paradigm toward interpretable, step-wise reasoning that closely reflects expert analysis. CUBIST can be understood as a language for analyz- ing human actions, providing structured, high-resolution representations that support detailed modeling of move- ment, technique, and coaching processes. Specifically, human actions are first identified and routed to activity-specific expert pathways (Stage I). Guided by the resulting activity prior and the under- lying ontology, the model then performs fine-grained, phase-aware activity decomposition and error reasoning over the action sequence (Stage I). These interpretable error probabilities are subsequently synthesized through an explicit compositional rule-based scoring engine to produce the final action quality assessment report and its associated coaching interpretation (Stage I). The full CUBIST model is visualized in Fig. 11. This progressive and compositional design is in- tended to reflect aspects of expert reasoning and refer- eeing practice, where judgments are formed by breaking down complex actions, inspecting them from multiple viewpoints, and synthesizing the resulting evidence. By grounding model outputs in an underlying ontology, CUBIST enables fine-grained attribution of both inter- mediate reasoning and final assessments. As a result, CUBIST provides a framework for im- proving interpretability in Action Quality Assessment (AQA), supporting transparency at both the reasoning and output levels. Now, in the following, we present each stage in detail. 20 Stage I: Action Perception-driven Expert Routing 퓠 pool Pose Vision 010000200003000040000 −150 0 150 010000200003000040000 −150 0 150 sEMG Pose Encoder Vision Encoder sEMG Encoder Attention Aggregation Global Feature 퓕 global Action Classification Action Logits � 퓨 action ∁ Expert 20 Expert 1 Expert Routing Predicted Errors Action Type Scoring Engine 풮=1− ∑ 푖=1 푁 푊 푖 ∑ 푗=1 푀 푊 푗 × 100 85/100 Stage I: Compositional Score Engine ∁ Pose sEMG 퓛 ASL 퓠 phase 퓠 phase Error Feature 퓕 error Error Query Decoder Diagnostic Prediction Error Probability 풑 n ∁ Temporal Adapter Temporal Feature 퓩 Implicit Phase Parser Phase Feature 퓟 Vision Encoder Joint Feature 퓒 Ontological Error Branch PreparationConcentricEccentric E T 0 6 1 E T 0 0 2 E T 0 0 7 E T 0 0 2 E T 0 5 2 E T 0 3 2 Stage I: Ontology-driven Error Reasoning Expert 16 퓛 cls 퓨 action Fig. 11 Illustration of proposed CUBIST models’ architectures (multimodal version) for the AQA task. 4.1 Stage I—Action Perception-driven Expert Routing In human expert action quality assessment, accurate identification of the action type constitutes the initial step of the evaluation process. Similarly, Stage I of the CUBIST model functions as a high-precision routing module that maps each input video to the appropriate action-specific expert network according to its predicted action category. This stage is built upon a spatiotempo- ral video encoderφ, which serves as the backbone of the CUBIST model. The encoder supports reliable action- category recognition while establishing a robust feature representation for the subsequent stages of fine-grained error reasoning. Specifically, we define the input video sequence as V ∈R B×T×C×H×W , whereBdenotes the batch size, Tis the number of sampled frames,Cis the number of input channels, andH=Wdenotes the spatial res- olution. The backbone partitions the input video into non-overlapping spatiotemporal tubelets with temporal sizet s =2 and spatial patch sizep=16. After tubelet partitioning and linear projection, the video is encoded into a one-dimensional token sequence: H = φ(V)∈R B×L×D ,(2) whereL=⌊T/t s ⌋×(H/p) 2 is the total number of tokens, and D is the embedding dimension. Since the 3D positional encoding of the pre-trained backbone is originally defined for a 16-frame input, cor- responding to a positional grid of (8,14,14), CUBIST adopts trilinear interpolation to extend the positional encoding to (51,14,14). This adaptation enables the backbone to process long video sequences of up to 103 frames while preserving the semantic priors encoded in the pre-trained weights as much as possible. To aggregate the high-dimensional token sequence Hinto a compact global representation, we introduce an attention-based pooling mechanism with a learnable query. Specifically, we define a learnable query vector Q pool ∈R 1×D , which attends to the entire token se- quence through multi-head attention: F global = Squeeze (LN (MHA(Q pool , H, H)))∈R B×D , (3) whereMHA(·) denotes an 8-head multi-head atten- tion module,LN(·) denotes layer normalization, and Squeeze(·) removes the singleton query dimension. Com- pared with simple mean pooling, this attention-based pooling mechanism can adaptively assign higher ag- gregation weights to more informative spatiotemporal regions. The resulting global featureF global is then passed through a lightweight multi-layer perceptron classifica- 21 tion head to produce action-category logits: ˆ Y action = MLP cls (F global )∈R B×C action ,(4) where C action denotes the number of action categories. During optimization, Stage I is trained using a cross- entropy loss with label smoothing: L cls = CE( ˆ Y action , Y action ; ε=0.1),(5) whereεis the smoothing coefficient. Label smoothing mitigates over-confidence on hard training labels and improves the generalization ability of the action classifi- cation module. 4.2 Stage I – Ontology-driven Error Reasoning Drawing on the refined analysis logic of human ex- perts after establishing the action category, the CU- BIST model uses the action category output from Stage I as a heuristic prior. It feeds the video spatiotemporal features into the ontology-driven reasoning module to decode the fine-grained error types within the action sequentially. Since the fitness knowledge graph defines highly heterogeneous error-evaluation branches, we inte- grate the mixture-of-experts (MoE) architecture with the ontology-driven concept. Specifically, this stage up- grades the traditional feed-forward neural network into a specific mixture of experts architecture. It instanti- ates and maintains an exclusive expert network module E a for each action category. This module predicts and outputs theN a dimensional binary error labels of the corresponding action branch in the fitness knowledge graph. 4.2.1 Temporal Adapter The high-dimensional token sequenceH ∈R B×L×D output by the backbone exhibits entangled spatial and temporal information. To effectively recover its native spatiotemporal topology and build a compact pure tem- poral representation, we design a temporal adapter module. This module first reshapes the input sequence along the spatiotemporal dimensions into a 4D tensor H grid ∈R B×T orig ×S×D , whereT orig =⌊T/t s ⌋represents the original number of time steps andS= (H/p) 2 repre- sents the number of spatial patches contained in a single time step. Subsequently, the adapter performs a global average pooling operation on the spatial dimension to eliminate intra-frame spatial redundancy, yielding the intermediate feature representation ̃ H. Building on this, it utilizes adaptive average pooling to compress the tem- poral dimension to a target fixed lengthT ′ , thereby obtaining the initial temporal feature Z (0) . To strengthen the capacity of this feature to model local temporal dependencies, this module cascadesN conv depthwise separable convolution blocks with residual connections after the pooling layer. A single convolution block sequentially encapsulates depthwise convolution, pointwise convolution, the GELU activation function, and layer normalization to ensure efficient gradient flow. The specific feature evolution process operates as follows: Z (i+1) = LN Z (i) + GELU PW DW Z (i) , i = 0,...,N conv − 1 (6) whereDW(·) andPW(·) denote the depthwise separa- ble convolution and pointwise convolution operations, respectively. After the aforementioned multi-level tem- poral modeling, the network outputs the final temporal representation denoted as Z =Z (N conv ) ∈R B×T ′ ×D . 4.2.2 Implicit Phase Parser From the kinematic perspective, complex sports actions naturally consist of several continuous phases. To enable the network to autonomously discover and capture such temporal structures without manual phase boundary annotations, we introduce an implicit phase parser mod- ule. This module definesKindependent and learnable phase query vectorsQ phase ∈R K×D . Then, it utilizes a cross-attention mechanism to perform selective focus- ing on the temporal featuresZ, achieving data-driven semantic-level temporal segmentation: P raw , W phase = MHA(Q phase , Z, Z)(7) whereP raw ∈R B×K×D represents the aggregated raw phase representation, andW phase ∈R B×K×T ′ is the cross-attention weight matrix. To further strengthen the representational capacity of the features, the raw phase representation passes through a residual connection and enters a feed-forward neural network for nonlinear refinement: P = LN ˆ P + FFN( ˆ P) , ˆ P = LN(P raw +Q phase ) (8) Through this cascaded computation, the final phase featureP ∈R B×K×D encodes the deep semantic rep- resentations of theKimplicit motion phases on the complete action timeline. 4.2.3 Phase-Aware Error Query Decoder In the implementation of Stage I, the error query de- coder instantiates an independent learnable query vector Q err ∈R N a ×D for each category of fine-grained action error requiring detection. Considering the occurrence 22 mechanism of action execution errors, their judgment relies heavily on the continuous dynamic temporal flow evolution and remains constrained by explicit motion phase semantic boundaries. Based on this physical prior, we concatenate the temporal featureZand phase fea- turePalong the sequence dimension. This creates a joint context matrixC= [Z;P]∈R B×(T ′ +K)×D that possesses phase-aware capabilities. Subsequently, each error query vector uses a cross-attention mechanism to perform adaptive retrieval on this joint context. This process autonomously anchors the temporal intervals and phase semantics that are highly correlated with the corresponding error type, finally outputting the decoded error feature representation F error ∈R B×N a ×D . ThenF error enters a feed-forward neural network with a residual connection for nonlinear refinement. A linear projection layer maps it into scalar predicted logits for various types of errors. The sigmoid activa- tion function transforms these logits into final predicted probabilities: p n = σ W ⊤ head · LN(F error + FFN(F error )) , n = 1,...,N a (9) whereσ(·) denotes the sigmoid activation function and W head ∈R D denotes the linear projection weight. For network optimization, since the fine-grained error detection is inherently a multi-label classification task with high-dimensional, sparse positive examples. There- fore, Stage I introduces an asymmetric loss function to guide gradient backpropagation: L ASL =− 1 N a N a X n y n (1− p n ) γ + logp n + (1− y n ) ˆp γ − n log(1− ˆp n ) (10) whereˆp n =max(p n − m,0) represents the predicted probability of negative samples after introducing a prob- ability margin offset mechanism andmis the truncation margin parameter. The termsγ + andγ − control the gradient focusing weights of the model on positive and negative samples, respectively. This asymmetric penalty mechanism effectively suppresses the absolute dominant role of massive simple negative samples on the global optimization gradient, significantly improving the detec- tion performance and generalization robustness of the model for rare motion errors. 4.3 Stage I – Compositional Scoring Engine After completing fine-grained error analysis, human ex- perts synthesize the identified errors into a final action- quality score. As the concluding component of the evalu- ation framework, Stage I emulates this scoring process through a deterministic neuro-symbolic module. Rather than learning a direct correlational mapping from vi- sual features to quality scores, Stage I derives the final score through an explicit scoring rule grounded in the annotation ontology. Specifically, it receives the multi-dimensional fine-grained error probability distri- bution predicted by Stage I and applies the annotation- defined error weights (see Sec. 3.4) together with the scoring formula in Equation 1. Through this composi- tional rule-based transformation, high-dimensional error probability representations are mapped into a final quan- titative score for action quality. This stage introduces no additional learnable parameters and does not require gradient-based optimization. Identical Stage I error probabilities therefore produce the same score under the fixed scoring rule. As a result, the final score is directly attributable to interpretable error concepts and their corresponding annotation-defined penalties, improving transparency at the output level. 4.4 Multimodal Fusion To fully exploit the heterogeneous information collected in MyoMechanix, CUBIST integrates RGB video, 3D pose, and sEMG through a modular late-fusion strat- egy. Direct end-to-end optimization of heterogeneous modalities may lead to cross-modal gradient interference, potentially hindering the learning of modality-specific discriminative representations. In our pilot studies, we observed suboptimal performance that may be asso- ciated with this issue. To mitigate it, we introduce a feature-disentangled late-fusion design that allows each modality to develop task-relevant representations before their integration, while limiting additional architectural and optimization complexity. The fusion procedure is organized as follows. –For action recognition, modality-specific encoders process the 3D pose and sEMG inputs to produce a compact non-visual multimodal feature, denoted asF multi . In parallel, the RGB stream generates a global representationF global . These modality-specific features are concatenated along the channel dimen- sion to form a joint action representation, which is then passed to the multi-layer perceptron classifica- tion head,MLP cls , to estimate the action-category probability ˆ Y action . –For fine-grained error detection, multimodal fu- sion is performed after task-specific feature abstrac- tion has been established. The Phase-Aware Error Query Decoder first derives an error-oriented repre- 23 sentationF error from the RGB stream by allowing the query vectorQ err to attend to the temporal fea- turesZand phase featuresP. This representation is then integrated with the synchronously extracted 3D pose and sEMG features to construct a joint error representation. The fused feature is passed through a feed-forward neural network with a residual con- nection, followed by the final linear prediction head W head , to estimate the probability of error occur- rence. This late-fusion mechanism avoids prematurely forc- ing heterogeneous modalities into a shared representa- tion space, thereby helping to preserve their modality- specific discriminative characteristics and reducing the risk of cross-modal gradient interference. It also enables complementary reasoning across modalities. For exam- ple, RGB video captures appearance and movement context, 3D pose characterizes the spatial configura- tion of the body, and sEMG reflects underlying muscle activation dynamics. Their integration is particularly valuable near ambiguous action boundaries or for sub- tle erroneous exertion patterns that may not be fully characterized by any single modality alone. As a re- sult, the proposed strategy leverages the complementary strengths of all three modalities while requiring only minimal architectural modification. 4.5 Cross-Action Shared Skill Modeling for Expert Adaptation MyoMechanix is a large-scale dataset comprising diverse yet related action categories. This structure provides a valuable opportunity to learn skill representations that are shared across actions and transferable to down- stream expert models. To capitalize on this property, we adopt a cross-action representation learning strategy that first captures robust and transferable spatiotempo- ral cues of action quality/skills, and then adapts them to action-specific expert models. In the first phase, a shared backbone network is trained jointly across all action categories to learn spatiotemporal skill represen- tations that generalize across actions. In the second phase, action-specific expert models are trained inde- pendently. To accelerate convergence and preserve the transferable knowledge acquired in the first phase, each expert backbone is initialized with the weights of the converged shared model and optimized using a pro- gressive unfreezing strategy. Specifically, the backbone parameters are initially frozen, and only the newly intro- duced downstream modules are trained, including the temporal adapter, implicit phase parser, and error query decoder. After a predefined warm-up period, several up- per Transformer blocks of the backbone are unfrozen and jointly fine-tuned with the downstream modules using a reduced learning rate. Both training phases use the AdamW optimizer together with a cosine-annealing learning rate scheduler. 5 Experiments To systematically validate the central claims of this work, we conduct comprehensive experiments on the three benchmark tasks introduced in Sec. 3.5. These eval- uations are designed not merely to report task-specific performance, but to examine whether the MyoMechanix ecosystem enables a richer and more challenging form of action understanding than conventional vision-centric benchmarks. In particular, we investigate three comple- mentary questions: (i) whether multimodal physiological sensing improves the assessment of action quality be- yond external motion cues alone; (i) whether structured annotations support fine-grained, semantically grounded reasoning about execution errors and corrective feed- back; and (i) whether latent neuromuscular dynamics can be inferred, at least partially, from observable visual motion. The remainder of this section is organized as follows: –First, we benchmark Action Quality Assessment under both standard and disjoint splits, evaluating existing methods and quantifying the contribution of multimodal fusion (Sec. 5.1); –Next, we assess structured multimodal reasoning through the Video Question Answering bench- mark, which targets action recognition, fine-grained error diagnosis, and corrective feedback generation (Sec. 5.2); – Finally, we investigate cross-modal physiological in- ference through the novel Video2EMG task, ex- amining the feasibility of predicting muscle acti- vation patterns directly from visual observations (Sec. 5.3). 5.1 AQA & Coaching As the cornerstone task within the MyoMechanix ecosys- tem, establishing performance benchmarks for AQA is paramount. To comprehensively measure the inher- ent challenges of the dataset and thoroughly analyze the representational capabilities of each modality, we systematically constructed a comprehensive evaluation framework that includes unimodal baseline experiments, multimodal fusion experiments, and diverse data split- ting protocols. We also compare our CUBIST model against SOTA models. 24 5.1.1 Baseline methods Vision Branch. Considering that vision features serve as the paramount core modality in the AQA field, we ad- hered to previously outlined research trends when estab- lishing benchmarks. We selected 7 representative SOTA heterogeneous network architectures—including I3D- MLP [48], CoRe [51], TPT [2], HGCN [4], MCoRe [40], T2CR [41], and DAE [50] – to ensure the comprehensive and objective evaluation. More information on these models has been provided in the Appendix. Pose Branch. Apart from visual features, skeletal pose—as a modality that accurately captures human kinematic trajectories—plays a significant role in AQA. Therefore, we selected and incorporated multiple SOTA representative models – including MS-G3D [83], CTR- GCN [84], ST-GCN++ [85], and SkateFormer [86] – as evaluation baselines to systematically assess the gain value of this modality within the MyoMechanix-AQA. More information on these models has been provided in the Appendix. EMG Branch For the EMG modality, we designed a specialized branch to process continuous natural signals. To mitigate the inter-subject variance inherent in raw physiological data, a comparative model is incorporated to explicitly compute the relative activation contribu- tions of distinct muscles during the execution of an ac- tion. This procedure distills the raw data into a robust representation of muscle coordination. Subsequently, an MLP is utilized to project these low-dimensional compar- ative features into a high-dimensional embedding space, thereby facilitating seamless integration with macro- scopic visual or skeletal features. 5.1.2 Implementation Details For data preparation, we segment each full action video into single-action clips based on annotated keyframes, uniformly sampling 103 frames per clip. Given the strict synchronization across modalities in the MyoMechanix dataset, skeletal and sEMG signals are cropped using the identical keyframe indices: each skeleton sample is formatted into a 103-frame sequence, while the sEMG data retains its native temporal resolution. For all mul- timodal baselines, we adopt a late-fusion strategy, simi- lar to our CUBIST model (Sec. 3.5), ensuring fairness. Each modality is processed by its respective independent branch (e.g., I3D for RGB, SkateFormer for skeleton), and the extracted representations are integrated strictly at the feature level prior to the final prediction head. All experiments were conducted on multiple NVIDIA RTX 4090 GPUs, totaling approximately 4000 GPU-hours. Further details provided in the Appendix. All the codes will be publicly released. 5.1.3 Performance Metrics To maintain consistency with prior research in the AQA field, we utilize the Spearman Rank Correlation (ρ) and the Relative L2-Distance (R− l2) as the primary evaluation metrics. Theρmetric, originally introduced by Pirsiavash et al. [9], is the most widely adopted performance indicator within the field of action quality assessment. This metric effectively quantifies the strength and direction of the monotonic relationship between the ground-truth and predicted rankings; however, it does not capture the absolute differences between their respective scores. The formula for ρ is defined as follows: ρ = 1− 6 P n i=1 (R i − c R i ) 2 n(n 2 − 1) (11) whereR i and b R i represent the ground truth and pre- dicted rankings of thei-th sample, respectively, andn denotes the total number of samples. The value ofρ ranges within [−1,1], with values approaching 1 indicat- ing superior predictive performance. To address this limitation, Yu et al. [51] proposed the Relative L2-Distance (R− l2), a metric based on the L2 norm that, unlikeρ, emphasizes the numerical discrepancy between the ground-truth and predicted scores. Furthermore, normalizing the score ranges across different categories of actions,R− l2 facilitates cross- category training in a manner that the traditional L2 distance cannot. The formulaR−l2 is defined as follows: R− ℓ 2 = 1 n n X i=1 ( |s i − bs i | s max − s min ) 2 (12) wheres i andbs i represent the ground truth and predicted scores of thei-th sample, respectively. The variables s max ands min represent the maximum and minimum scores within the respective category of action, andn represents the total number of samples. The value of R− l2 ranges within [0,1], with values approaching 0 indicating superior predictive performance. 5.1.4 Results All the results are summarized in Table. 5. Performance on individual actions across various splits are reported in Table. 6, Table. 7, Table. 8, and Table. 9. Qualitative comparative examples are provided in Table. 10. 25 Table 5 Average performance of various AQA models across various protocols.R − ℓ 2 (×100). Prepended ’B’ indicates best model in a particular modality. Prepended S and M indicate single and multiple views, respectively. All the multimodal and multiview models have been developed by us. Notice the improvement as modalities and views are incorporated. CUBIST paradigm achieves the new SOTA. ProtocolModelρ↑ R−ℓ 2 ↓ Vanilla I3D-MLP [48]0.48373.6586 CoRe [51]0.68562.2941 TPT [2]0.58323.1595 HGCN [4]0.71442.6387 MCoRe [40]0.69362.7761 T2CR [41]0.46023.8081 DAE [50]0.74312.4180 Ours CUBIST0.7829 2.1542 MS-G3D [83]0.55123.4048 CTR-GCN [84]0.58733.1140 ST-GCN++ [85]0.60722.8277 SkateFormer [86]0.65082.7112 EMG0.20464.9312 BSV+BPose0.79541.4689 BSV+EMG0.76072.3050 BPose+EMG0.66922.4187 BSV+BPOSE+EMG0.82971.1403 Ours CUBIST (SV+Pose)0.82961.0528 Ours CUBIST (SV+Pose+EMG) 0.8455 1.0198 BMV0.78611.8593 BMV+BPose0.84980.9915 BMV+EMG0.80161.1588 BMV+BPose+EMG0.85640.8865 Ours CUBIST (MV+Pose+EMG) 0.8836 0.7418 Cross Subject BSV0.34844.1568 BPose0.34174.1753 EMG0.13757.7435 BSV+BPose+EMG0.38534.0956 BMV+BPose+EMG0.4323 3.9717 Cross View I3D-MLP [48]0.31364.4885 CoRe [51]0.39814.0688 TPT [2]0.35384.1956 HGCN [4]0.43723.9519 MCoRe [40]0.41273.9628 T2CR [41]0.29844.3975 DAE [50]0.4663 3.7596 Mix View I3D-MLP [48]0.34614.1619 CoRe [51]0.44093.8987 TPT [2]0.40894.0452 HGCN [4]0.48763.6371 MCoRe [40]0.45933.8137 T2CR [41]0.31584.2100 DAE [50]0.5111 3.5171 Performance of Uni-Modal Models.1We con- duct unimodal comparison experiments on the Vanilla split. We first train and evaluate 7 baseline models from the vision branch alongside the CUBIST model under a unified experimental configuration. The results in Table. 5 show clear hierarchical differences in per- formance among vision baseline models with diverse topological architectures.2The performance of foun- dational models such as I3D-MLP remains relatively limited. Conversely, temporal modeling or contrastive learning networks explicitly tailored for the AQA task, such as CoRe and TPT, achieve significant performance gains. Further analysis demonstrates that DAE, which relies on fine-grained feature representations, attains highly competitive results (ρ= 0.7431). Notably, our proposed CUBIST model outperforms all baselines – in- cluding highly competitive ones – achieving new SOTA performance (ρ= 0.7829,R− ℓ 2 = 2.1542), thereby validating the advantages of our paradigm and the FKG structure.3For the pose modality, we evaluated mul- tiple classical architectures based on skeletal keypoints. Experimental results indicate that models employing advanced graph convolutional networks (GCNs) and Transformer architectures demonstrate superior repre- sentation capabilities. Specifically, SkateFormer achieved a performance ofρ= 0.6508, significantly outperforming MS-G3D (0.5512), CTR-GCN (0.5873), and ST-GCN++ (0.6072), highlighting its advantage in capturing pure kinematic topological features of the human body.4 Furthermore, we conducted an independent evaluation to characterize the performance of the sEMG modality (ρ= 0.2046). As expected, the standalone performance of the sEMG modality appears relatively limited from a purely numerical perspective. Nevertheless, the valuable physiological insights it offers extend beyond visual or kinematic modalities. Specifically, sEMG provides direct and objective information about the activation patterns of targeted muscle groups, including co-contraction, fa- tigue signatures, force imbalances, and compensatory muscle recruitment. Beyond its numerical contribution to overall prediction accuracy, this physiological infor- mation deserves deeper consideration for the comple- mentary perspective it provides within a multimodal understanding framework. Effectiveness of Multi-Modal Fusion. 1 Based on the quantitative findings from the unimodal experiments, we selected DAE [50] as the visual backbone for our subsequent multimodal baseline models (BSV/BMV), and SkateFormer [86] as the skeletal backbone (BPose). We then systematically ablate the multimodal config- urations by progressively combining different modali- ties and evaluating their performance. In parallel, we 26 Table 6 Performance of AQA models on each action in the Vanilla split (ρ↑). Panel A: A01–A10 Modality ModelA01A02A03A04A05A06A07A08A09A10 Vision I3D-MLP0.43820.36860.46120.38930.54750.49280.50380.50040.48880.4532 CoRe0.66620.56150.66600.62130.71060.73970.73920.76690.64800.6498 TPT0.51360.74670.51700.57170.55050.70800.54060.48510.59540.5363 HGCN0.64300.67660.76240.66800.62160.67430.73730.74360.69510.6994 MCoRe0.70350.63150.66470.44310.77890.48040.57500.97050.81330.6891 T2CR0.32550.40580.54740.38320.54930.52170.37480.55990.41800.3778 DAE0.58700.70960.72640.68320.68280.74070.78180.78020.74430.7696 CUBIST0.81120.78840.78200.78750.79140.72880.83750.89130.73720.7720 Pose MS-G3D0.33170.64220.47450.46050.54640.46890.59520.54420.53840.4906 CTR-GCN0.37030.49130.56560.46550.65500.55690.50800.61110.66400.5312 ST-GCN++0.53550.53090.58300.44930.47000.57740.55790.62720.53310.3796 SkateFormer0.65300.69130.61260.54460.69480.49150.46500.68230.67540.5453 EMG DAE0.13540.14260.15990.14860.31200.18650.27130.05980.35220.2538 Single View BSV+BPOSE0.64090.77870.79480.73400.74970.78260.82750.83020.79130.8201 BSV+EMG0.60270.72660.74620.69870.70520.75970.80130.80020.76310.7868 BPOSE+EMG0.64050.69890.66190.58310.72780.54860.50730.70790.73040.5786 BSV+BPOSE+EMG0.67740.81320.83040.76820.78550.82030.86420.86520.82810.8566 CUBIST (SV+P)0.79600.79650.84840.79180.82360.87790.85720.85680.86640.7781 CUBIST (SV+P+E)0.81500.87800.79700.83150.89830.86000.91270.77030.91980.7924 Multi View BMV0.64470.75810.77810.68750.75450.80360.83170.81110.79540.7946 BMV+BPOSE0.71540.83020.84180.76280.83270.86210.88720.87540.86800.8590 BMV+EMG0.66160.77320.79520.70190.77150.81860.84760.82550.80940.8110 BMV+BPOSE+EMG0.72960.83750.84720.77370.83800.86880.89510.88150.87210.8653 CUBIST (MV+P+E)0.86780.87570.84600.87920.93210.90000.90870.81440.97630.8349 Panel B: A11–A20 Modality ModelA11A12A13A14A15A16A17A18A19A20 Vision I3D-MLP0.42170.43350.42070.49830.59510.50470.52250.52820.50900.5966 CoRe0.67500.63170.43200.69890.79190.69630.70070.76580.77950.7706 TPT0.51630.65520.53470.49180.65860.52310.60160.79070.57630.5509 HGCN0.63910.69380.63270.75470.83870.75690.71260.76720.75920.8110 MCoRe0.59390.47250.70200.76560.96310.53070.86790.79990.60970.8173 T2CR0.34080.48020.27320.49810.68940.54020.53940.40730.50050.4707 DAE0.73360.68810.72100.74660.81810.73070.82330.81450.74340.8368 CUBIST0.80600.74150.74700.73710.71920.82330.71990.86490.79670.7753 Pose MS-G3D0.47320.66950.57660.46200.70120.40050.67300.60970.64390.7212 CTR-GCN0.49380.64600.66060.50350.74780.58640.66960.69130.63760.6910 ST-GCN++0.59590.68030.63390.72070.90370.51730.63340.68990.72650.7983 SkateFormer0.57250.60600.64540.66150.85050.64740.71910.77170.68440.8010 EMG DAE0.12710.26410.23080.13860.37530.16300.18770.15920.10540.3188 Single View BSV+BPOSE0.78850.74360.78380.79670.87210.79220.85970.86130.77510.8851 BSV+EMG0.75020.70890.74490.76020.83450.74570.83780.83330.75490.8524 BPOSE+EMG0.61370.61550.67040.69360.88250.65140.64610.73280.69320.7997 BSV+BPOSE+EMG0.82310.77030.81520.83320.90480.82360.89360.89030.81150.9195 CUBIST (SV+P)0.84250.86060.76890.85430.83080.76810.85680.84970.84690.8211 CUBIST (SV+P+E)0.85380.88460.77450.82970.84200.78880.86120.89490.85130.8531 Multi View BMV0.74300.73770.78150.78100.85910.77950.85250.85470.79920.8748 BMV+BPOSE0.81020.79100.84120.83620.92290.84820.90910.91120.85010.9408 BMV+EMG0.75590.75170.79800.79690.87420.79390.86980.87030.81560.8907 BMV+BPOSE+EMG0.81810.79520.84550.84430.92940.85530.91450.91460.85690.9460 CUBIST (MV+P+E)0.91110.91830.80830.86810.89280.84240.90460.94830.85880.8838 27 Table 7 Performance of AQA models on each action in the Cross-Subject splits (ρ↑). Panel A: A01–A10 ModelA01A02A03A04A05A06A07A08A09A10 Single-View0.05890.35580.09830.30180.46180.16860.34360.66520.42010.6200 SkateFormer0.24080.13250.32720.52970.58230.42600.39460.13780.13670.2333 EMG0.16720.05300.08630.00650.25970.15460.06380.25480.14080.1573 BSV+BPOSE+EMG0.10050.36540.21410.33270.58140.37910.11930.51700.31950.5290 BMV+BPOSE+EMG0.20380.38660.26600.54990.69630.34660.24880.51410.27260.4432 Panel B: A11–A20 ModelA11A12A13A14A15A16A17A18A19A20 Single-View0.21530.23030.26590.11050.65090.11430.51910.76600.14640.4551 SkateFormer0.13820.18700.17670.16680.59280.40200.64470.57080.11510.6995 EMG0.12060.00870.09800.06380.34840.18950.10780.17940.11950.1710 BSV+BPOSE+EMG0.18020.21520.24920.24250.62210.50230.69670.63000.18160.7284 BMV+BPOSE+EMG0.125400.18870.35120.31240.57690.59710.72160.72070.23730.7589 Table 8 Performance of AQA models on each action in the Cross-View splits (ρ↑). Panel A: A01–A10 ModelA01A02A03A04A05A06A07A08A09A10 I3D-MLP0.32640.55100.37990.26940.43290.28060.05780.50870.33410.1246 CoRe0.38030.31520.35300.30710.38850.41100.49080.48740.42020.3733 TPT0.36720.45360.31010.28690.37430.43920.28480.37180.28850.3244 HGCN0.39530.71660.45000.40060.35130.38930.51690.55200.32500.4726 MCoRe0.38650.45770.44530.38740.44860.42750.33390.45850.42670.3307 T2CR0.23080.41120.24170.34960.17730.17270.12290.56180.32700.3148 DAE0.57750.53680.18810.34630.67170.42680.44970.55420.56900.4500 Panel B: A11–A20 ModelA11A12A13A14A15A16A17A18A19A20 I3D-MLP0.26880.45790.25800.21310.63980.36910.37200.05450.17950.1944 CoRe0.38340.33810.12730.36740.48370.41380.36340.51010.54600.5012 TPT0.39740.32070.34150.38810.30180.35900.33840.49030.30130.3364 HGCN0.41240.40510.18010.54470.78010.44960.45570.40390.14610.3973 MCoRe0.33460.34990.33630.53960.44260.41190.44860.40710.48150.3989 T2CR0.24690.17240.23980.13840.12290.37390.65580.25080.40740.4497 DAE0.32230.46350.19520.31310.71740.52400.57420.56390.56740.3153 Table 9 Performance of AQA models on each action in the Mix-View splits (ρ↑). Panel A: A01–A10 ModelA01A02A03A04A05A06A07A08A09A10 I3D-MLP0.30060.23100.32360.25170.40990.35520.36620.36280.35120.3156 CoRe0.42150.31680.42130.37660.46590.49500.49450.52220.40330.4051 TPT0.33930.57240.34270.39740.37620.53370.36630.31080.42110.3620 HGCN0.41620.44980.53560.44120.39480.44750.51050.51680.46830.4726 MCoRe0.46920.39720.43040.20880.54460.24610.34070.73620.57900.4548 T2CR0.18110.26140.40300.23880.40490.37730.23040.41550.27360.2334 DAE0.35500.47760.49440.45120.45080.50870.54980.54820.51230.5376 Panel B: A11–A20 ModelA11A12A13A14A15A16A17A18A19A20 I3D-MLP0.28410.29590.28310.36070.45750.36710.38490.39060.37140.4590 CoRe0.43030.38700.18730.45420.54720.45160.45600.52110.53480.5259 TPT0.34200.48090.36040.31750.48430.34880.42730.61640.40200.3766 HGCN0.41230.46700.40590.52790.61190.53010.48580.54040.53240.5842 MCoRe0.35960.23820.46770.53130.72880.29640.63360.56560.37540.5830 T2CR0.19640.33580.12880.35370.54500.39580.39500.26290.35610.3263 DAE0.50160.45610.48900.51460.58610.49870.59130.58250.51140.6048 28 Table 10 Qualitative comparison of outputs from CUBIST and SOTA AQA Model. Compared to DAE, which outputs only a single holistic score, CUBIST not only achieves higher prediction accuracy but also introduces interpretable mechanisms for error detection and compositional scoring, thereby significantly enhancing the transparency of the assessment. Furthermore, by leveraging the FKG, CUBIST provides actionable, expert-level feedback corresponding to specific errors. ComponentModel Decline Barbell Bench PressSquat Exercise True Score 70.2372.39 Predicted Score DAE 61.2478.41 CUBIST 63.3668.66 Interpretable Error Detection DAE ✗ CUBIST ✓ Decline Bench Press E T 0 3 6 E T 0 0 2 E T 0 0 4 E T 0 3 7 E T 0 0 6 E T 0 0 7 E T 0 0 9 E T 0 1 7 E T 0 1 8 E T 0 3 8 E T 0 3 9 E T 0 1 9 E T 0 1 4 E T 0 4 0 E T 0 4 1 E T 0 4 2 E T 0 4 3 E T 0 4 4 Preparation PhaseConcentric PhaseEccentric Phase E T 0 0 8 CUBIST Precise Skill Error Detection ✓ Squat ActionVideo E T 0 6 1 E T 0 0 2 E T 0 0 3 E T 0 0 4 E T 0 0 6 E T 0 0 7 E T 0 1 0 E T 0 1 7 E T 0 0 2 E T 0 0 5 E T 0 5 6 E T 0 5 5 E T 0 3 2 E T 0 3 9 E T 0 1 4 E T 0 0 2 E T 0 5 2 E T 0 3 2 Preparation PhaseConcentric PhaseEccentric Phase CUBIST Precise Skill Error Detection Compositional Scoring DAE ✗ CUBIST [1-(8+8+...+9+9)/131]×100=63.36[1-(8+5+...+9+10)/134]×100=68.66 Feedback DAE ✗ CUBIST - Actively depress the shoulders. Imagine pulling the scapulae down toward the hips, tightening them back and downward to prevent the shoulders from elevating toward the ears. - Maintain an angle of about 45 ◦ -60 ◦ between the upper arms and torso, forming an arrow shape. - Lower slowly until the target muscles are adequately stretched. Reduce the load if necessary. - Adjust the barbell to an appropriate finishing position. - Actively retract and depress the scapulae against the bench, maintaining a chest up, shoulders down posture. As you press, visualize compressing the target muscles. - Imagine bending the barbell, maintaining symmetrical force output in both hands. - Ensure the knees always track over the toes, avoiding inward collapse. - Slightly turn the toes outward and distribute your weight evenly across the entire foot. Reduce the load if necessary. - Sit the hips back while the knees move forward slightly to maintain balance but avoid excessive forward knee travel. Keep the heels firmly on the ground, spine straight, and the bar’s center of gravity between the forefoot and heel. - Engage the abdominal muscles and slightly contract the glutes, maintaining the lumbar spine in a neutral position. - Rise by extending the hips and knees together, lifting both the hip and shoulder joints simultaneously while keeping the back stable. conduct the same multimodal evaluations using our CUBIST model. 2 Under single-view (SV) conditions, purely visual unimodal networks are inherently lim- ited by the absence of explicit depth information and their susceptibility to self-occlusion. When skeletal in- formation (BSV+BPOSE, 7.04% gain) or EMG signals (BSV+EMG, 2.37% gain) are incorporated into BSV (DAE), the model achieves clear performance improve- ments. Furthermore, tri-modal fusion (BSV + BPOSE + EMG) raises performance toρ= 0.8297, correspond- ing to an 11.65% improvement over the existing uni- modal SOTA, DAE. More importantly, CUBIST demon- strates strong multimodal modeling capability, achieving ρ= 0.8296 with only bimodal inputs. After integrat- ing all three modalities in CUBIST(SV+POSE+EMG), performance further improves toρ= 0.8455, represent- ing a 7.99% gain over unimodal CUBIST, whileR− ℓ 2 decreases to 1.0198.3To alleviate the spatial ambi- guity inherent in single-view inputs, we further extend the evaluation to multi-view (BMV) settings. The multi- view input itself (ρ= 0.7861) yields a 5.7% improvement over the single-view SOTA. The benefits of multimodal fusion remain evident: relative to DAE, incorporating skeletal information (BMV+BPOSE) and all modali- ties (BMV+BPOSE+EMG) improvesρby 14.36% and 15.25%, respectively. Ultimately, CUBIST(MV + POSE + EMG) achieves the best overall performance, with ρ= 0.8836 andR − ℓ 2 = 0.7418. Compared with unimodal CUBIST, this corresponds to a 12.86% im- provement inρ; compared with DAE, the improvement reaches 18.91%. These progressive ablation results vali- date the SOTA performance of CUBIST and highlight the importance of complementary information sources for AQA. Accurate action quality assessment benefits not only from external kinematic observations, such as multi-view appearance and skeletal structure, but also from internal physiological cues provided by sEMG. By complementing visual representations with muscle- activation information, sEMG helps the model capture aspects of action execution that are not directly ob- servable from appearance alone. 4 In summary, the multimodal ablation experiments show that progres- 29 sively introducing complementary modalities leads to consistent performance gains on the AQA task. Across the evaluated settings, skeleton keypoints (POSE) and multiple views (MV) contribute notable improvements, withρincreases of 7.76% and 6.32%, respectively, while sEMG provides a stable additional gain of 2.27%. Al- though the absolute improvement from EMG is smaller, its consistent contribution supports its complementary role in multimodal AQA. Overall, these results demon- strate a clear cross-modal synergy between external kinematic representations and internal physiological sig- nals. Performance on Generalization-oriented Splits. 1To thoroughly investigate the model’s generalization capability under data distribution shifts, we conduct benchmark testing across three generalization-oriented data-splitting protocols: cross-subject, cross-view, and mix-view.2As shown in Table. 5, all single-modality baseline models exhibit substantial performance degra- dation when encountering unseen subjects or novel camera views. For example, in the highly challenging cross-subject split, where subject-specific differences in movement patterns become especially pronounced, theρvalue of the single-view visual model drops to 0.3484, while SkateFormer similarly declines to 0.3417. These results highlight the limitations of unimodal mod- els in generalizing across subject-dependent appear- ance and motion variations.3In contrast, the mul- timodal fusion paradigm provides improved robustness under such distribution shifts. By incorporating struc- tural pose cues (POSE) and appearance-independent physiological signals (EMG), the multimodal configu- ration helps alleviate the sharp decline in performance. Notably, in the cross-subject setting, the combined BMV+BPOSE+EMG model raises the correlation met- ric to 0.4323, significantly outperforming each single- modality baseline. Furthermore, comparing the quanti- tative results of cross-view and mix-view reveals better overall performance under the mix-view setting. This improvement is likely attributable to the availability of video streams from multiple views during mix-view training, which enables the network to learn more ro- bust view-invariant representations. In contrast, the cross-view protocol requires inference on entirely unseen viewpoints, making it a substantially more demanding generalization setting. Summary of the Results. Our research presents the following four core findings: – CUBIST achieves the best performance: In pure vi- sion unimodal evaluations, CUBIST demonstrates strong spatiotemporal modeling capability and achie- ves new SOTA performance. Under multimodal fu- sion settings, it further exhibits effective cross-modal feature integration and again achieves new SOTA performance, highlighting the effectiveness of My- oMechanix’s FKG design, multimodal information, and the CUBIST paradigm. –Multimodal data delivers clear performance gains: Experimental results demonstrate that the transi- tion from single-view to multi-view inputs, together with the incorporation of complementary modalities spanning external kinematics (vision and skeleton) and internal physiology (surface electromyography), yields substantial synergistic benefits. These findings highlight the value of multidimensional information for accurate AQA. –Limited generalization under challenging splits em- phasizes the dataset difficulty: When confronted with substantial data distribution shifts, such as those across subjects and viewpoints, advanced models exhibit noticeable performance degradation. This finding underscores the challenge of disentangling action quality from subject- and view-specific varia- tions, while also motivating future research toward more robust real-world AQA systems. –Considerable room for improvement remains across all splits: The results indicate substantial opportuni- ties for future research in model design, multimodal fusion, generalization, interpretability, and related directions. 5.2 Video Question Answering In this experiment, we use MyoMechanix-VideoQA data- set (Sec. 3.5.2) to assess whether domain-specific fine- tuning can equip current vision-language models (VLMs) and multimodal large language models (MLLMs) with query-conditioned, fine-grained biomechanical action understanding in fitness scenarios. 5.2.1 Baseline Methods We select representative state-of-the-art open-source VLMs as baseline models, including MiniCPM-O-2.6- 8B [87], InternVL3-2B [88], and Qwen2.5-VL-3B [89]. To comprehensively and objectively evaluate model per- formance, we consider three distinct testing configura- tions: 1) Off-the-shelf mode: directly using the origi- nal pre-trained weights for inference; 2) Rules-based prompting mode: incorporating standardized, struc- tured action execution criteria as contextual guidance during evaluation; and 3) Supervised fine-tuning mode: applying LoRA-based fine-tuning to adapt the pre-trained models to our dataset and task. 30 Table 11 Performance of VLMs on MyoMechanix-VideoQA dataset. R− ℓ 2 (×100) Metrics PretrainedPromptSFT InternVLMiniCPMQwenInternVLMiniCPMQwenQwen BLEU ↑0.00030.01540.03880.00070.03120.04150.1970 ROUGE-L ↑0.16890.29020.20120.20020.31020.21280.4010 METEOR ↑0.14050.30240.30490.17360.33450.32060.4688 CHRF++ ↑8.722820.418036.134017.307828.486539.530852.0005 BERTScore F1 ↑0.10170.25350.16640.15500.27430.16660.4390 R− ℓ 2 ↓3.990010.55393.11833.27578.64483.09423.0592 Table 12 Qualitative comparison of models in natural language feedback and error analysis generation. DimensionQwen (SFT)Qwen (Prompt)MiniCPM (Prompt)InternVL (Prompt) Coarse Action Recognition✓ Fine-grained AnalysisBetterGoodLimitedLimited Error Diagnosis5 faults flagged2 faultsNoneNone Key-point Coverage (5 total)5433 WordingConcise, technicalVerbose, exhaustiveModerateDirect, brief Table 13 Qualitative comparison of models in score prediction. Sample IDGT ScoreQwen (SFT)Qwen (Prompt)MiniCPM (Prompt)InternVL (Prompt) P08-A09-09100.0065659585 N04-A15-0122.1665659575 A05-A20-0468.1555657575 5.2.2 Implementation Details We performed supervised fine-tuning of Qwen2.5-VL using the LLaMA-Factory framework [90]. Training was conducted in bfloat16 precision with Low-Rank Adap- tation (LoRA), using a rank of 8, an alpha of 16, and a dropout rate of 0.05. We optimized the model with AdamW, using an initial learning rate of 5×10 −5 , co- sine learning rate decay, and gradient clipping with a maximum norm of 1.0. The per-device batch size was set to 2, with gradient accumulation over 4 steps. Training took ∼12 hours on a single NVIDIA H20 GPU. 5.2.3 Performance Metrics To rigorously evaluate the quality of the open-ended an- swers against the ground-truth annotations provided by experts, we employ a multi-granular evaluation frame- work. This framework spans from surface-level lexical fidelity (BLEU [91], ROUGE-L [92]) and morphologi- cal robustness (METEOR [93], CHRF++ [94]) to deep semantic consistency (BERTScore-F1 [95]). Specifically, BLEU and ROUGE-L assess lexical agreement between generated responses and reference answers, including the accurate use of task-relevant anatomical terms such as body-part mentions via the precision and recall of n-grams. METEOR further ac- counts for stemming and synonym matching, while CHRF++ captures character- and word-level n-gram similarity, making it less sensitive to minor surface-form variations. To complement these overlap-based measures, BERTScore-F1 leverages pre-trained contextual embed- dings to estimate semantic similarity between generated and reference answers. Collectively, these metrics pro- vide a more comprehensive assessment of both phrasing fidelity and meaning preservation. For all metrics, higher values indicate better performance. 5.2.4 Results Overall.1As shown in Table. 11, the three testing strategies and three evaluated models exhibit a clear hierarchy in performance. Pre-trained models evaluated in a zero-shot setting provide a basic baseline; however, they generally struggle to produce the structured diag- nostic responses required by our MyoMechanix-VideoQA benchmark. Importantly, this suggests that our dataset requires perception and reasoning not covered under existing datasets used for training foundational models. Thus, incorporating our dataset would help develop more powerful and holistic foundational models. Incorporat- ing domain-specific scoring rubrics through prompting yields moderate yet consistent improvements across most metrics, suggesting that inference-time guidance helps 31 align model outputs with task expectations. This also suggests that our dataset supports novel tasks, not yet covered by existing foundational datasets. Supervised fine-tuning (SFT) further enhances task-specific gener- ation quality. Compared with the prompt-only setting, SFT on the Qwen model produces substantial gains in text generation metrics, including BLEU, which in- creases from 0.0415 to 0.1970, and CHRF++, which rises from 39.5308 to 52.0005. Despite these improve- ments in language generation, theR−ℓ 2 error decreases by less than 2% across adaptation stages. This suggests that while weight adaptation is more effective than ex- plicit textual prompting for internalizing task-specific response patterns, its impact on score prediction accu- racy remains limited.2Beyond the trends across adap- tation strategies, the comparison of individual models highlights the importance of video-oriented pretraining for fine-grained biomechanical action understanding en- abled by our MyoMechanix-VideoQA dataset. Across all baseline configurations, Qwen2.5-VL-3B consistently achieves the strongest overall performance. Despite its relatively small parameter count, Qwen appears to bene- fit from extensive multi-source video pretraining, which may provide stronger motion priors and contribute to its balanced performance across both lexical and semantic metrics. In contrast, MiniCPM-O-2.6-8B benefits from a stronger language backbone, enabling competitive per- formance on semantic-similarity metrics after prompting. However, its weaker results on stricter lexical match- ing metrics suggest that limitations in video-specific temporal understanding may reduce its fine-grained di- agnostic precision. InternVL3-2B, which is designed for more lightweight deployment, shows lower overall per- formance, potentially reflecting combined constraints in visual perception and language modeling capacity. For score prediction, as measured byR− ℓ 2 , Qwen consis- tently achieves the lowest error, MiniCPM exhibits the highest error, and InternVL falls in between. Qualitative Text Comparison. To provide a detailed view of VLMs’ performance, we select a representative example (P07-A17-01) from the test set to showcase text generation quality (see Table. 12).1For sample P07-A17-01, all evaluated VLMs correctly categorize the action as a barbell bent-over row. This indicates that coarse-grained action recognition is already robustly han- dled by current models. However, when the task shifts to the distinctive requirements of MyoMechanix-VideoQA – joint-level motion diagnosis, temporal reasoning, and ex- ecution scoring – the capabilities of these models sharply diverge. The SFT Qwen delivers the most diagnosti- cally valuable narrative. It precisely pinpoints multiple biomechanical faults, such as spinal curvature, shoulder elevation, and improper eccentric tempo. Furthermore, it expresses these observations with succinct, discipline- specific diction, successfully blending technical precision with error-focused depth. 2 In contrast, the same model under a prompt-only configuration achieves alignment with the reference checklist but adopts a largely affirma- tive stance devoid of critical feedback. This exhaustive yet uncurated recitation inflates the textual volume without yielding actionable coaching insights. Similarly, the MiniCPM and InternVL baselines reproduce core setup cues but verge on generating generic rehearsal manuals. They confine themselves to superficial form validation and entirely omit the identification of faults, which limits their utility for quantitative evaluation. 3 Ultimately, this contrast underscores the core challenge and novelty of the MyoMechanix-VideoQA benchmark. While current VLMs have largely mastered the classifica- tion of action types, the modeling of kinematic nuances and the delivery of fine-grained quality assessment re- main significant open challenges. Stagnant Score Prediction. During the experiments (see Table. 11 and Table. 13), we observed that although text generation quality improved markedly, the accuracy of scoring did not exhibit a corresponding improvement, which lagged behind specialized AQA models. We be- lieve this can be primarily attributed to the inherent limitations of large pretrained language models in han- dling numerical data. During pretraining, these models rely on autoregressive prediction over discrete tokens – splitting numbers into subword units and treating them as standard vocabulary. Consequently, they lack dis- tance constraints on the continuous numerical axis and do not employ specialized regression losses to reinforce gradient signals for magnitude differences. Furthermore, the training objective of these models is to maximize likelihood rather than to achieve precise counting or numerical regression, with the optimization process fa- voring overall syntactic and semantic coherence. As a result, these models struggle to represent fine-grained visual details or distinctions in scores, yielding approxi- mate rather than accurately calibrated scoring outputs. To improve on numerical task, we suggest future work could begin with a unified modeling framework that integrates discrete generation and continuous regres- sion. By leveraging numerically aware representation learning and multi-scale supervisory signals, numerical values would be mapped to continuous vectors endowed with dimensional and ordinal constraints. An explicit regression head or score-distance regularization could then be employed to simultaneously minimize language reconstruction error and numerical deviation during autoregressive generation. At the same time, introduc- 32 ing cross-modal alignment through pose, temporal, or physical constraints would enforce a consistent metric structure among vision, language, and numeric modali- ties in the latent space, thereby enhancing the model’s sensitivity to fine-grained quantitative differences and its generalization ability. Summary. Our experimental results demonstrate that the MyoMechanix-VideoQA dataset presents a signif- icant challenge to SOTA VLMs. Although supervised fine-tuning substantially improves model performance and establishes a promising baseline for future research, the multimodal understanding capability remains far from fully realized. In the text generation task, existing VLMs still have approximately 50% room for perfor- mance improvement. In the action quality score pre- diction task, the performance of current VLMs still lags significantly behind that of domain-specific AQA algorithms by up to 80%. 5.3 Novel Technology—Video2EMG ResNet ViT LSTM SVR Weight Layer ReLUReLU Weight Layer Norm MHA Norm MLP L× Linear Projection Transformer 퐻 푡−1 퐻 푡 퐶 푡 퐶 푡−1 푋 푡−1 휎 tanh tanh 퐹 푡 퐼 푡 푂 푡 � 퐶 푡 Normal Sample Support Vector 푥 푦 푓 푥+휀 푓(푥) 푓 푥 −휀 sEMG PredictorVisionEncoder Fig. 12 Illustration of proposed baseline models’ ar- chitectures for the Video2EMG task. As introduced earlier, while surface electromyogra- phy (sEMG) signals provide indispensable, precise evi- dence of muscle functional states, their acquisition relies on costly and obtrusive wearable sensors. To mitigate this limitation, we envision a future where fine-grained, muscle-level feedback can be estimated directly from ordinary videos. To pave the way for this novel cost- effective technology, we implement and evaluate the proposed Video2EMG task. By leveraging the strictly synchronized multimodal recordings within the My- oMechanix dataset, this benchmark challenges models to predict continuous sEMG sequences solely from visual observations. In this section, we establish the founda- tional baselines for this novel cross-modal translation task, demonstrating the feasibility of inferring internal, latent physiological dynamics from observable kinematic cues. 5.3.1 Our Video2EMG Baseline Design Given the novelty of the Video2EMG task, established models are currently unavailable for direct usage or adoption. Thus, we design sound baseline models for the Video2EMG task. For this, we draw upon the core princi- ples of time-series waveform analysis and relevant studies such as emg2pose [75] to construct streamlined end-to- end spatiotemporal, visual regression models to predict EMG signals from videos (see Fig. 12). Our models in- tegrate visual encoders (ResNet [96] or ViT [97]) with an sEMG sequence predictor (LSTM [98] or SVR [99]). LSTM processes features sequentially, while for SVR, features are concatenated along the temporal direction (a parallel implementation). We have provided further modeling and implementation details in the following. 5.3.2 Video2EMG Modeling Details For LSTM-based models, the frame-level features ex- tracted by the visual encoders were processed by a single-layer LSTM configured with a hidden size of 256. Subsequently, the final hidden state was passed through fully connected layers for the final prediction. In the case of kernel-based SVR machines, features from 16 frames were concatenated into high-dimensional vectors and were regressed to the space of the EMG signals uti- lizing SVRs with RBF kernels. The models were trained and evaluated on 18 muscles according to the collected sEMG signals. The training process utilized the mean squared error loss function alongside the Adam opti- mizer, which was configured with an initial learning rate of 1×10 −3 . Furthermore, a batch size of 32 was employed. This process required approximately 6 hours on a single NVIDIA RTX4090 graphics processing unit. Beyond the vailla split, we also introduce two rigorous evaluation protocols: cross-view and cross-subject config- urations. In total, the experiments for the Video2EMG task consumed approximately 72 hours of computation. 5.3.3 Task Standardization and Performance Metrics Since the task is novel, we establish a standard practice for it and propose multiple suitable metrics to measure the performance on the Video2EMG task. In accor- dance with standard evaluation protocols in time-series waveform analysis, we adopt the Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Cross- Correlation Peak (CCP) as the primary evaluation met- rics for this task. The mean absolute error (MAE) and 33 Table 14 Performance of Video2EMG models. Model VanillaCross-ViewCross-Subject CCP ↑MAE ↓RMSE ↓CCP ↑MAE ↓RMSE ↓CCP ↑MAE ↓RMSE ↓ ResNet+LSTM0.37060.16550.21330.27230.25500.30540.28720.29340.3442 ResNet+SVR0.41740.14660.18780.27930.15080.20020.30620.17050.2170 ViT+LSTM0.40400.16480.20690.31040.20790.26190.30860.21480.2655 ViT+SVR0.43450.14650.18820.33110.14830.19850.33110.16310.2104 the root mean square error (RMSE), which quantify the average error and the variability of the error, respec- tively, are reported. Furthermore, the cross-correlation peak (CCP) is utilized to quantify the morphological similarity between the predicted and the target wave- forms. The CCP is formally defined as the maximum value of the normalized cross-correlation function: CCP = max m P n [y(n)− ̄y]·[ˆy(n−m)− ̄ ˆy] √ P n [y(n)− ̄y] 2 · P n [ˆy(n−m)− ̄ ˆy] 2 (13) wherey(n) andˆy(n) represent the discrete target and predicted waveforms, respectively, ̄yand ̄ ˆy denote their respective mean values; andmindicates the time lag. The value of the CCP is bounded between−1 and 1, with larger positive values indicating a stronger alignment of the respective waveforms. Specifically, a value of CCP approaching 1 demonstrates the successful capture of the essential morphology of the target sequence by the computational model, even in the presence of temporal shifts. 5.3.4 Results Action Video sEMG 010000200003000040000 −150 0 150 GroundTruth-L 010000200003000040000 −150 0 150 Predicted-L 010000200003000040000 −150 0 150 GroundTruth-R 010000200003000040000 −150 0 150 Predicted-R Left vastus lateralis Right vastus lateralis Fig. 13 Qualitative result of the Video2EMG task based on the ViT+SVR. Overall. Following 150 epochs under 3 evaluation pro- tocols, the baselines for the Video2EMG task achieve promising performance on the vanilla split. As demon- strated in Table. 14, the best configuration reaches an MAE of 0.1465, an RMSE of 0.1882, and a CCP of 0.4345. We have shown qualitative results of the Video2EMG task in Fig. 13. To provide a deeper under- standing of the model’s behavior, we present a detailed analysis of the network architectures and their robust- ness to distribution shifts below. Architectural Insights. Notably, the results in Ta- ble. 14 reveal only a marginal performance gap between the two visual backbones. This similarity is largely due to our specific experimental setup: by utilizing frozen, medium-scale models solely as feature extractors, both the ResNet-50 and ViT-S/16 architectures provide com- parably mature and stable high-level semantic represen- tations. Given identical training data and supervisory signals, the quality of the extracted features results in highly similar MAE outcomes. In contrast, the choice of the regression head exerts a much more pronounced effect on the final predictions. Under the exact same visual representations, the SVR-based temporal mod- eling consistently outperforms the LSTM, achieving a relative MAE reduction of approximately 10%. This disparity clearly indicates that visual feature extraction is not the primary limiting factor; rather, the complex cross-modal mapping from visual features to continuous physiological signals remains the major bottleneck in the current pipeline. Robustness Analysis. To comprehensively measure the morphological alignment of the predicted EMG wave- forms, we introduce the cross-correlation peak (CCP) as an intuitive evaluation metric. In practical EMG analysis, signals are typically nonstationary and noisy. Values above 0.7 generally indicate strong morphological similarity, 0.3 to 0.7 represent moderate similarity, and values below 0.3 suggest poor alignment or mismatched morphology. As reported in Table. 14, while the base- line methods achieve moderate waveform alignment on the vanilla split (CCP = 0.4345), their performance degrades noticeably under the disjoint cross-subject and cross-view protocols. This sharp decline reveals that the Video2EMG task on the MyoMechanix dataset is highly sensitive to individual physiological variations and viewpoint shifts. 34 Summary. These findings demonstrate that while esti- mating muscle-level feedback from videos is feasible, the task is highly challenging and far from being fully solved. The current baselines establish a promising foundation but leave considerable room (about 56%, because CCP has an upper bound of 1.0) for future research to cover. 6 Conclusion In this work, we introduced MyoMechanix, a biomechani- cally grounded multimodal ecosystem for skilled activity understanding and coaching, and established a com- positional paradigm for action quality assessment that moves beyond purely visual, monolithic action modeling. To support this direction, we presented MyoMechanix, a large-scale multimodal ecosystem for weight-loaded ac- tions that synchronizes multiview RGB video, 3D pose, sEMG, and complementary physiological signals under strong internal–external alignment. Together with ex- pert annotations, MyoMechanix provides a rich founda- tion for studying fine-grained, physically meaningful ac- tion understanding. Building on this sensing framework, we developed the Fitness Knowledge Graph (FKG) to encode structured relationships among actions, phases, execution steps, errors, and corrective feedback. This enables a more interpretable and semantically grounded formulation of action quality assessment. We further pro- posed CUBIST, a compositional reasoning framework that decomposes actions into structured units, analyzes their quality, and recomposes them for holistic scoring, error attribution, and feedback generation. Across the newly established MyoMechanix-AQA, MyoMechanix- VideoQA, and MyoMechanix-Video2EMG benchmarks, extensive experiments demonstrate that multimodal sensing, structured representations, and compositional reasoning jointly improve performance, interpretabil- ity, and diagnostic capability, with our CUBIST model achieving state-of-the-art results. Our findings highlight three broader implications. First, effective action qual- ity assessment benefits substantially from integrating latent biomechanical signals with visible motion cues. Second, language-grounded reasoning over structured ac- tion knowledge can enhance fine-grained understanding in vision–language models. Third, the promising results on Video2EMG suggest a path toward estimating inter- nal muscle dynamics from accessible visual observations, potentially reducing dependence on specialized sens- ing hardware. Despite these advances, MyoMechanix remains a challenging benchmark, and several direc- tions merit further exploration, including more robust cross-subject generalization, richer causal modeling of biomechanics, broader action coverage, and stronger in- tegration with embodied and interactive systems. Over- all, this work establishes a new foundation for action understanding that is multimodal, biomechanically in- formed, and compositionally interpretable, opening op- portunities for next-generation Physical AI in fitness, rehabilitation, healthcare, and representation learning. Acknowledgement This work is supported by the Suzhou Basic Scientific Research Project under the Grant SSD2024013, the Doctoral Student Program of the Young S&T Talents Cultivation Project, CAST. Data Availability The full dataset and codebase are released under the C BY-NC-SA 4.0 license and are publicly available at the Github. References 1.Paritosh Parmar and Brendan Tran Morris. Learning to score olympic events. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 20–28, 2017. 2. Yang Bai, Desen Zhou, Songyang Zhang, Jian Wang, Errui Ding, Yu Guan, Yang Long, and Jingdong Wang. Action quality assessment with temporal parsing transformer. In European Conference on Computer Vision, pages 422–438. Springer, 2022. 3. Sijie Yan, Yuanjun Xiong, and Dahua Lin. Spatial tempo- ral graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018. 4.Kanglei Zhou, Yue Ma, Hubert PH Shum, and Xiao- hui Liang. Hierarchical graph convolutional networks for action quality assessment. IEEE Transactions on Circuits and Systems for Video Technology, 33(12):7749– 7763, 2023. 5.Hao Yin, Paritosh Parmar, Daoliang Xu, Yang Zhang, Tianyou Zheng, and Weiwei Fu. A decade of action qual- ity assessment: Largest systematic survey of trends, chal- lenges, and future directions. International Journal of Computer Vision, 134(2):73, 2026. 6.Veysel ALCAN and Murat Z ̇ INNURO ̆ GLU. Current de- velopments in surface electromyography. Turkish Journal of Medical Sciences, 53(5):1019–1031, 2023. 7. FASTMOVE. Fastmove wireless emg, 2024. 8. FASTMOVE. Fastmove 3d motion for realtime, 2024. 9.Hamed Pirsiavash, Carl Vondrick, and Antonio Torralba. Assessing the quality of actions. In European Conference on Computer Vision, pages 556–571. Springer, 2014. 10. Paritosh Parmar and Brendan Morris. Action quality assessment across multiple actions. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1468–1476. IEEE, 2019. 11. Paritosh Parmar and Brendan Tran Morris. What and how well you performed? a multitask learning approach to ac- tion quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 304–313, 2019. 12.Ryoji Ogata, Edgar Simo-Serra, Satoshi Iizuka, and Hi- roshi Ishikawa. Temporal distance matrices for squat classification. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition Workshops, pages 0–0, 2019. 35 13. Chengming Xu, Yanwei Fu, Bing Zhang, Zitian Chen, Yu- Gang Jiang, and Xiangyang Xue. Learning to score figure skating sport videos. IEEE Transactions on Circuits and Systems for Video Technology, 30(12):4578–4590, 2019. 14. Jibin Gao, Wei-Shi Zheng, Jia-Hui Pan, Chengying Gao, Yaowei Wang, Wei Zeng, and Jianhuang Lai. An asym- metric modeling for action assessment. In European Con- ference on Computer Vision, pages 222–238. Springer, 2020. 15.Ling-An Zeng, Fa-Ting Hong, Wei-Shi Zheng, Qi-Zhi Yu, Wei Zeng, Yao-Wei Wang, and Jian-Huang Lai. Hybrid dynamic-static context-aware attention network for action assessment in long videos. In Proceedings of the ACM In- ternational Conference on Multimedia, pages 2526–2534, 2020. 16.Faegheh Sardari, Adeline Paiement, Sion Hannuna, and Majid Mirmehdi. Vi-net—view-invariant quality of human movement assessment. Sensors, 20(18):5258, 2020. 17. Shunli Wang, Dingkang Yang, Peng Zhai, Chixiao Chen, and Lihua Zhang. Tsa-net: Tube self-attention network for action quality assessment. In Proceedings of the ACM In- ternational Conference on Multimedia, pages 4902–4910, 2021. 18.Paritosh Parmar, Amol Gharat, and Helge Rhodin. Do- main knowledge-informed self-supervised representations for workout form assessment. In European Conference on Computer Vision, pages 105–123. Springer, 2022. 19.Jinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen, Jie Zhou, and Jiwen Lu. Finediving: A fine-grained dataset for procedure-aware action quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2949–2958, 2022. 20.Shiyi Zhang, Wenxun Dai, Sujia Wang, Xiangwei Shen, Jiwen Lu, Jie Zhou, and Yansong Tang. Logo: A long- form video dataset for group action quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2405–2414, 2023. 21.Yanli Ji, Lingfeng Ye, Huili Huang, Lijing Mao, Yang Zhou, and Lingling Gao. Localization-assisted uncertainty score disentanglement network for action quality assessment. In Proceedings of the ACM International Conference on Multimedia, pages 8590–8597, 2023. 22.Jibin Gao, Jia-Hui Pan, Shao-Jie Zhang, and Wei-Shi Zheng. Automatic modelling for interactive action as- sessment. International Journal of Computer Vision, 131(3):659–679, 2023. 23. Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jing-Ke Meng, and Wei-Shi Zheng. Egoexo-fitness: Towards egocentric and exocentric full-body action under- standing. In European Conference on Computer Vision, 2024. 24.Linfeng Dong, Wei Wang, Yu Qiao, and Xiao Sun. Lu- cidaction: A hierarchical and multi-model dataset for comprehensive action quality assessment. In Advances in Neural Information Processing Systems, 2024. 25.Kaili Zheng, Kaiwen Wang, Xun Zhu, Qingyuan Yang, Chenyi Guo, and Ji Wu. Fitaqa: A benchmark of fitness action quality assessment for multimodal large language models. arXiv preprint arXiv:2608.08736, 2026. 26. Yixin Gao, S Swaroop Vedula, Carol E Reiley, Narges Ahmidi, Balakrishnan Varadarajan, Henry C Lin, Lingling Tao, Luca Zappella, Benjamın B ́ejar, and David D Yuh. Jhu-isi gesture and skill assessment working set (jigsaws): A surgical activity dataset for human motion modeling. In Medical Image Computing and Computer Assisted Intervention Workshop, volume 3, page 3, 2014. 27. Zijian Chen, Wei Sun, Yuan Tian, Jun Jia, Zicheng Zhang, Jiarui Wang, Ru Huang, Xiongkuo Min, Guangtao Zhai, and Wenjun Zhang. Gaia: Rethinking action quality assessment for ai-generated videos. In Advances in Neural Information Processing Systems, 2024. 28.Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21096–21106, 2022. 29.Hazel Doughty, Dima Damen, and Walterio Mayol-Cuevas. Who’s better? who’s best? pairwise deep ranking for skill determination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6057– 6066, 2018. 30.Hazel Doughty, Walterio Mayol-Cuevas, and Dima Damen. The pros and cons: Rank-aware temporal attention for skill determination in long videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7862–7871, 2019. 31. Shijia Feng, Michael Wray, Brian Sullivan, Youngkyoon Jang, Casimir Ludwig, Iain Gilchrist, and Walterio Mayol- Cuevas. Are you struggling? dataset and baselines for struggle determination in assembly videos: S. feng et al. International Journal of Computer Vision, 133(11):7817– 7854, 2025. 32.Yuyang Ji, Yixuan Shen, Shengjie Zhu, Yu Kong, and Feng Liu. From 3d pose to prose: Biomechanics-grounded vision- language coaching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23506–23515, 2026. 33.Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19383–19400, 2024. 34. Ministry of Human Resources and Social Security of the People’s Republic of China and General Administration of Sport of China. National occupational skill standard — social sports instructor (occupational code: 4-13-04- 01). Standard, Ministry of Human Resources and Social Security of the People’s Republic of China and General Administration of Sport of China, 2020. 35. Beijing Sport University. Fitness and Bodybuilding Tuto- rial. Beijing Sport University Press, 2013. 36.Human Resources Development Center of the General Ad- ministration of Sport of China. Occupational Competency Training Textbook for Social Sports Instructors—Fitness Coaches (with Technical Action Videos). Higher Educa- tion Press, 2023. 37.Joe Weider. Joe Weider’s Bodybuilding System. Weider Pubns, 1998. 38.Hong-Bo Zhang, Li-Jia Dong, Qing Lei, Li-Jie Yang, and Ji-Xiang Du. Label-reconstruction-based pseudo-subscore learning for action quality assessment in sporting events. Applied Intelligence, 53(9):10053–10067, 2023. 39. Angchi Xu, Ling-An Zeng, and Wei-Shi Zheng. Likert scor- ing with grade decoupling for long-term action assessment. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 3232–3241, 2022. 40. Qi An, Mengshi Qi, and Huadong Ma. Multi-stage con- trastive regression for action quality assessment. In IEEE 36 International Conference on Acoustics, Speech and Signal Processing, pages 4110–4114. IEEE, 2024. 41.Xiao Ke, Huangbiao Xu, Xiaofeng Lin, and Wenzhong Guo. Two-path target-aware contrastive regression for ac- tion quality assessment. Information Sciences, 664:120347, 2024. 42.Jinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang, and Yuxin Peng. Fineparser: A fine-grained spatio-temporal action parser for human-centric action quality assessment. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14628–14637, 2024. 43.Jia-Hui Pan, Jibin Gao, and Wei-Shi Zheng. Action assessment by joint relation graphs. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6331–6340, 2019. 44. Xin Chen, Anqi Pang, Wei Yang, Yuexin Ma, Lan Xu, and Jingyi Yu. Sportscap: Monocular 3d human motion capture and fine-grained understanding in challenging sports videos. International Journal of Computer Vision, 129:2846–2864, 2021. 45.Lauren Okamoto and Paritosh Parmar. Hierarchical neurosymbolic approach for comprehensive and explain- able action quality assessment. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 3204–3213, 2024. 46.Konstantinos Roditakis, Alexandros Makris, and Anto- nis Argyros. Towards improved and interpretable action quality assessment with self-supervised alignment. In Proceedings of the PErvasive Technologies Related to As- sistive Environments Conference, pages 507–513, 2021. 47.Paritosh Parmar, Eric Peh, and Basura Fernando. Learn- ing to visually connect actions and their effects. arXiv preprint arXiv:2401.10805, 2024. 48.Yansong Tang, Zanlin Ni, Jiahuan Zhou, Danyang Zhang, Jiwen Lu, Ying Wu, and Jie Zhou. Uncertainty-aware score distribution learning for action quality assessment. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 9839–9848, 2020. 49.Caixia Zhou, Yaping Huang, and Haibin Ling. Uncertainty-driven action quality assessment.arXiv preprint arXiv:2207.14513, 2022. 50. Boyu Zhang, Jiayuan Chen, Yinfei Xu, Hui Zhang, Xu Yang, and Xin Geng. Auto-encoding score distri- bution regression for action quality assessment. Neural Computing and Applications, 36(2):929–942, 2024. 51.Xumin Yu, Yongming Rao, Wenliang Zhao, Jiwen Lu, and Jie Zhou. Group-aware contrastive regression for action quality assessment. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 7919– 7928, 2021. 52. Paritosh Parmar, Jaiden Reddy, and Brendan Morris. Piano skills assessment. In IEEE International Workshop on Multimedia Signal Processing, pages 1–5. IEEE, 2021. 53.Qing Lei, Huiying Li, Hongbo Zhang, Jixiang Du, and Shangce Gao. Multi-skeleton structures graph convolu- tional network for action quality assessment in long videos. Applied Intelligence, 53(19):21692–21705, 2023. 54.Ling-An Zeng and Wei-Shi Zheng. Multimodal action quality assessment. IEEE Transactions on Image Pro- cessing, 2024. 55.Chi Hsuan Wu, Kumar Ashutosh, and Kristen Grauman. Skillsight: Efficient first-person skill assessment with gaze. arXiv preprint arXiv:2511.19629, 2025. 56.Shiyi Zhang, Sule Bai, Guangyi Chen, Lei Chen, Jiwen Lu, Junle Wang, and Yansong Tang. Narrative action evaluation with prompt-guided multimodal interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18430–18439, 2024. 57.Kumie Gedamu, Yanli Ji, Yang Yang, Jie Shao, and Heng Tao Shen. Visual-semantic alignment temporal parsing for action quality assessment. IEEE Transactions on Circuits and Systems for Video Technology, 2024. 58.James Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau, Alejandro Lozano, Lisa Dunlap, Trevor Darrell, and Ser- ena Yeung-Levy. Video action differencing. arXiv preprint arXiv:2503.07860, 2025. 59.Hitoshi Matsuyama, Nobuo Kawaguchi, and Brian Y Lim. Iris: Interpretable rubric-informed segmentation for action quality assessment. In Proceedings of the International Conference on Intelligent User Interfaces, pages 368–378, 2023. 60.Abrar Majeedi, Viswanatha Reddy Gajjala, Satya Sai Srinath Namburi GNVV, and Yin Li. Rica2: Rubric- informed, calibrated assessment of actions. arXiv preprint arXiv:2408.02138, 2024. 61.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question an- swering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017. 62.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal rea- soning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 2758–2766, 2017. 63.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yuet- ing Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answer- ing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9127–9134, 2019. 64. Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. Tvqa+: Spatio-temporal grounding for video question answering. arXiv preprint arXiv:1904.11574, 2019. 65.Paritosh Parmar, Eric Peh, Ruirui Chen, Ting En Lam, Yuhan Chen, Elston Tan, and Basura Fernando. Causalchaos! dataset for comprehensive causal action question answering over longer causal chains grounded in dynamic visual scenes. Advances in Neural Information Processing Systems, 37:92769–92802, 2024. 66.Yi Zhang, Peiyang Li, Xuyang Zhu, Steven W Su, Qing Guo, Peng Xu, and Dezhong Yao. Extracting time- frequency feature of single-channel vastus medialis emg signals for knee exercise pattern recognition. PloS one, 12(7):e0180526, 2017. 67.Fran ̧cois Hug, Cl ́ement Vogel, Kylie Tucker, Sylvain Dorel, Thibault Deschamps, ́ Eric Le Carpentier, and Lilian La- courpaille. Individuals have unique muscle activation signatures as revealed during gait and pedaling. Journal of Applied Physiology, 127(4):1165–1174, 2019. 68.N ́estor J Jarque-Bou, Manfredo Atzori, and Henning M ̈uller. A large calibrated database of hand movements and grasps kinematics. Scientific data, 7(1):12, 2020. 69. Asad Mansoor Khan, Sajid Gul Khawaja, Muhammad Us- man Akram, and Ali Saeed Khan. semg dataset of routine activities. Data in brief, 33:106543, 2020. 70.Hristo Dimitrov, Anthony M. J. Bull, and Dario Farina. High-density EMG, IMU, kinetic, and kinematic open- source data for comprehensive locomotion activities. Sci- entific Data, 10(1):1–10, 2023. 71. Huawei Wang, Akash Basu, Guillaume Durandau, and Massimo Sartori. A wearable real-time kinetic measure- ment sensor setup for human locomotion. Wearable tech- nologies, 4:e11, 2023. 37 72. Manfredo Atzori, Arjan Gijsberts, Claudio Castellini, Barbara Caputo, Anne-Gabrielle Mittaz Hager, Simone Elsig, Giorgio Giatsidis, Franco Bassetto, and Henning M ̈uller. Electromyography data for non-invasive naturally- controlled robotic hand prostheses. Scientific data, 1(1):1– 13, 2014. 73.Yilin Liu, Shijia Zhang, and Mahanth Gowda. Neuropose: 3d hand pose tracking using emg wearables. In Proceedings of the Web Conference, pages 1471–1482, 2021. 74. Mehmet Akif Ozdemir, Deniz Hande Kisa, Onan Guren, and Aydin Akan. Dataset for multi-channel surface elec- tromyography (semg) signals of hand gestures. Data in brief, 41:107921, 2022. 75.Sasha Salter, Richard Warren, Collin Schlager, Adrian Spurr, Shangchen Han, Rohin Bhasin, Yujun Cai, Peter Walkington, Anuoluwapo Bolarinwa, Robert J Wang, et al. emg2pose: A large and diverse benchmark for surface electromyographic hand pose estimation. Advances in Neural Information Processing Systems, 37:55703–55728, 2024. 76.Kunyu Peng, David Schneider, Alina Roitberg, Kailun Yang, Jiaming Zhang, Chen Deng, Kaiyu Zhang, M. Saquib Sarfraz, and Rainer Stiefelhagen. Towards video-based activated muscle group estimation in the wild. In Proceedings of the 32nd ACM International Confer- ence on Multimedia, M ’24, page 4495–4504, New York, NY, USA, 2024. Association for Computing Machinery. 77. ZCAM. Zcam e2-m4, 2020. 78. OLYMPUS. M.zuiko digital ed 14-150m f4.0-5.6, 2022. 79. OnePlus. Oneplus 7, 2019. 80.Marianna Capecci, Maria Gabriella Ceravolo, Francesco Ferracuti, Sabrina Iarlori, Andrea Monteriu, Luca Romeo, and Federica Verdini. The kimore dataset: Kinematic assessment of movement and clinical scores for remote monitoring of physical rehabilitation. IEEE Transac- tions on Neural Systems and Rehabilitation Engineering, 27(7):1436–1448, 2019. 81.Yansong Tang, Jinpeng Liu, Aoyang Liu, Bin Yang, Wenxun Dai, Yongming Rao, Jiwen Lu, Jie Zhou, and Xiu Li. Flag3d: A 3d fitness activity dataset with language instruction. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pages 22106–22117, 2023. 82.Mengshi Qi, Yeteng Wu, Xianlin Zhang, and Huadong Ma. Explainable action form assessment by exploiting multimodal chain-of-thoughts reasoning. arXiv preprint arXiv:2512.15153, 2025. 83.Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang. Disentangling and unifying graph convolutions for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 143–152, 2020. 84. Yuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li, Ying Deng, and Weiming Hu. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 13359–13368, 2021. 85.Haodong Duan, Jiaqi Wang, Kai Chen, and Dahua Lin. Pyskl: Towards good practices for skeleton action recog- nition. In Proceedings of the 30th ACM international conference on multimedia, pages 7351–7354, 2022. 86.Jeonghyeok Do and Munchurl Kim. Skateformer: skeletal- temporal transformer for human action recognition. In European Conference on Computer Vision, pages 401–420. Springer, 2024. 87.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024. 88.Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479, 2025. 89.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. 90.Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. Lla- mafactory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372, 2024. 91.Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002. 92.Chin-Yew Lin and Franz Josef Och. Automatic evaluation of machine translation quality using longest common sub- sequence and skip-bigram statistics. In Proceedings of the 42nd annual meeting of the association for computational linguistics (ACL-04), pages 605–612, 2004. 93.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005. 94. Maja Popovi ́c. chrf++: words helping character n-grams. In Proceedings of the second conference on machine trans- lation, pages 612–618, 2017. 95.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Wein- berger, and Yoav Artzi. Bertscore: Evaluating text gener- ation with bert. In International Conference on Learning Representations. 96.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 97. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 98.Sepp Hochreiter and J ̈urgen Schmidhuber. Long short- term memory. Neural computation, 9(8):1735–1780, 1997. 99.Harris Drucker, Christopher J Burges, Linda Kaufman, Alex Smola, and Vladimir Vapnik. Support vector regres- sion machines. Advances in neural information processing systems, 9, 1996.