Paper deep dive
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 2:18:48 PM
Summary
This study evaluates the feasibility of using Vision-Language Models (VLMs) to non-invasively estimate horizontal (H) and vertical (V) hand distances required by the Revised NIOSH Lifting Equation (RNLE) from RGB video. Two multi-stage pipelines were developed: a detection-only pipeline (GD-Dv2) and a detection-plus-segmentation pipeline (GD-SAM-Dv2). The results indicate that the segmentation-based, multi-view pipeline consistently yields smaller errors, achieving mean absolute errors of approximately 6-8 cm for H and 5-8 cm for V. Pixel-level segmentation reduced estimation error by 20-30% for H and 35-40% for V compared to the detection-only approach.
Entities (10)
Relation Signals (6)
Revised NIOSH Lifting Equation â requires â Horizontal Distance
confidence 95% ¡ The RNLE relies on six task-specific parametersâincluding horizontal and vertical hand distances
Revised NIOSH Lifting Equation â requires â Vertical Distance
confidence 95% ¡ The RNLE relies on six task-specific parametersâincluding horizontal and vertical hand distances
Vision-Language Models â usedfor â Ergonomic Risk Assessment
confidence 95% ¡ We evaluated the feasibility of using innovative vision-language models (VLMs) to non-invasively estimate H and V from RGB video streams.
GD-SAM-Dv2 â outperforms â GD-Dv2
confidence 92% ¡ the segmentation-based, multi-view pipeline consistently yielding the smallest errors... pixel-level segmentation reduced estimation error by approximately 20â30% for H and 35â40% for V relative to the detection-only pipeline.
Segment Anything Model â componentof â GD-SAM-Dv2
confidence 90% ¡ In the GDâSAMâDv2 pipeline, detected bounding boxes were further refined using the SAM
Grounding DINO â componentof â GD-Dv2
confidence 90% ¡ GDâDv2, a Grounding DINOâbased detection-only pipeline
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Manual lifting tasks are a major contributor to work-related musculoskeletal disorders, and effective ergonomic risk assessment is essential for quantifying physical exposure and informing ergonomic interventions. The Revised NIOSH Lifting Equation (RNLE) is a widely used ergonomic risk assessment tool for lifting tasks that relies on six task variables, including horizontal (H) and vertical (V) hand distances; such distances are typically obtained through manual measurement or specialized sensing systems and are difficult to use in real-world environments. We evaluated the feasibility of using innovative vision-language models (VLMs) to non-invasively estimate H and V from RGB video streams. Two multi-stage VLM-based pipelines were developed: a text-guided detection-only pipeline and a detection-plus-segmentation pipeline. Both pipelines used text-guided localization of task-relevant regions of interest, visual feature extraction from those regions, and transformer-based temporal regression to estimate H and V at the start and end of a lift. For a range of lifting tasks, estimation performance was evaluated using leave-one-subject-out validation across the two pipelines and seven camera view conditions. Results varied significantly across pipelines and camera view conditions, with the segmentation-based, multi-view pipeline consistently yielding the smallest errors, achieving mean absolute errors of approximately 6-8 cm when estimating H and 5-8 cm when estimating V. Across pipelines and camera view configurations, pixel-level segmentation reduced estimation error by approximately 20-30% for H and 35-40% for V relative to the detection-only pipeline. These findings support the feasibility of VLM-based pipelines for video-based estimation of RNLE distance parameters.
Tags
Links
- Source: https://arxiv.org/abs/2602.20658v1
- Canonical: https://arxiv.org/abs/2602.20658v1
Trouble viewing inline? Open PDF directly â
Full Text
79,201 characters extracted from source content.
Expand or collapse full text
1 VisionâLanguage Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video a Mohammad Sadra Rajabi, https://orcid.org/0000-0002-9100-3973 b Aanuoluwapo Ojelade, https://orcid.org/0000-0001-9715-3254 a Sunwook Kim, https://orcid.org/0000-0003-3624-1781 a Maury A. Nussbaum, https://orcid.org/0000-0002-1887-8431 Affiliations: a Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg VA 24061, USA b St. Jude Children's Research Hospital, Memphis, TN 38105, USA Corresponding author: Corresponding address: Maury A. Nussbaum Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg VA 24061, USA. Phone: 540-231-6053. Email: nussbaum@vt.edu 2 Abstract Manual lifting tasks are a major contributor to work-related musculoskeletal disorders, and effective ergonomic risk assessment is essential for quantifying physical exposure and informing ergonomic interventions. The Revised NIOSH Lifting Equation (RNLE) is a widely used ergonomic risk assessment tool for lifting tasks that relies on six task variables, including horizontal (H) and vertical (V) hand distances; such distances are typically obtained through manual measurement or specialized sensing systems and are difficult to use in real-world environments. We evaluated the feasibility of using innovative visionâlanguage models (VLMs) to non-invasively estimate H and V from RGB video streams. Two multi-stage VLM-based pipelines were developed: a text-guided detection-only pipeline and a detection-plus- segmentation pipeline. Both pipelines used text-guided localization of task-relevant regions of interest, visual feature extraction from those regions, and transformer-based temporal regression to estimate H and V at the start and end of a lift. For a range of lifting tasks, estimation performance was evaluated using leave-one-subject-out validation across the two pipelines and seven camera view conditions. Results varied significantly across pipelines and camera view conditions, with the segmentation-based, multi-view pipeline consistently yielding the smallest errors, achieving mean absolute errors of approximately 6â8 cm when estimating H and 5â8 cm when estimating V. Across pipelines and camera view configurations, pixel-level segmentation reduced estimation error by approximately 20â30% for H and 35â40% for V relative to the detection-only pipeline. These findings support the feasibility of VLM-based pipelines for video- based estimation of RNLE distance parameters. Keywords: VisionâLanguage Models (VLMs); Revised NIOSH Lifting Equation (RNLE); Computer vision; Ergonomic risk assessment; Segmentation; Multi-view video 3 1.0 Introduction Work-related musculoskeletal disorders (WMSDs) remain a major occupational health concern worldwide and are among the leading causes of lost workdays, reduced productivity, and substantial economic burden for employers and workers (Govaerts et al., 2021; Liberty Mutual Insurance, 2023; U.S. Bureau of Labor Statistics, 2024). Manual material handling (MMH) tasksâsuch as lifting, lowering, carrying, pushing, and pullingâare a primary contributor to WMSDs, as they expose workers to biomechanical risk factors including forceful exertions, non- neutral postures, repetitive motions, and prolonged task durations (Da Costa & Vieira, 2010; Waters et al., 1993). Epidemiological evidence consistently demonstrates elevated incidence rates of low back and upper-extremity disorders among workers engaged in MMH-intensive occupations across manufacturing, logistics, healthcare, and construction sectors (Punnett & Wegman, 2004). Reducing the burden of WMSDs therefore requires accurate and practical methods for quantifying physical exposures and injury risks during MMH tasks in real-world work environments. Ergonomic risk assessment tools, such as the Revised NIOSH Lifting Equation (RNLE; Waters et al., 1993), are widely used to evaluate the physical demands of manual lifting tasks and to guide workplace interventions. The RNLE relies on six task-specific parametersâincluding horizontal and vertical hand distances, load weight, asymmetry, coupling quality, and task frequencyâthat directly affect biomechanical loading and WMSD risk (Waters et al., 1993). Accurate measurement of these parameters is critical, as measurement errors have been shown to affect RNLE outputs and risk estimates, potentially influencing risk classification and subsequent intervention decisions (Fox et al., 2019; Waters et al., 1998). In practice, though, RNLE parameters are commonly obtained through manual measurement, which is time- consuming, subject to observer bias, and difficult to apply consistently across large or complex work environments (David, 2005; Dempsey et al., 2001, 2005; Lu et al., 2016). Although wearable sensors (Hlucny & Novak, 2020; Lu et al., 2020; Mudiyanselage et al., 2021; Ranavolo et al., 2024) and marker-based motion capture systems (Gutierrez et al., 2024; Ranavolo et al., 2017) can provide accurate measurements, their cost, intrusiveness, and logistical complexity may limit their feasibility for practical field-based ergonomic assessments (Sabino et al., 2024; Schall et al., 2018). To address these limitations, computer visionâbased approaches have been increasingly explored for ergonomic assessment. Markerless human pose estimation methods using RGB or RGB-D videoâsuch as OpenPose (Zhao et al., 2023) or BlazePose (Bazarevsky et al., 2020)âcan estimate body joint locations and postures without physical instrumentation (Cao et al., 2016). These approaches have shown promise for posture classification (Chen, 2019; Hsu et al., 2025; Jung et al., 2022; Liu & Chang, 2022) and ergonomic risk assessment (Forgione et al., 2025; Lou et al., 2025; Pires et al., 2025), but their reported application to distance-based ergonomic parameters remains limited. Like most vision-based methods, pose-based systems are sensitive to occlusion, camera perspective, and lighting conditions; additionally, they often produce temporally unstable body joint keypoint trajectories across frames, which can degrade frame- level distance estimates (Cheng et al., 2019; Mehta et al., 2017; Pavllo et al., 2018; Veges & Lorincz, 2020; Zahabi et al., 2026). Moreover, such systems primarily focus on skeletal representations and do not explicitly encode object-centric spatial relationshipsâsuch as those between the worker, load, and environmentâthat are fundamental to ergonomic risk modeling 4 (Bezzini et al., 2023; Parsa et al., 2019; Paudel et al., 2022). As a result, pose-based approaches may struggle to provide robust estimates of the geometric parameters required for ergonomic risk assessment tools such as the RNLE. Recent advances in visionâlanguage models (VLMs) offer clear potential to overcome these challenges by enabling joint reasoning about objects, actions, and spatial relationships (e.g., relative position and distance) between workers, handled loads, and the environment within complex scenes (Guran et al., 2024; Liao et al., 2024; Zhang et al., 2024). VLMs integrate visual encoders with language-based representations, allowing models to be prompted to directly identify and localize semantically meaningful entitiesâsuch as the worker, relevant body segments, handled loads, and toolsâusing natural-language queries, rather than relying solely on skeletal representations (Chen et al., 2024; Kirillov et al., 2023; Liu et al., 2024). This combination of text-guided object localization and object-centric visual representations suggests that VLM-based approaches may be well suited for estimating distance-based ergonomic parameters, since these representations facilitate explicit modeling of humanâobject and objectâ environment relationships that are difficult to capture using pose-based methods alone (Kang et al., 2025; Wen et al., 2025). By combining text-guided detection and segmentation with geometric reasoning, VLMs can provide spatially grounded representations of key regions of interest that correspond directly to the definitions of horizontal and vertical distances in the RNLE. However, as visibility of task-relevant landmarks can vary with camera viewpoint, it is important to evaluate the effect of camera view conditions on VLM-based distance estimation. Despite the integration of VLMs in general computer vision and emerging applications in construction safety (Fan, Mei, Wang, et al., 2024) and posture and task classification (Fan, Mei, & Li, 2024; Rajabi et al., 2025; Yong et al., 2024), the use of VLMs for quantitative estimation of ergonomic exposure metrics, particularly RNLE parameters remains unexplored. Furthermore, the effect of camera view conditions on such estimates has not yet been reported to our knowledge. In the current study, we evaluated the feasibility of using VLMs to estimate the horizontal and vertical distances required by the RNLE from RGB video streams. We examined the estimation performance of VLM-based pipelines that integrate text-guided object detection and segmentation, visual feature extraction, and a transformer-based distance regression model across varied camera view conditions. We hypothesized that estimation performance would vary across VLM-based pipelines and camera view conditions. While our immediate goal was to assess VLM performance under simulated lifting tasks, this work also provides insight into the potential of video-based approaches for estimating key parameters used in ergonomic risk assessment tools across occupational settings. 2.0 Methods 2.1 Overview of the Dataset We used a subset of data obtained from a prior study by Ojelade et al. (2025), in which 32 healthy young adults (19 males and 13 females; see Appendix A.1 for participant demographics and inclusion criteria) performed eight manual handling tasks. In the subset of data used here, participants performed symmetric lifts from two distinct Lift Originsâthe floor and individual knee heightâup to individual hip height. Hip height was defined as the vertical distance from the floor to the greater trochanter. Participants completed these lifts using two Hand Configurations (broad = 52 cm handle spacing and narrow = 33 cm; see Appendix A.2 for 5 symmetric lifting task details and Figure A.1 for an illustration of the hand configurations), three Box Mass levels (6, 9, and 12 kg), with each condition performed twice (2 Lift Origins Ă 2 Hand Configurations Ă 3 Box Masses Ă 2 replications = 24 trials per participant; see Appendix A.3 for experimental procedures). Whole-body kinematics were recorded using a multimodal instrumentation setup, including three synchronized Azure Kinect⢠cameras (Microsoft Corporation, Seattle, WA, USA) and a wearable inertial measurement unit (IMU)-based motion capture (Noraxon Ultium, Noraxon, Scottsdale, AZ, USA). The cameras recorded at 30 Hz and were positioned approximately 1.74 m from edge of the work area (see Appendix A.4 for camera system instrumentation). The three cameras provided distinct viewpoints of the lifting task and are hereafter referred to as Camera View 1 (V1), Camera View 2 (V2), and Camera View 3 (V3). V1 and V2 corresponded to two oblique views of the workspace from opposite sides, whereas V3 provided a frontal view of the participants. To assess the effect of Camera View Condition, seven camera view conditions were defined based on the available RGB streams: three single-view conditions (i.e., V1, V2, and V3) and four multi-view conditions formed by combining synchronized views (i.e., V1+V2, V1+V3, V2+V3, and V1+V2+V3). For multi-view conditions, the synchronized views were processed in parallel and combined for model input and evaluation. The IMU system consisted of 16 sensors, positioned on the head, upper and lower thoracic spines, pelvis, and bilaterally on the upper and lower arms, hands, thighs, shanks, and feet; IMU data processing is described in Appendix A.5. For the present study, only the RGB video streams (1280 Ă 720 pixels) from the Kinects were used as input to the VLM-based pipeline, while the IMU data were used exclusively to derive reference kinematic measurements for ground-truth distance labeling (see Data Labeling section). 2.2 Overview of the VLM-Based Pipelines for Horizontal and Vertical Distance Estimation We developed and evaluated two VLM-based, multi-stage pipelines to estimate RNLE horizontal (H) and vertical (V) distances from RGB video (Figure 1). The two pipelines are hereafter referred to as GDâDv2 (Grounding DINO with DINOv2 features; detection-only) and GDâ SAMâDv2 (detection followed by segmentation). For ground-truth labeling, frame-level H and V values were derived from processed IMU-based kinematic data and aligned to RGB video frames. The VLM-based pipelines then consisted of three primary steps: Step 1 was region-of- interest processing, in which task-relevant regions of interest (ROIs) are detectedâand, when applicable, segmentedâfrom RGB video frames. Step 2 was feature extraction, in which detected ROIs are transformed into feature representations. Step 3 was distance estimation, in which temporally ordered feature sequences are processed using a transformer-based regression model to estimate H and V at the frame level. Each of these steps is described in detail in the following sections. All processing was performed offline using Python (v3.10.11; https://w.python.org) on secure Virginia Tech Advanced Research Computing resources (VT ARC; https://arc.vt.edu/), with no participant videos, images, or derived data uploaded to or processed on external or cloud-based systems. RGB video handling and frame extraction were implemented using OpenCV (https://opencv.org/), and model inference and training were conducted using PyTorch (https://pytorch.org/, Paszke et al., 2019). 6 Figure 1: Overview of two pipelines for estimating the horizontal (H) and vertical (V) distances defined in the RNLE at the start and end of a lift. A text-guided detection process was used to localize the lifter and task-relevant regions of interest (ROIs) using Grounding DINO. In the GDâSAMâDv2 pipeline, detected ROIs were further refined using bounding-boxâguided segmentation with SAM. ROI features extracted using DINOv2 were processed by a transformer-based temporal regression model to estimate H and V at the start and end of a lift. 7 2.3 Data Labeling As introduced earlier, the RNLE combines six lifting task-related parameters (Appendix A.6), including the horizontal and vertical distances of the hands (i.e., H and V; Waters et al., 1993; Figure 2). H is defined as the horizontal distance between the midpoint of the hands and the midpoint of the ankles, projected onto the floor plane (Waters et al., 1993, 1998). Here, for each time sample in the IMU data, the horizontal position of the hands was calculated as the midpoint of the left and right hand-tip trajectories in the anteriorâposterior and mediolateral directions, while the horizontal position of the ankles was calculated as the midpoint of the left and right medial malleoli. H was measured based on the Euclidean distance between these two midpoints in the horizontal plane. V is defined as the vertical position of the hands relative to the floor (Waters et al., 1993, 1998). Here, the vertical position of the hands was calculated as the mean of the left and right hand-tip vertical coordinates, since the floor vertical coordinates were set to zero. Each RGB video frame was assigned corresponding ground-truth H and V values by temporally aligning the distances computed from IMU-derived kinematic data with the RGB video recordings at the frame level. These distances (i.e., H and V) were used exclusively as reference labels for model training and evaluation. Figure 2: Illustration of the horizontal (H) and vertical (V) distances defined in the RNLE at the start (left) and end (right) of a lift. Distance H is measured from the midpoint between the ankles to a point projected on the floor directly below the midpoint between the hands, whereas distance V is measured from the floor to the midpoint between the hands (Image generated by ChatGPT, OpenAI, https://openai.com/index/chatgpt/). 2.4 Detection and Segmentation of Regions of Interest (ROIs) To identify task-relevant ROIs required for estimating RNLE distance parameters from RGB video, two VLMâbased pipelines were evaluated. Both pipelines relied on text-guided object localization to detect the primary âlifterâ, as well as relevant body parts and objects, but differed in whether pixel-level segmentation was applied following detection. The two pipelines consisted of: (1) GDâDv2, a Grounding DINOâbased detection-only pipeline (Liu et al., 2024); and (2) GDâSAMâDv2, a Grounding DINOâbased detection followed by segmentation using the Segment Anything Model (SAM: ViT-H backbone; Kirillov et al., 2023). Grounding DINO was used for zero-shot, text-guided object detection, whereas SAM provided zero-shot, promptable 8 segmentation to refine detected ROIs at the pixel level, with neither model requiring task- specific retraining (see Figures 3, A.2, and A.3 for examples of the ROI detection and segmentation process across camera views). In both pipelines, a two-stage detection strategy was applied for each RGB video frame. First, the âlifterâ was identified using the text prompt âperson liftingâ, and the detection with the highest confidence score was selected when multiple candidates were present. Second, a spatial crop, centered on the detected lifter, was generated to reduce background interference and to improve localization of fine-grained elements. Within this cropped region, task-relevant details were detected using the prompt âhand . wrist . shoe . wooden box . crate . holding object â, allowing detection and localization of ROIs. These prompts were selected to align with RNLE- relevant anatomical landmarks and task objects, which was intended to support zero-shot detection. In the GDâDv2 pipeline, the resulting bounding boxes from this detection process were used directly as ROIs for subsequent feature extraction and distance estimation. This approach relied solely on box-level localization to represent task-relevant regions. In the GDâSAMâDv2 pipeline, detected bounding boxes were further refined using the SAM (Kirillov et al., 2023). For each detected ROI, SAM generated a pixel-level segmentation mask guided by the bounding box generated by Grounding DINO (Liu et al., 2024). These masks were used to isolate the detected ROIs from the surrounding background, excluding background pixels and more precisely capturing the shape and spatial extent of each ROI. The segmented regions were then used for feature extraction. 9 Figure 3: Example visualization of the region-of-interest (ROI) detection and segmentation process for a representative lifting frame of Camera View 1 (V1). Top left: Original RGB video frame. Top right: Detection of the primary lifter using Grounding DINO, with the bounding box corresponding to the highest-confidence person detection. Bottom left: Detection of task-relevant ROIs within a cropped region centered on the lifter using Grounding DINO. Bottom right: Pixel-level segmentation of the detected ROIs produced by the GDâSAMâDv2 pipeline. 10 2.5 Feature Extraction of Regions of Interest For each RGB video frame, each ROI was cropped from the original RGB frame based on the detected bounding box or segmentation mask, resized to 224 Ă 224 pixels, and then normalized using ImageNet statistics (Deng et al., 2009). ImageNet channel-wise statistics were used to normalize RGB pixel values, which standardizes input intensity distributions across frames and ensures compatibility with the pretrained feature extractor (Deng et al., 2009). Visual features were extracted from all identified ROIs using the DINOv2 vision transformer (ViT-Base; Oquab et al., 2024) in a zero-shot configuration without task-specific fine-tuning. DINOv2 is a self- supervised vision transformer trained to produce general-purpose visual representations without reliance on labeled training data (Oquab et al., 2024). Feature extraction was performed identically for both pipelines, to ensure that any observed differences in performance were attributable to differences in ROI representation rather than feature encoding. The resulting 768- dimensional DINOv2 feature vector was extracted for each ROI. To incorporate task-relevant geometric context, DINOv2 visual features were augmented with object-related geometric features derived from the handled box used in the lifting task. Specifically, five additional features were appended to the DINOv2 feature vector: the width and height of the handled-object bounding box produced by Grounding DINO, normalized by the original frame width and height, respectively, and the known physical dimensions (width, depth, and height) of the box. The normalized image-plane dimensions were computed consistently in both pipelines and captured the apparent size within each frame, while the known physical dimensions provided a consistent real-world scale reference across frames. This reference was necessary to relate image-based representations to physical distances, given that monocular RGB images lack absolute scale information without an object of known size (JĂźngel et al., 2008). Together, these features resulted in a 773-dimensional representation for each frame, consisting of DINOv2 visual features augmented with object-related geometric information. 2.6 Regression-Based Distance Estimation Regression-based models were trained to estimate H and V from temporally ordered visual features extracted from RGB video frames. The regression process involved designing transformer-based model architecture, defining data handling procedures, training and validation strategies, and evaluating model performance, each of which is described below. 2.6.1 Model Architecture A transformer-based regression model was used to estimate H and V distances while explicitly modeling temporal dependencies across sequences of video frames. The model architecture consisted of a bidirectional transformer encoder designed to capture both short- and long-range temporal relationships within lifting sequences. Incorporating temporal context allowed the model to leverage information from neighboring frames, which can reduce the impact of transient occlusions, intermittent ROI detection or segmentation failures, and frame-level noise, while enforcing temporally consistent distance estimates across the lifting sequence (e.g., Arnab et al., 2019). Transformer-based architectures are well suited for this task, because they use self- attention mechanisms to model dependencies across an entire sequence, rather than relying solely on local temporal context (Dosovitskiy, 2020; Vaswani et al., 2017). Each input frame was represented by a 773-dimensional feature vector, consisting of DINOv2 visual features augmented with box-related geometric information (Section 2.5). These frame- 11 level feature vectors were first projected into a 512-dimensional latent space using a linear embedding layer. Sinusoidal positional encodings were then added to preserve temporal ordering within each sequence. The embedded sequences were processed by a stack of six transformer encoder layers, each comprising eight attention heads, feedforward sublayers, dropout regularization, and residual connections. The transformer output at each time step was passed through a regression head consisting of fully connected layers with nonlinear activation and dropout, producing simultaneous estimations of H and V for each frame. Although transformer- based regression model produces frame-level estimates of H and V for all frames in each sequence during training, estimation performance was evaluated only for the start and end frames of each lift, consistent with RNLE parameter definitions. 2.6.2 Data Handling, Model Training, and Validation Input data were organized into fixed-length sequences of 100 consecutive frames. For training, overlapping temporal windows were generated using a 50% stride (i.e., each window advanced by half its length) to increase the number of training samples and improve temporal generalization. For validation, fixed-length sequences corresponding to the start and end of each lift were extracted to align with RNLE parameter definitions. Variable-length sequences were handled through zero padding, and attention masks were applied to ensure that padded frames did not contribute to the loss. Ground-truth H and V values derived from IMU data (Section 2.3) were normalized prior to training by dividing by a fixed scalar (2000 m), which exceeds the maximum expected range of H and V in the dataset and ensured that normalized values remained within a consistent and numerically stable range across participants. Model training was performed for up to 100 epochs using the AdamW optimizer, with a learning-rate scheduler that reduced the learning rate when validation loss plateaued (e.g., Loshchilov & Hutter, 2019; Senior et al., 2013). Early stopping was applied to prevent overfitting (Prechelt, 1998); training was terminated if validation loss failed to improve over 15 consecutive epochs. When a new minimum validation loss was achieved, the model state was saved for subsequent evaluation. A leave-one-subject-out (LOSO) cross-validation strategy was used for model validation. Model training and validation were performed across 32 folds, with each fold using data from 31 participants for training and data from the remaining one participant held out for validation. This validation approach accounts for inter-individual variability and reflects deployment scenarios in which models trained on known individuals are applied to an unseen worker (Gholamiangonabadi et al., 2020). 2.6.3 Evaluation Metrics Regression performance was evaluated separately for the start and end of a lift by comparing predicted and ground-truth H and V values. Model performance was quantified using three standard regression metrics for H and V (see Appendix A.7 for evaluation metric definitions and computation details): mean absolute error (MAE), root mean square error (RMSE), and maximum absolute error (MaxAE). For each camera-view condition and VLM-based pipeline, performance metrics were computed for each LOSO fold. 2.7 Statistical Analyses To assess the effects of Pipeline and Camera View Condition on H and V estimation performance at the start and end of a lift, separate two-way repeated-measures analyses of variance (RANOVAs) were performed for each error metric. For each RANOVA model, biological sex 12 (Sex) was included as a blocking effect. Where relevant, significant interaction effects were explored using simple-effects testing, and post hoc paired comparisons were completed using the Tukeyâs HSD procedure. All statistical analyses were performed with JMP Pro 18 (SAS, Cary, NC) using the restricted maximum likelihood (REML) method. To obtain normal distributions of model residuals, logarithmic transformations were applied to six dependent variablesâMAE of H at the start and end of a lift, MAE and RMSE of V at the end of a lift, and MaxAE of H and V at the end of a lift. Statistical significance was determined when p < 0.05, and summary data are reported as least-square means based on the statistical model fits. 3.0 Results All RANOVA results are summarized in Tables A.1 and A.2. Significant main or interaction effects of Pipeline and Camera View Condition were found across all regression performance metrics. More detailed results are provided below. 3.1 Start of a Lift: Estimation Errors of Horizontal and Vertical Distances Significant main effects of Pipeline and Camera View Condition were found across all error metrics (i.e., MAE, RMSE, and MaxAE) for both H and V estimation (Table A.1). Additionally, significant Pipeline Ă Camera View Condition interaction effects were found for MAE of both H and V estimations, RMSE of V estimation, and MaxAE of V estimation. Across all metrics, the GDâSAMâDv2 pipeline yielded significantly smaller errors vs. the GDâDv2 pipeline. When estimating H, mean values were MAE = ~7.2 cm vs. ~9.25 cm; RMSE = ~9.6 cm vs. ~12.1 cm; and MaxAE = ~20.7 cm vs. ~23.6 cm. When estimating V, mean values were MAE = ~14.5 cm vs. ~23.0 cm; RMSE = ~18.3 cm vs. ~27.7 cm; and MaxAE = ~40.7 cm vs. ~49.0 cm. Across all metrics, the V1+V2+V3 multi-view condition resulted in significantly smaller errors vs. single-view condition for both H and V estimation. When estimating H, V1+V2+V3 produced the smallest errors: MAE = ~6.2 cm; RMSE = ~8.1 cm; MaxAE = ~19.8 cm, while single-view conditions led to the largest errors, with V1 producing: MAE = ~10.68 cm; RMSE = ~13.5 cm; MaxAE = ~24.6 cm. When estimating V, V1+V2+V3 similarly generated the smallest errors: MAE = ~7.78 cm; RMSE = ~11.0 cm; MaxAE = ~37.2 cm, while single-view conditions yielded the greatest errors, with V2 producing MAE = ~27.73 cm and RMSE = ~32.6 cm, and V1 producing MaxAE = ~51.6 cm. Generally, the magnitude of the difference in estimation error between the two pipelines varied, and in some cases significantly, across camera view configurations (Figures 4, A.4, and A.5). When estimating H, the GDâSAMâDv2 pipeline with multi-view conditions consistently produced the smallest errors: MAE = ~6.0 cm for V1+V3, V2+V3, and V1+V2+V3 (Figure 4); RMSE = ~7.9â8.0 cm for V2+V3, V1+V3, V1+V2, and V1+V2+V3; and MaxAE = ~18.8â19.4 cm for V2+V3, V1+V2+V3, V1+V2, and V1+V3. In contrast, the GDâDv2 pipeline with the V1 single-view condition yielded the largest errors, with MAE = ~12.39 cm (Figure 4), RMSE = ~15.1 cm, and MaxAE = ~26.4 cm. 13 Figure 4: Significant interaction effect of Pipeline Ă Camera View Condition on MAE of H (top) and V (bottom) estimation at the start of a lift. Hereafter, for each condition, mean estimation errors were computed for each participant across all trials within each fold, and the box plots summarize the distribution of these participant-level mean errors. Based on post hoc paired comparisons, conditions that do not share a common letter are significantly different; this lettering convention is used for all subsequent figures. 14 When estimating V, the GDâSAMâDv2 pipeline with multi-view conditions yielded the smallest errors, with MAE = ~7.35 cm for V1+V2+V3 (Figure 4), RMSE = ~10.7 cm for V1+V2+V3 (Figure A.4), and MaxAE = ~35.6 cm for V1+V3 (Figure A.5). In contrast, the GDâDv2 pipeline with the V1 single-view condition produced the largest errors, with MAE = ~32.39 cm (Figure 4), RMSE = ~38.3 cm (Figure A.4), and MaxAE = ~58.0 cm (Figure A.5). The absolute difference between the smallest- and largest-error pipelineâcamera combinations was substantially larger for V estimation vs. for H estimation, with MAE = ~25.0 cm vs. ~6.4 cm (Figure 4), RMSE = ~27.6 cm vs. ~7.2 cm, and MaxAE = ~22.4 cm vs. ~7.6 cm, indicating greater variability in V estimation performance. Across camera view conditions, the magnitude of error reduction associated with the GDâSAMâ Dv2 pipeline was generally larger for single-view conditions vs. multi-view configurations. When estimating H, error reductions were larger for single-view vs. multi-view conditions, with MAE = ~2â3 cm vs. ~1â2 cm, RMSE = ~2â4 cm vs. ~1â3 cm, and MaxAE = ~3â5 cm vs. ~2â3 cm. When estimating V, error reductions were likewise larger for single-view vs. multi-view conditions, with MAE = ~8â15 cm vs. ~4â8 cm, RMSE = ~9â16 cm vs. ~5â10 cm, and MaxAE = ~8â13 cm vs. ~4â10 cm. 3.2 End of a Lift: Estimation Errors of Horizontal and Vertical Distances Significant main effects of Pipeline and Camera View Condition, as well as significant Pipeline Ă Camera View Condition interaction effects, were found across all error metrics for both H and V estimation, except for the main effect of Camera View Condition and the Pipeline Ă Camera View Condition interaction effect on MaxAE for H estimation (Table A.2). Across all metrics, the GDâSAMâDv2 pipeline yielded significantly smaller errors vs. the GDâDv2 pipeline. When estimating H, mean values were MAE = ~9.3 cm vs. ~13.2 cm, RMSE ~13.1 cm vs. ~17.0 cm, and MaxAE ~27.0 cm vs. ~30.5 cm. When estimating V, mean errors were MAE = ~7.5 cm vs. ~12.2 cm, RMSE = ~8.8 cm vs. ~13.5 cm, and MaxAE = ~18.0 cm vs. ~22.3 cm. Where significant main effects of Camera View Condition were found, the V1+V2+V3 multi- view configuration generally yielded smaller errors vs. single-view configurations for both H and V estimation. When estimating H, V1+V2+V3 resulted in the smallest errors, with MAE = ~8.1 cm and RMSE = ~10.9 cm, whereas single-view conditions produced larger errors, with the largest MAE and RMSE occurring for the V2 condition, in which MAE = ~14.4 cm and RMSE = ~18.5 cm. When estimating V, V1+V2+V3 produced the smallest errors across all metrics, with MAE = ~4.8 cm, RMSE = ~6.0 cm, and MaxAE = ~15.0 cm, whereas single-view conditions produced the largest errors, especially for the V1 condition, with MAE = ~15.5 cm, RMSE = ~16.7 cm, and MaxAE = ~24.9 cm. Generally, the magnitude of the difference in estimation error between the two pipelines varied, and in some cases significantly, across camera view configurations. When estimating H, the GDâ SAMâDv2 pipeline with multi-view configurations generally produced smaller errors, with MAE = ~7.4â7.6 cm for the V1+V2+V3 and V1+V2 conditions (Figure 5) and RMSE = ~10.6â 10.9 cm for the V1+V2+V3, V1+V3, V2+V3, and V1+V2 conditions (Figure A.6). In contrast, the GDâDv2 pipeline with single-view conditions produced the largest errors, most notably for the V1 condition, with MAE = ~16.7 cm (Figure 5) and RMSE = ~20.2 cm (Figure A.6). 15 Figure 5: Significant interaction effect of Pipeline Ă Camera View Condition on MAE of H (top) and V (bottom) estimation at the end of a lift. When estimating V, the GDâSAMâDv2 pipeline with multi-view configurations generally produced the smallest errors, with MAE = ~4.9â5.9 cm across all multi-view conditions (Figure 5), RMSE = ~6.0â7.2 cm across all multi-view conditions (Figure A.7), and MaxAE = ~14.9â 15.1 cm for the V1+V2+V3 and V1+V3 conditions (Figure A.8). In contrast, the GDâDv2 16 pipeline with single-view conditions produced the largest errors, most notably for the V1 condition, with MAE = ~22.7 cm (Figure 5), RMSE = ~23.1 cm (Figure A.7), and MaxAE = ~29.4 cm (Figure A.8). The absolute difference between the smallest- and largest-error pipelineâcamera combinations was larger for V estimation vs. H estimation, with MAE = ~17.8 cm vs. ~9.3 cm (Figure 5), RMSE = ~17.1 cm vs. ~9.6 cm (Figures A.6 and A.7), and MaxAE = ~14.6 cm vs. ~3.0 cm, indicating greater variability in V estimation performance. Across camera configurations, the magnitude of error reduction for the GDâSAMâDv2 pipeline was generally larger for single-view conditions vs. multi-view configurations. When estimating H, larger error reductions were found for single-view vs. multi-view conditions, with MAE = ~4â6 cm vs. ~1â2 cm, RMSE = ~5â9 cm vs. ~2â4 cm, and MaxAE = ~1â5 cm vs. ~1â4 cm. When estimating V, larger error reductions were likewise found for single-view vs. multi-view conditions, with MAE = ~8â13 cm vs. ~2â5 cm, RMSE = ~7â11 cm vs. ~2â6 cm, and MaxAE = ~5â9 cm vs. ~1â5 cm. 4.0 Discussion Our goal here was to evaluate the feasibility of using VLM-based pipelines to estimate H and V required by the RNLE from RGB video. Our findings provide new results that demonstrating that VLM-based pipelines combining text-guided object detection, pixel-level segmentation, and temporal modeling can provide effective and ambient estimates of RNLE distance parameters. Estimation performance varied across VLM-based pipelines and camera view conditions, though, with particularly notable differences between the start and end phases of lifting tasks. Below, we discuss key findings in relation to pipeline design, camera view condition, temporal dynamics, and practical applications for ergonomic assessment. 4.1 Effect of Pixel-Level Segmentation on H and V Estimation Performance The GDâSAMâDv2 pipeline, which used pixel-level segmentation, generally yielded smaller errors vs. the GDâDv2 pipeline across most error metrics for both H and V estimation at the start and end of a lift. Overall, segmentation produced consistent error reductions of ~20â30% when estimating H and ~35â40% when estimating V (Figures 4 and 5). These results are consistent with the notion that pixel-level segmentation provides more precise spatial representations of task-relevant regions than does bounding boxâbased localization alone (Kirillov et al., 2023; Ren et al., 2024). Bounding boxes inherently include background pixels and may fail to accurately capture the spatial extent of irregularly shaped objects or body segments, particularly when the lifter's posture varies substantially across frames or when the handled load partially occludes body segments. By contrast, SAM-generated segmentation masks isolate object boundaries at the pixel level (Kirillov et al., 2023), enabling the feature extractor (i.e., DINOv2) to focus on semantically relevant visual content while suppressing background interference. Recent work has demonstrated similar benefits of integrating SAM with object detection models for fine- grained spatial localization tasks (e.g., Han et al., 2023; Ren et al., 2024; Wang et al., 2024), suggesting that pixel-level segmentation can improve feature quality for downstream classification or regression tasks. Consistent with those earlier findings, the GDâSAMâDv2 pipeline here yielded the most substantial segmentation-related improvements for V estimation relative to the GDâDv2 pipeline, with errors ~8.5 cm smaller at the start and ~4.7 cm smaller at the end of a lift. The 17 greater benefit for V estimation may reflect the sensitivity of vertical distance measures to precise localization of the hands and feet, consistent with principles of perspective projection, in which small vertical displacements in image space can correspond to substantial differences in physical height when projected into real-world coordinates (Hartley & Zisserman, 2003). The advantage of pixel-level segmentation was most evident under single-view camera conditions, in which the GDâSAMâDv2 pipeline yielded substantially larger error reductions vs. multi-view conditions. For instance, at the start of a lift, the addition of segmentation yielded V estimation MAE values ~8â15 cm smaller for single-view conditions, but ~4â8 cm smaller for multi-view conditions. This pattern suggests that segmentation improves object localization within individual frames, which is particularly beneficial when viewpoint diversity is limited (Bae et al., 2022; Pavllo et al., 2018). In comparison, under multi-view conditions the added benefit of segmentation is smaller, but still meaningful, with MAE reductions of ~1â2 cm when estimating H and ~4â8 cm when estimating V. 4.2 Effect of Camera View Condition on H and V Estimation Performance Camera view conditions varied with H and V estimation performance, with multi-view conditions generally yielding smaller errors vs. single-view conditions. At the start of a lift, the V1+V2+V3 condition produced MAE values of ~6.2 cm when estimating H and ~7.8 cm when estimating V, which were significantly smaller vs. the worst-performing single-view conditions, with MAE = ~10.7 cm for the V1 condition when estimating H and MAE = ~27.7 cm for the V2 condition when estimating V (Figure 4). This performance advantage of the V1+V2+V3 condition was maintained at the end of a lift, with MAE = ~8.1 cm when estimating H and MAE = ~4.8 cm when estimating V, which were significantly smaller vs. the worst-performing single- view conditions, with MAE = ~14.4 cm for the V2 condition when estimating H and MAE = ~15.5 cm for the V1 condition when estimating V (Figure 5). These results suggest that multi- view capture provides geometric redundancy that reduces localization uncertainty caused by occlusion and perspective effects in monocular video (Hartley & Zisserman, 2003; Remondino & ElâHakim, 2006; Szeliski, 2022). Similar benefits of using multiple synchronized views have been reported in multi-view human pose estimation, in which incorporating information from more than one camera yields more accurate spatial localization of body landmarks than single- view approaches, particularly in conditions involving occlusion or unfavorable viewing angles (Qiu et al., 2019). This pattern is consistent with prior evidence that camera placement meaningfully affects vision-based ergonomic assessment outputs, with viewpoint-dependent visibility and depth ambiguity contributing to variability in estimated RNLE-related quantities (Murugan et al., 2025; Neupane et al., 2024; Wang et al., 2019; Zahabi et al., 2026; Zhang et al., 2022). In the present dataset, V1 and V2 corresponded to two oblique views of the workspace from opposite sides, whereas V3 provided a frontal view of the participant. When synchronized views with differing perspectives were used, multi-view conditions showed more consistent localization of task-relevant landmarks used to estimate RNLE distances, including the hands, shoes, and handled load, across frames. This redundancy was particularly relevant for V estimation, which exhibited greater variability under single-view conditions and showed larger error reductions under multi-view conditions. Notably, prior validation work indicates that RNLE distance-related inputs are less accurate under single-view capture and benefit from frontal or sagittal camera perspectives, supporting the inclusion of a frontal view (V3) in multi- 18 view conditions (Neupane et al., 2024; Wang et al., 2019; Zahabi et al., 2026; Zhang et al., 2022). Interestingly, certain two-camera conditions, particularly V1+V3 and V2+V3, produced errors comparable to the V1+V2+V3 condition when estimating H at the start of a lift, most notably for MAE and RMSE. When estimating V, these two-camera conditions likewise showed comparable performance to V1+V2+V3 for MAE at the start of a lift, although differences became more pronounced at the end of a lift. Notably, both V1+V3 and V2+V3 combine a frontal view with an oblique view, suggesting that these pairing can capture much of the task-relevant geometric information required for RNLE distance estimation, particularly during the initial phase of a lift. In these conditions, the frontal view supports localization of vertical displacement, while the oblique view preserves information related to anteriorâposterior and mediolateral positioning. However, the V1+V2+V3 condition consistently yielded the smallest errors across metrics and lift phases, indicating more stable performance even when differences in mean error were relatively small. These findings suggest that multi-view camera setups can provide more stable H and V estimation, particularly for V, although the magnitude of improvement depends on the specific combination of views. This stability matters, because errors in these inputs can propagate to risk outputs; for example, there is recent evidence of systematic bias in RNLE metrics derived using a commercial tool (e.g., RWL overestimation and LI underestimation on average), implying that reducing view-dependent uncertainty in H and V can affect risk classification rather than merely improving kinematic fidelity (Zahabi et al., 2026). 4.3 Temporal Differences in H and V Estimation Performance Between Start and End of Lift Estimation errors for H and V exhibited opposite temporal patterns between the start and end of a lift. Specifically for the GDâSAMâDv2 pipeline, V estimation errors decreased substantially from the start to the end of a lift, with respective MAE of ~14.5 vs. ~7.5 cm, whereas H errors increased, with respective values of ~7.2 vs. ~9.3 cm (see Figure A.9 for MAE comparison between start and end of a lift). This divergence persisted across both pipelines and most camera view conditions, indicating that the underlying error sources for H and V might vary across the lifting phases. At the start of a lift, participants adopted a posture involving torso flexion, with their hands near the floor, and the hands and lower body segments were more likely to be partially occluded by the torso and handled load, particularly in the oblique single-view conditions (i.e., V1 and V2) and in multi-view conditions that relied primarily on oblique perspectives (i.e., V1+V2). These occlusions can reduce the consistency of ROI localization and degrade frame-level estimation, particularly for V. At the end of a lift, participants were more upright with their hands near hip height, which might reduce occlusion and improve visibility of the hands, especially when a frontal perspective was available (i.e., V3, V1+V3, V2+V3, and V1+V2+V3), supporting smaller V estimation errors. In contrast, the increase in H estimation error from the start to the end of a lift may reflect greater difficulty localizing ankle-related landmarks at the end of a lift. As the load is raised, it may partially occlude the lower body or reduce the visibility of shoe-related ROIs used to localize the ankles, which would disproportionately affect H estimation. This ankle-visibility limitation is most relevant for single-view conditionsâparticularly V1 and V2âand was mitigated, but not eliminated, when synchronized views include the frontal perspective (i.e., V1+V3, V2+V3, and V1+V2+V3). Overall, these findings indicate that H and V estimation performance is not 19 uniform across the start and end of a lift and is sensitive to posture-dependent occlusion patterns. This pattern suggests that a single regression model may not optimally capture both lift phases. 4.4 Practical Applications for VLM-Based Ergonomic Assessment The present findings have several implications for future applications of VLM-based systems in occupational ergonomic assessment. First, the demonstrated feasibility of estimating RNLE distance parameters from RGB videoâwithout requiring manual measurement, wearable sensors, or marker-based motion capture systemsâsupports the use of VLM-based approaches to enable video-based ergonomic risk assessments. Second, multi-view camera conditions consistently yielded smaller estimation errors than single-view conditions, indicating that practical implementations should prioritize the use of multiple synchronized views when feasible. While single-view conditions reduced hardware and installation cost, they resulted in substantially larger estimation errorsâparticularly for Vâwhich may limit their suitability for applications requiring reliable distance estimates. Also, two-camera conditions (e.g., V1+V3) yielded estimation performance comparable to three-camera configurations in several cases, indicating a favorable trade-off between estimation performance and system complexity that may be appropriate in resource-constrained environments. Third, the consistent advantage of pixel-level segmentation across camera view conditions suggests that segmentation-augmented pipelines should be prioritized in future implementations when estimation performance is a primary concern. Although the detection-only pipeline may reduce computational complexity, the segmentation-based approach yielded substantial reductions in estimation error (approximately 22â39% across metrics), which may justify the additional computational cost in many ergonomic assessment applications. As segmentation models continue to improve in efficiency (e.g., Ravi et al., 2024), this trade-off is likely to become increasingly favorable. 4.5 Limitations and Future Research Directions Several limitations of our study should be acknowledged. Although the estimation performance was evaluated across multiple participants of both biological sexes with varied anthropometries and across a range of lifting conditions (e.g., lift origin, hand configuration, and box mass), all tasks were performed by relatively young (18â39 years old), healthy participants in controlled laboratory settings with uniform backgrounds and consistent lighting. Thus, caution should be taken in generalizing the findings to other populations and to real-world occupational settings. In practice, work environments often involve variable lighting, cluttered backgrounds, dynamic occlusions from coworkers or equipment, and more diverse lifting postures, all of which may affect model performance. Although we relied on zero-shot VLMs for detection and segmentation of ROIs and feature extraction (i.e., Grounding DINO, SAM, and DINOv2) without task-specific fine-tuningâsuggesting some degree of generalizabilityâfuture work should explicitly evaluate model performance in real environments to better assess external validity. In addition, ground-truth distance labels were derived from wearable IMU sensors, which introduce their own measurement uncertainty and may not perfectly align with the anatomical landmarks specified by the RNLE. While IMU-based motion capture has been validated for kinematic measurements (Schall et al., 2018), small errors related to sensor placement relative to true joint centers may propagate into the reference labels. Future studies could incorporate marker-based optical motion capture or direct physical measurements to provide independent validation of VLM-based estimates. 20 Further, while our results demonstrate the feasibility of estimating H and V, the RNLE includes additional parametersâsuch as asymmetry angle, coupling quality, and task frequencyâthat were not addressed here. Future work could examine whether VLM-based pipelines can be extended to estimate these additional RNLE parameters directly from video, enabling ambient RNLE-based risk assessment. Finally, differences in estimation performance between the start and end phases of lifts suggest that a single regression model may not optimally capture the full range of lifting postures. Future work could explore phase-aware modeling strategies or architectures that explicitly account for posture changes across the lifting tasks. 5.0 Conclusions Efficiently and accurately estimating RNLE distance parameters (i.e., H and V) during manual lifting is essential for assessing and controlling the risk of WMSDs. While a range of approaches have been used for ergonomic assessment, including observation-based methods, wearable sensors, marker-based motion capture systems, and pose-based vision methods, these approaches are often time-consuming, intrusive, or difficult to scale in real-world work environments. We developed and evaluated VLM-based pipelines to non-invasively estimate H and V from RGB video streams, including a text-guided, detection-only pipeline and a detection-plus- segmentation pipeline, each of which was assessed across seven camera view conditions using leave-one-subject-out validation. Estimation performance varied across pipelines and camera view conditions, with the segmentation-based, multi-view pipeline consistently yielding the lowest errors, achieving mean absolute errors of approximately 6â8 cm when estimating H and 5â8 cm when estimating V. Overall, these findings demonstrate the feasibility of using VLM- based pipelines to provide video-based estimates of RNLE distance parameters, supporting progress toward ergonomic risk assessment without wearable sensors or manual measurements. However, further research is needed to evaluate pipeline performance across diverse worker populations and in real occupational settings. 6.0 Declaration of generative AI and AI-assisted technologies in the manuscript preparation process During the preparation of this work, we used ChatGPT to refine some sentences, improve the clarity of the text, and generate a schematic illustration (i.e., Figure 2). After using this tool/service, we reviewed and edited the content as needed, and we take full responsibility for the content of this publication. 7.0 Acknowledgements The first author was supported by a predoctoral training program grant (T03 OH008613) from CDC/NIOSH. The current contents are solely the authorsâ responsibility and do not necessarily represent the official views of CPWR, NIOSH, or the CDC. The authors thank Advanced Research Computing at Virginia Tech for providing computational resources and technical support that contributed to the results reported in this paper. URL: https://arc.vt.edu/ References Arnab, A., Doersch, C., & Zisserman, A. (2019). Exploiting temporal context for 3D human pose estimation in the wild. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3395â3404. 21 Bae, G., Budvytis, I., & Cipolla, R. (2022). Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View Geometry. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2832â2841. https://doi.org/10.1109/CVPR52688.2022.00286 Bazarevsky, V., Grishchenko, I., Raveendran, K., Zhu, T., Zhang, F., & Grundmann, M. (2020). BlazePose: On-device Real-time Body Pose tracking (Version 1). arXiv. https://doi.org/10.48550/ARXIV.2006.10204 Bezzini, R., Crosato, L., Teppati Losè, M., Avizzano, C. A., Bergamasco, M., & Filippeschi, A. (2023). Closed-Chain Inverse Dynamics for the Biomechanical Analysis of Manual Material Handling Tasks through a Deep Learning Assisted Wearable Sensor Network. Sensors, 23(13), 5885. https://doi.org/10.3390/s23135885 Cao, Z., Simon, T., Wei, S.-E., & Sheikh, Y. (2016). Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields (Version 2). arXiv. https://doi.org/10.48550/ARXIV.1611.08050 Chen, K. (2019). Sitting Posture Recognition Based on OpenPose. IOP Conference Series: Materials Science and Engineering, 677(3), 032057. https://doi.org/10.1088/1757- 899X/677/3/032057 Chen, W., Mees, O., Kumar, A., & Levine, S. (2024). Vision-Language Models Provide Promptable Representations for Reinforcement Learning (arXiv:2402.02651). arXiv. https://doi.org/10.48550/arXiv.2402.02651 Cheng, Y., Yang, B., Wang, B., Yan, W., & Tan, R. T. (2019). Occlusion-aware networks for 3d human pose estimation in video. Proceedings of the IEEE/CVF International Conference on Computer Vision, 723â732. Da Costa, B. R., & Vieira, E. R. (2010). Risk factors for workârelated musculoskeletal disorders: A systematic review of recent longitudinal studies. American Journal of Industrial Medicine, 53(3), 285â323. https://doi.org/10.1002/ajim.20750 David, G. C. (2005). Ergonomic methods for assessing exposure to risk factors for work-related musculoskeletal disorders. Occupational Medicine, 55(3), 190â199. https://doi.org/10.1093/occmed/kqi082 Dempsey, P. G., Burdorf, A., Fathallah, F. A., Sorock, G. S., & Hashemi, L. (2001). Influence of measurement accuracy on the application of the 1991 NIOSH equation. Applied Ergonomics, 32(1), 91â99. https://doi.org/10.1016/S0003-6870(00)00026-0 Dempsey, P. G., McGorry, R. W., & Maynard, W. S. (2005). A survey of tools and methods used by certified professional ergonomists. Applied Ergonomics, 36(4), 489â503. https://doi.org/10.1016/j.apergo.2005.01.007 Deng, J., Dong, W., Socher, R., Li, L.-J., Kai Li, & Li Fei-Fei. (2009). ImageNet: A large-scale hierarchical image database. 2009 IEEE Conference on Computer Vision and Pattern Recognition, 248â255. https://doi.org/10.1109/CVPR.2009.5206848 Dosovitskiy, A. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv Preprint arXiv:2010.11929. 22 Fan, C., Mei, Q., & Li, X. (2024). Assisting in the identification of ergonomic risks for workers: A large vision-language model approach. ISARC. Proceedings of the International Symposium on Automation and Robotics in Construction, 41, 1010â1017. SciTech Premium Collection (3092410938). Fan, C., Mei, Q., Wang, X., & Li, X. (2024). ErgoChat: A Visual Query System for the Ergonomic Risk Assessment of Construction Workers (arXiv:2412.19954). arXiv. https://doi.org/10.48550/arXiv.2412.19954 Forgione, C., Coruzzolo, A. M., Lolli, F., Balugani, E., & Gamberini, R. (2025). Leveraging OpenPose and Kinect: Cutting-edge technologies for Ergonomic Risk Assessment. IFAC- PapersOnLine, 59(10), 1313â1318. https://doi.org/10.1016/j.ifacol.2025.09.221 Fox, R. R., Lu, M.-L., Occhipinti, E., & Jaeger, M. (2019). Understanding outcome metrics of the revised NIOSH lifting equation. Applied Ergonomics, 81, 102897. https://doi.org/10.1016/j.apergo.2019.102897 Gholamiangonabadi, D., Kiselov, N., & Grolinger, K. (2020). Deep Neural Networks for Human Activity Recognition With Wearable Sensors: Leave-One-Subject-Out Cross-Validation for Model Selection. IEEE Access, 8, 133982â133994. https://doi.org/10.1109/ACCESS.2020.3010715 Govaerts, R., Tassignon, B., Ghillebert, J., Serrien, B., De Bock, S., Ampe, T., El Makrini, I., Vanderborght, B., Meeusen, R., & De Pauw, K. (2021). Prevalence and incidence of work- related musculoskeletal disorders in secondary industries of 21st century Europe: A systematic review and meta-analysis. BMC Musculoskeletal Disorders, 22(1), 751. https://doi.org/10.1186/s12891-021-04615-9 Guran, N. B., Ren, H., Deng, J., & Xie, X. (2024). Task-oriented Robotic Manipulation with Vision Language Models (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2410.15863 Gutierrez, M., Gomez, B., Retamal, G., PeĂąa, G., Germany, E., Ortega-Bastidas, P., & Aqueveque, P. (2024). Comparing Optical and Custom IoT Inertial Motion Capture Systems for Manual Material Handling Risk Assessment Using the NIOSH Lifting Index. Technologies, 12(10), 180. https://doi.org/10.3390/technologies12100180 Han, X., Wei, L., Yu, X., Dou, Z., He, X., Wang, K., Sun, Y., Han, Z., & Tian, Q. (2023). Boosting segment anything model towards open-vocabulary learning. arXiv Preprint arXiv:2312.03628. Hartley, R., & Zisserman, A. (2003). Multiple view geometry in computer vision. Cambridge university press. Hlucny, S. D., & Novak, D. (2020). Characterizing Human Box-Lifting Behavior Using Wearable Inertial Motion Sensors. Sensors, 20(8), 2323. https://doi.org/10.3390/s20082323 Hsu, G.-S. J., Wu, J. S., Huang, Y.-K. D., Chiu, C.-C., & Kang, J.-H. (2025). Automatic Detect Incorrect Lifting Posture with the Pose Estimation Model. Life, 15(3), 358. https://doi.org/10.3390/life15030358 23 Jung, S., Su, B., Wang, H., Lu, L., Xie, Z., Xu, X., & Fitts, E. P. (2022). A computer vision-based lifting task recognition method. Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 66(1), 1210â1214. https://doi.org/10.1177/1071181322661507 JĂźngel, M., Mellmann, H., & Spranger, M. (2008). Improving Vision-Based Distance Measurements Using Reference Objects. In U. Visser, F. Ribeiro, T. Ohashi, & F. Dellaert (Eds.), RoboCup 2007: Robot Soccer World Cup XI (Vol. 5001, p. 89â100). Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-540-68847-1_8 Kang, D., Jeong, D., Lee, H., Park, S., Park, H., Kwon, S., Kim, Y., & Paik, J. (2025). VLM- HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis. In A. Del Bue, C. Canton, J. Pont-Tuset, & T. Tommasi (Eds.), Computer Vision â ECCV 2024 Workshops (Vol. 15634, p. 218â235). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031- 92591-7_14 Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., DollĂĄr, P., & Girshick, R. (2023). Segment Anything (arXiv:2304.02643). arXiv. https://doi.org/10.48550/arXiv.2304.02643 Liao, Y.-H., Mahmood, R., Fidler, S., & Acuna, D. (2024). Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models (arXiv:2409.09788). arXiv. https://doi.org/10.48550/arXiv.2409.09788 Liberty Mutual Insurance. (2023). 2023 Workplace Safety Index. 2023 Workplace Safety Index: The Top 10 Causes of Disabling Injuries. https://business.libertymutual.com/insights/2023- workplace-safety-index/ Liu, P.-L., & Chang, C.-C. (2022). Simple method integrating OpenPose and RGB-D camera for identifying 3D body landmark locations in various postures. International Journal of Industrial Ergonomics, 91, 103354. Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, Jie, Jiang, Q., Li, C., Yang, Jianwei, Su, H., Zhu, J., & Zhang, L. (2024). Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection (arXiv:2303.05499). arXiv. https://doi.org/10.48550/arXiv.2303.05499 Loshchilov, I., & Hutter, F. (2019). Decoupled Weight Decay Regularization (arXiv:1711.05101). arXiv. https://doi.org/10.48550/arXiv.1711.05101 Lou, Z., Zhan, Z., Xu, H., Li, Y., Hu, Y. H., Lu, M.-L., Werren, D. M., & Radwin, R. G. (2025). A Single-Camera Method for Estimating Lift Asymmetry Angles Using Deep Learning Computer Vision Algorithms. IEEE Transactions on Human-Machine Systems, 55(2), 309â314. https://doi.org/10.1109/THMS.2025.3539187 Lu, M.-L., Barim, M. S., Feng, S., Hughes, G., Hayden, M., & Werren, D. (2020). Development of a Wearable IMU System for Automatically Assessing Lifting Risk Factors. In V. G. Duffy (Ed.), Digital Human Modeling and Applications in Health, Safety, Ergonomics and Risk Management. Posture, Motion and Health (Vol. 12198, p. 194â213). Springer International Publishing. https://doi.org/10.1007/978-3-030-49904-4_15 24 Lu, M.-L., Putz-Anderson, V., Garg, A., & Davis, K. G. (2016). Evaluation of the Impact of the Revised National Institute for Occupational Safety and Health Lifting Equation. Human Factors: The Journal of the Human Factors and Ergonomics Society, 58(5), 667â682. https://doi.org/10.1177/0018720815623894 Mehta, D., Sridhar, S., Sotnychenko, O., Rhodin, H., Shafiei, M., Seidel, H.-P., Xu, W., Casas, D., & Theobalt, C. (2017). VNect: Real-time 3D human pose estimation with a single RGB camera. ACM Transactions on Graphics, 36(4), 1â14. https://doi.org/10.1145/3072959.3073596 Mudiyanselage, S. E., Nguyen, P. H. D., Rajabi, M. S., & Akhavian, R. (2021). Automated Workersâ Ergonomic Risk Assessment in Manual Material Handling Using sEMG Wearable Sensors and Machine Learning. Electronics, 10(20), 2558. https://doi.org/10.3390/electronics10202558 Murugan, A. S., Noh, G., Jung, H., Kim, E., Kim, K., You, H., & Boufama, B. (2025). Optimising computer vision-based ergonomic assessments: Sensitivity to camera position and monocular 3D pose model. Ergonomics, 68(1), 120â137. https://doi.org/10.1080/00140139.2024.2304578 Neupane, R. B., Li, K., & Boka, T. F. (2024). A survey on deep 3D human pose estimation. Artificial Intelligence Review, 58(1), 24. Ojelade, A., Rajabi, M. S., Kim, S., & Nussbaum, M. A. (2025). A Data-Driven Approach to Classifying Manual Material Handling Tasks Using Markerless Motion Capture and Recurrent Neural Networks. International Journal of Industrial Ergonomics. https://doi.org/10.1016/j.ergon.2025.103755 Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.- Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., ... Bojanowski, P. (2024). DINOv2: Learning Robust Visual Features without Supervision (arXiv:2304.07193). arXiv. https://doi.org/10.48550/arXiv.2304.07193 Parsa, B., Samani, E. U., Hendrix, R., Devine, C., Singh, S. M., Devasia, S., & Banerjee, A. G. (2019). Toward ergonomic risk prediction via segmentation of indoor object manipulation actions using spatiotemporal convolutional networks. IEEE Robotics and Automation Letters, 4(4), 3153â3160. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., & Antiga, L. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32. Paudel, P., Kwon, Y.-J., Kim, D.-H., & Choi, K.-H. (2022). Industrial Ergonomics Risk Analysis Based on 3D-Human Pose Estimation. Electronics, 11(20), 3403. https://doi.org/10.3390/electronics11203403 Pavllo, D., Feichtenhofer, C., Grangier, D., & Auli, M. (2018). 3D human pose estimation in video with temporal convolutions and semi-supervised training (Version 2). arXiv. https://doi.org/10.48550/ARXIV.1811.11742 25 Pires, J. D. D., Oliveira, M. C., Fonseca, B., Ribeiro, M., Giglio, B. A. F., & Calzado, A. X. (2025). Ergonomic Assessment Using Human Pose Estimation: A Real-Time Approach with YOLO and BlazePose. SimpĂłsio Brasileiro de Computação Aplicada Ă SaĂşde (SBCAS), 293â 298. Prechelt, L. (1998). Early StoppingâBut When? In G. B. Orr & K.-R. MĂźller (Eds.), Neural Networks: Tricks of the Trade (p. 55â69). Springer. https://doi.org/10.1007/3-540-49430-8_3 Punnett, L., & Wegman, D. H. (2004). Work-related musculoskeletal disorders: The epidemiologic evidence and the debate. Journal of Electromyography and Kinesiology, 14(1), 13â23. https://doi.org/10.1016/j.jelekin.2003.09.015 Qiu, H., Wang, C., Wang, J., Wang, N., & Zeng, W. (2019). Cross view fusion for 3d human pose estimation. Proceedings of the IEEE/CVF International Conference on Computer Vision, 4342â 4351. Rajabi, M. S., Ojelade, A., Kim, S., & Nussbaum, M. A. (2025). Vision-Language Models for Occupational Physical Exposure Assessment: Classification and Temporal Segmentation of Manual Material Handling Tasks. Available at SSRN 5667590. Ranavolo, A., Ajoudani, A., Chini, G., Lorenzini, M., & Varrecchia, T. (2024). Adaptive Lifting Index (aLI) for Real-Time Instrumental Biomechanical Risk Assessment: Concepts, Mathematics, and First Experimental Results. Sensors, 24(5), 1474. https://doi.org/10.3390/s24051474 Ranavolo, A., Varrecchia, T., Rinaldi, M., Silvetti, A., Serrao, M., Conforto, S., & Draicchio, F. (2017). Mechanical lifting energy consumption in work activities designed by means of the ârevised NIOSH lifting equation". INDUSTRIAL HEALTH, 55(5), 444â454. https://doi.org/10.2486/indhealth.2017-0075 Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., & Gustafson, L. (2024). Sam 2: Segment anything in images and videos. arXiv Preprint arXiv:2408.00714. Remondino, F., & ElâHakim, S. (2006). Imageâbased 3D Modelling: A Review. The Photogrammetric Record, 21(115), 269â291. https://doi.org/10.1111/j.1477-9730.2006.00383.x Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., & Yan, F. (2024). Grounded sam: Assembling open-world models for diverse visual tasks. arXiv Preprint arXiv:2401.14159. Sabino, I., Fernandes, M. D. C., Cepeda, C., Quaresma, C., Gamboa, H., Nunes, I. L., & Gabriel, A. T. (2024). Application of wearable technology for the ergonomic risk assessment of healthcare professionals: A systematic literature review. International Journal of Industrial Ergonomics, 100, 103570. https://doi.org/10.1016/j.ergon.2024.103570 Schall, M. C., Sesek, R. F., & Cavuoto, L. A. (2018). Barriers to the Adoption of Wearable Sensors in the Workplace: A Survey of Occupational Safety and Health Professionals. Human Factors: The Journal of the Human Factors and Ergonomics Society, 60(3), 351â362. https://doi.org/10.1177/0018720817753907 26 Senior, A., Heigold, G., Ranzato, M., & Yang, K. (2013). An empirical study of learning rates in deep neural networks for speech recognition. 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, 6724â6728. https://doi.org/10.1109/ICASSP.2013.6638963 Szeliski, R. (2022). Computer Vision: Algorithms and Applications. Springer International Publishing. https://doi.org/10.1007/978-3-030-34372-9 U.S. Bureau of Labor Statistics. (2024). Survey of Occupational Injuries and Illnesses Data. Nonfatal Occupational Injuries and Illnesses Requiring Days Away from Work. https://w.bls.gov/data/home.htm Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ĺ., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30. Veges, M., & Lorincz, A. (2020). Temporal Smoothing for 3D Human Pose Estimation and Localization for Occluded People (arXiv:2011.00250). arXiv. https://doi.org/10.48550/arXiv.2011.00250 Wang, J., Zhu, M., Li, Y., Li, H., Yang, L., & Woo, W. L. (2024). Detect2Interact: Localizing Object Key Field in Visual Question Answering with LLMs. IEEE Intelligent Systems, 39(3), 35â44. Wang, X., Hu, Y. H., Lu, M.-L., & Radwin, R. G. (2019). The accuracy of a 2D video-based lifting monitor. Ergonomics, 62(8), 1043â1054. https://doi.org/10.1080/00140139.2019.1618500 Waters, T. R., Baron, S. L., & Kemmlert, K. (1998). Accuracy of measurements for the revised NIOSH lifting equation. Applied Ergonomics, 29(6), 433â438. https://doi.org/10.1016/S0003- 6870(98)00015-5 Waters, T. R., Putz-Anderson, V., Garg, A., & Fine, L. J. (1993). Revised NIOSH equation for the design and evaluation of manual lifting tasks. Ergonomics, 36(7), 749â776. https://doi.org/10.1080/00140139308967940 Wen, D., Peng, K., Yang, K., Chen, Y., Liu, R., Zheng, J., Roitberg, A., Paudel, D. P., Van Gool, L., & Stiefelhagen, R. (2025). RoHOI: Robustness Benchmark for Human-Object Interaction Detection (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2507.09111 Yong, G., Liu, M., & Lee, S. (2024). Automated Captioning for Ergonomic Problem and Solution Identification in Construction Using a Vision-Language Model and Caption Augmentation. Construction Research Congress 2024, 709â718. https://doi.org/10.1061/9780784485293.071 Zahabi, S. J. N., Kim, S., Nussbaum, M. A., Porto, R., & Lim, S. (2026). Evaluating the Accuracy and Feasibility of a Commercial AI-Powered Ergonomic Assessment system for Automotive Assembly Work. https://doi.org/10.2139/ssrn.6045455 Zhang, S., Wang, C., Dong, W., & Fan, B. (2022). A survey on depth ambiguity of 3D human pose estimation. Applied Sciences, 12(20), 10591. 27 Zhang, Z., Hu, F., Lee, J., Shi, F., Kordjamshidi, P., Chai, J., & Ma, Z. (2024). Do Vision- Language Models Represent Space And How? Evaluating Spatial Frame Of Refer- Ence Under Ambiguities. Zhao, W., Yang, G., Zhang, R., Jiang, C., Yang, C., Yan, Y., Hussain, A., & Huang, K. (2023). Open-Pose 3D Zero-Shot Learning: Benchmark and Challenges (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2312.07039 1 Appendix A.1 Participants A convenience sample of 32 young adults (19 males and 13 females) completed the study and was recruited from the university and local community. Respective means (SD) of age, body mass, and height were 26.8 (4.5) years, 77.1 (12.2) kg, and 176.0 (5.7) cm for the males; and 27.0 (5.6) years, 68.5 (8.4) kg, and 169.3 (6.7) cm for the females. All participants self-reported being right-handed, physically active (i.e., exercising at least twice per week), and having no musculoskeletal disorders within the past year. The research reported herein compiled with the tenets of the Declaration of Helsinki, and the study protocol was approved by the Institutional Review Board at Virginia Tech (#23-095). Informed consent was obtained from all participants prior to any data collection (Ojelade et al., 2025; Ojelade, 2024). A.2 Symmetric Lifting Task Details A single wood box (width = 26.0 cm; depth = 41.0 cm; and height = 23.5 cm; Figure A.1) was used for all symmetric lifting trials. Figure A.1. Illustration of the two hand configurations: broad (left), and narrow (right; adapted from Ojelade et al. (2025) A.3 Experimental Procedures A repeated-measures experimental design was implemented. Participants completed a training phase to practice the tasks using their comfortable work strategies and speed, simulating an industrial setting. During the experimental phase, the order of Hand Configurations and Box Mass presentation was counterbalanced using balanced Latin square designs to reduce potential bias. To minimize physical fatigue, a mandatory rest period of at least four minutes was provided between trials involving different hand configurations (Ojelade et al., 2025). A.4 Azure Kinect⢠Camera Instrumentation Whole-body motion was recorded using three synchronized Azure Kinect⢠markerless camera systems (Microsoft Corporation, Seattle, WA, USA), sampled at 30 Hz. The cameras were positioned approximately 1.74 m from the edge of the work area; this configuration was established during pilot testing to improve coverage given the narrow field of view of the camera units. The three cameras were time-synchronized using a 3.5-m auxiliary cable connected in a daisy-chain configuration, with one camera designated as the primary device and the other two as secondary devices (Ojelade et al., 2025). 2 A.5 Wearable IMU Data Processing Whole-body kinematics were captured from a wearable inertial motion capture system (Noraxon Ultium, Noraxon, Scottsdale, AZ, USA) and were exported using Noraxon myoResearch. To minimize temporal offsets among motion capture systems, the IMU system was synchronized with other data sources using an analog output signal generated from LabView. Whole-body kinematics were exported in Biovision Hierarchy (BVH) format from Noraxon myoResearch. Exported kinematics were low-pass filtered (6 Hz cutoff; 4th-order Butterworth; bidirectional), and kinematics data were downsampled to 30 Hz to match the sampling rate of the Azure Kinect⢠cameras (Ojelade et al., 2025). A.6 Revised NIOSH Lifting Equation Definition The Revised NIOSH Lifting Equation (RNLE) is a biomechanical risk assessment model to evaluate the physical demands of two-handed manual lifting tasks and to estimate the associated risk of low-back musculoskeletal injury (Waters et al., 1993). The RNLE provides a recommended weight limit (RWL) for a given lifting task based on task geometry, load characteristics, and temporal factors, and it forms the basis for computing the Lifting Index (LI), a commonly used indicator of lifting risk. The RNLE expresses the recommended weight limit as: í ííż=íżíśĂíťíĂííĂíˇíĂí´íĂíšíĂíśí where íżíś is the load constant and the remaining multipliers account for task-specific characteristics, including horizontal location of the hands (HM), vertical location of the hands (VM), vertical travel distance (DM), asymmetry (AM), lift frequency (FM), and coupling quality (CM). Among these factors, horizontal distance (H) and vertical distance (V) are fundamental geometric inputs that directly affect several multipliers in the equation and estimated biomechanical loading. In the RNLE, H and V describe the spatial relationship between the workerâs hands, body, and the floor at the start and end of each lift (Fox et al., 2019; Waters et al., 1993). A.7 Evaluation Metrics Details H and V estimation performance was evaluated separately for the start and end of the lift by comparing predicted (íŚĚ í ) and ground-truth (íŚ í ) values. For each lift trial, error metrics were computed across all start and end of the lift frames (indexed by í = 1,..., í) for each participant fold, camera-view condition, and VLM-based pipeline. íííí í´íí ííí˘íĄí í¸ííí (íí´í¸)= 1 í ââŁíŚĚ í âíŚ í ⣠í í=1 í ííĄ íííí ííí˘ííí í¸ííí (í ííí¸)= â 1 í â ( íŚĚ í âíŚ í ) 2 í í=1 íííĽííí˘í í´íí ííí˘íĄí í¸ííí (íííĽí´í¸)=íííĽ íâ1,...,í âŁíŚĚ í âíŚ í ⣠3 Figure A.2: Example visualization of the ROI detection and segmentation process for a representative lifting frame of Camera View 2 (V2). Top left: Original RGB video frame. Top right: Detection of the primary lifter using Grounding DINO, with the bounding box corresponding to the highest-confidence person detection. Bottom left: Detection of task-relevant ROIs within a cropped region centered on the lifter using Grounding DINO. Bottom right: Pixel-level segmentation of the detected ROIs produced by the Grounding DINO + SAM pipeline. 4 Figure A.3: Example visualization of the ROI detection and segmentation process for a representative lifting frame of Camera View 3 (V3). Top left: Original RGB video frame. Top right: Detection of the primary lifter using Grounding DINO, with the bounding box corresponding to the highest-confidence person detection. Bottom left: Detection of task-relevant ROIs within a cropped region centered on the lifter using Grounding DINO. Bottom right: Pixel-level segmentation of the detected ROIs produced by the Grounding DINO + SAM pipeline. 5 Table A.1. ANOVA results assessing the effect of Pipeline, Camera View Condition, and Sex on estimation errors for horizontal and vertical distances at the start of the lift. Entries are F values (p values), and significant effects are highlighted in bold font. Effect MAE â Horizontal MAE â Vertical RMSE â Horizontal RMSE â Vertical MaxAE â Horizontal MaxAE â Vertical Pipeline (P) 40.21 (<.0001) 81.94 (<.0001) 39.19 (<.0001) 79.21 (<.0001) 25.37 (<.0001) 34.08 (<.0001) Camera View Condition (C) 15.38 (<.0001) 38.05 (<.0001) 14.89 (<.0001) 35.39 (<.0001) 5.84 (<.0001) 8.77 (<.0001) Sex (S) 0.10 (0.7566) 2.62 (0.1161) 0.10 (0.7563) 1.90 (0.1786) 0.00 (0.9911) 0.00 (0.9495) S Ă P 0.54 (0.4646) 0.12 (0.7253) 0.34 (0.5587) 0.20 (0.6518) 0.58 (0.4485) 1.43 (0.2319) S Ă C 1.10 (0.3600) 0.68 (0.6690) 1.08 (0.3727) 0.69 (0.6547) 1.21 (0.2993) 0.96 (0.4507) P Ă C 2.43 (0.0257) 4.62 (0.0002) 1.89 (0.0821) 4.65 (0.0001) 1.30 (0.2560) 2.85 (0.0100) S Ă P Ă C 1.72 (0.1146) 1.99 (0.0664) 1.69 (0.1231) 2.14 (0.0483) 1.60 (0.1464) 1.93 (0.0744) 6 Table A.2. ANOVA results assessing the effect of Pipeline, Camera View Condition, and Sex on estimation errors for horizontal and vertical distances at the end of the lift. Entries are F values (p values), and significant effects are highlighted in bold font. Effect MAE â Horizontal MAE â Vertical RMSE â Horizontal RMSE â Vertical MaxAE â Horizontal MaxAE â Vertical Pipeline (P) 100.79 (<.0001) 59.45 (<.0001) 65.56 (<.0001) 58.69 (<.0001) 25.18 (<.0001) 27.47 (<.0001) Camera View Condition (C) 20.94 (<.0001) 29.92 (<.0001) 18.24 (<.0001) 27.84 (<.0001) 1.45 (0.1952) 11.84 (<.0001) Sex (S) 0.29 (0.5968) 2.20 (0.1483) 0.03 (0.8734) 1.83 (0.1857) 0.12 (0.7346) 0.29 (0.5973) S Ă P 0.45 (0.5049) 1.45 (0.2288) 0.32 (0.5722) 1.98 (0.1605) 0.90 (0.3437) 5.15 (0.0238) S Ă C 0.96 (0.4493) 0.70 (0.6528) 0.71 (0.6410) 0.63 (0.7039) 0.67 (0.6764) 0.76 (0.6012) P Ă C 4.83 (<.0001) 4.21 (0.0004) 4.13 (0.0005) 4.14 (0.0005) 1.74 (0.1104) 2.42 (0.0260) S Ă P Ă C 1.86 (0.0858) 2.09 (0.0531) 1.78 (0.1024) 2.34 (0.0309) 1.31 (0.2534) 2.73 (0.0131) 7 Figure A.4: Significant interaction effect of Pipeline Ă Camera View Condition on RMSE of V estimation at the start of the lift. Figure A.5: Significant interaction effect of Pipeline Ă Camera View Condition on MaxAE of V estimation at the start of the lift. 8 Figure A.6: Significant interaction effect of Pipeline Ă Camera View Condition on RMSE of H estimation at the end of the lift. Figure A.7: Significant interaction effect of Pipeline Ă Camera View Condition on RMSE of V estimation at the end of the lift. 9 Figure A.8: Significant interaction effect of Pipeline Ă Camera View Condition on MaxAE of V estimation at the end of the lift. 10 Figure A.9: Comparison of MAE for H (top) and V (bottom) estimation between the start and end of a lift for the GDâSAMâDv2 pipeline.