Source-linked AI summary
Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video
Mohammad Sadra Rajabi, Aanuoluwapo Ojelade, Sunwook Kim, Maury A. Nussbaum
TL;DR
The paper addresses the need for continuous, non-invasive hand-force measurements in occupational material handling by evaluating a VLM-based RGB-video pipeline. The approach produced feasible triaxial, bilateral force estimates, with object-aware ROIs and multi-camera capture offering selective benefits.
Problem
Continuous external hand-force measurement for occupational exposure assessment typically requires instrumented objects or specialized sensing.
Method
A VLM-based pipeline combined text-guided ROI localization, pretrained vision-transformer features, known box mass, and temporal regression, evaluated with leave-one-subject-out validation.
Results
Overall RMSE was ~4.7–5.6 N for horizontal and mediolateral forces and ~10.6–11.0 N for vertical force, while object ROIs generally improved estimation and multi-camera capture most clearly benefited peak-force estimation.
Takeaways & Limitations
RGB video with known box mass can provide detailed, continuous external hand-force estimates with substantially reduced instrumentation compared with conventional approaches.
Takeaways & Limitations
Generalization beyond young healthy participants and controlled laboratory settings, including unseen viewpoints, different objects, and workplace environments, was not evaluated.
Abstract
from arXiv · showhide
External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sensing. We evaluated a vision-language model (VLM)-based pipeline that combines task-specific textual cues, visual representations, and known box mass to estimate dynamic, triaxial, bilateral external hand forces from RGB video. Thirty-five healthy young adults performed five MMH tasks involving lifting, carrying, pushing, and pulling with box masses of 6, 9, and 12 kg. The pipeline used text-guided localization of participant and handled-object regions of interest (ROIs), pretrained vision-transformer feature extraction, and transformer-based temporal regression. Performance was evaluated using leave-one-subject-out validation across seven camera-view conditions (three single-view and four multi-view conditions) and four ROI strategies. Overall, root mean square error was ~4.7-5.6 N for the horizontal and mediolateral force components and ~10.6-11.0 N for the vertical component. Including the handled object as a second ROI generally improved force estimation, with some of the largest benefits under single-camera conditions, whereas pixel-level segmentation provided little additional improvement. Multi-camera capture provided the clearest benefit for peak-force estimation, particularly for the vertical component, whereas differences in overall frame-level error among camera configurations were comparatively modest. These findings demonstrate the feasibility of estimating continuous, bilateral, directional hand-force estimates from RGB video and known load mass without requiring sensors on the worker or handled objects as model inputs, supporting the development of more scalable occupational physical exposure and risk assessments.
2.0 Methods
The study developed a VLM-based pipeline that detects task-relevant worker and object regions, extracts visual features, and predicts bilateral hand forces from RGB video. It evaluated four ROI strategies across five MMH tasks and varied experimental conditions using instrumented-box force measurements as targets.
- Dataset and experimental conditions: 35 healthy young adults performed five MMH tasks under varied masses, hand configurations, lift origins, and repeated trials.The tasks represented lifting, carrying, pushing, and pulling activities.
- Dataset and experimental conditions: Dynamic force measurements from bilateral triaxial load cells served as the ground-truth targets for model development and evaluation.
- VLM-based pipeline: The pipeline detected and segmented task-relevant participant and handled-object ROIs, extracted visual features, and used temporal modeling to estimate frame-level bilateral external hand forces.The main stages were ROI processing, feature extraction, and transformer-based prediction.
- VLM-based pipeline: Task-specific textual prompts guided GroundingDINO detection of the participant and handled object before ROI processing.
- VLM-based pipeline: Four ROI strategies varied participant-only versus participant-plus-object regions and bounding boxes versus SAM-refined segmentation masks.The strategies were illustrated across the five MMH tasks using View 1 frames.
2.4 Feature Extraction from ROIs
ROI visual features were converted into temporally ordered inputs and paired with synchronized six-channel load-cell targets. A transformer model incorporated ROI, mass, camera-view, and temporal information and was validated with leave-one-subject-out cross-validation.
- Feature extraction: DINOv2 produced frozen 768-dimensional visual feature vectors for each detected ROI from every RGB video frame.Single-ROI strategies produced [T, 1, 768] tensors, while dual-ROI strategies produced [T, 2, 768] tensors.
- Feature extraction: Synchronized feature–force pairs used six bilateral targets: 𝐹𝑥, 𝐹𝑦, and 𝐹𝑧 for each hand.
- Temporal force prediction: A transformer encoder modeled temporal dependencies across consecutive frames to estimate dynamic bilateral external hand forces.The approach used temporal context rather than treating frames independently.
- Temporal force prediction: Dual-ROI inputs concatenated participant and handled-object features before projection into a 256-dimensional frame embedding.Mass and camera-view embeddings were added to frame tokens as contextual inputs.
- Training and evaluation: 300-frame sequences, attention masking, early stopping, and per-channel normalization supported model training and evaluation.Leave-one-subject-out validation trained across 35 folds, holding out one participant per fold.
- Training and evaluation: RMSE, MAE, and Peak Error measured force estimation performance, with RMSE emphasized for statistical results.Peak Error compared 95th-percentile absolute actual and predicted force magnitudes.
3.0 Results
Results examined camera-view and ROI-strategy effects on force estimation across force directions and hands. Object-inclusive ROIs generally helped, while camera-view differences in frame-level error were often modest and task-dependent.
- Overall evaluation: Continuous actual and predicted bilateral forces were examined across horizontal, mediolateral, and vertical directions for all four ROI strategies.
- Overall evaluation: RMSE and Peak Error were emphasized for camera-view and ROI-strategy effects, with task and mass included as context.
- Mass and task effects: RMSE increased consistently from 6 to 9 to 12 kg across all force directions and both hands.At 6 kg versus 12 kg, RMSE was roughly 26–36% smaller across force channels and hands.
- 𝐹𝑥 effects: Camera-view differences in 𝐹𝑥 RMSE were generally small in Tasks 1–4 and larger at ~10–13% in Task 5, without a consistent camera advantage.
- 𝐹𝑥 effects: Dual-ROI strategies S2 and S4 generally produced smaller 𝐹𝑥 RMSE than single-ROI strategies S1 and S3 in Tasks 4 and 5.Differences among ROI strategies were small in Tasks 1–3.
- 𝐹𝑥 effects: 𝐹𝑥 Peak Error differences among camera conditions were generally small in Tasks 1–3 and larger in Tasks 4–5, with no consistently superior condition.
- 𝐹𝑥 effects: For right-hand 𝐹𝑥 Peak Error, S2 and S4 were ~2.7 N versus ~2.9 N for S1 and S3.
3.3 𝑭𝒚: Effects of Camera View Condition and ROI Strategy
Camera view and ROI strategy affected force-estimation errors, with camera effects varying by task and object-inclusive dual-ROI strategies generally improving vertical-force RMSE.
- ~4.6–4.8 N right-hand 𝐹𝑦 RMSE occurred across camera views, while left-hand values were ~5.6–5.7 N without a significant camera effect.
- ~2.3–2.5 N versus ~2.7 N right-hand 𝐹𝑦 PE occurred under multi-camera versus single-camera conditions, with smaller differences for the left hand.
- ~3–11% differences in 𝐹𝑧 RMSE across camera views varied by task, with no camera condition consistently advantageous and larger differences in Tasks 2 and 3.
- ~5–7% lower 𝐹𝑧 RMSE was generally achieved by dual-ROI strategies S2 and S4 versus single-ROI strategies S1 and S3, especially in Tasks 1 and 2.
- ~4.6–4.9 N versus ~1.9–2.3 N 𝐹𝑧 PE occurred under single-camera versus multi-camera conditions, most prominently in Tasks 2 and 4.
- ~20–26% smaller right-hand 𝐹𝑧 PE occurred with S2 and S4 versus S1 and S3 under single-camera conditions, while multi-camera differences were small and inconsistent.
4.0 Discussion
The discussion interprets the VLM pipeline’s force estimates relative to instrumented approaches and examines effects of mass, task, camera view, and ROI representation. Overall errors were comparable to more instrumented methods, while peak-force estimation benefited most from multiple views and object-inclusive ROIs.
- ~4.7–11 N per-hand RMSE was broadly comparable to estimates from approaches using substantially more instrumentation.
- ~26–36% smaller RMSE occurred at 6 kg versus 12 kg across force channels and both hands, while PE was also smaller at the lower mass.
- Task-dependent errors affected all force channels, with relatively larger 𝐹𝑥 errors during pushing and pulling matching findings from markerless motion capture with in-shoe pressure data.
- ~4.6–4.9 N versus ~1.9–2.3 N 𝐹𝑧 PE occurred under single-camera versus multi-camera conditions in Tasks 2 and 4, whereas RMSE differences were modest.
- ~5–7% lower force RMSE resulted from including the handled object ROI, especially for 𝐹𝑥 in pushing and pulling and 𝐹𝑧 in lifting.
- ~20–26% smaller 𝐹𝑧 PE from object-inclusive ROIs appeared under single-camera conditions, suggesting object information may partly compensate for limited viewpoint information.
- Pixel-level segmentation added little beyond detection-only ROIs, with near-identical right-hand 𝐹𝑥 PE and 𝐹𝑦 RMSE across representations.
4.5 Practical Implications for VLM-Based Ergonomic Assessment
The VLM-based pipeline estimated detailed bilateral hand forces from RGB video and known box mass, but its practical use depends on task, force quantity, and system configuration. Object ROIs and multi-camera capture offered selective benefits, while broader validation remains necessary.
- Including the handled object as a second ROI generally improved estimation, particularly for pushing and pulling.The authors identify object localization as a potentially useful default when feasible.
- The study used young, healthy participants and controlled laboratory conditions, so generalization to diverse workers, objects, viewpoints, and workplaces remains unevaluated.
- RMSE was ~4.7–5.6 N for horizontal and mediolateral forces and ~10.6–11.0 N for vertical force components.
- Pixel-level segmentation provided little additional benefit beyond selecting informative regions of interest.
- Multi-camera capture most clearly improved peak-force estimation, while overall frame-level RMSE differences among camera conditions were modest.
A.1 Participants
The study used a convenience sample of 35 young adults recruited from a university and local community. Participants were right-handed and physically active.
- 35 young adults—21 males and 14 females—participated in the study.The sample was recruited from the university and local community.
- Participants self-reported being right-handed and physically active.
A.2 MMH Task Details and Experimental Procedures
Participants performed five simulated manual material handling tasks using one wood box in a repeated-measures design. Hand configurations and box-mass presentations were counterbalanced.
- One wood box measuring 26.0 cm wide, 41.0 cm deep, and 23.5 cm high was used across five simulated MMH tasks.
- The experiment used a repeated-measures design with training followed by experimental trials.Participants practiced using comfortable work strategies and speed to simulate an industrial setting.
- The order of hand configurations and box-mass presentations was counterbalanced using a balanced Latin square design.
- Hand configurations included broad and narrow arrangements.Figure A.1 illustrates the broad configuration on the left and narrow configuration on the right.
A.3 Azure Kinect™ Camera Instrumentation
Whole-body kinematics were recorded with three synchronized Azure Kinect markerless camera systems. Cameras operated at 30 Hz and were positioned to improve coverage of the work area.
- Three synchronized Azure Kinect markerless camera systems recorded whole-body kinematics at 30 Hz.
- The cameras were positioned approximately 1.74 m from the edge of the work area.
- Camera synchronization used a 3.5-mm auxiliary cable connected in a daisy-chain configuration.
A.4 Load Cell Instrumentation and Hand Force Measurements
External hand forces were measured with bilateral triaxial load cells attached near the box top and expressed in a box-centered coordinate system.
- Triaxial load-cell signals were sampled at 200 Hz, filtered with a 300-ms moving root-mean-square window, and downsampled to 30 Hz.
- The six force targets comprised Fx, Fy, and Fz for each hand, represented in the local box-centered coordinate system.The coordinate system defines +X as horizontal/anterior–posterior, +Y as lateral/medial–lateral, and +Z as vertical.
- Load-cell measurements served as the ground-truth targets for estimating external hand forces during manual material handling tasks.
A.5 Transformer-Based Force Prediction Model: Additional Details
The force-prediction pipeline combines ROI processing, pretrained visual features, temporal modeling, and masked windowed evaluation. Supplementary analyses compare ROI strategies, camera conditions, tasks, masses, and force channels using mixed-effects results and figure-based comparisons.
- A modified transformer encoder produced simultaneous per-frame estimates of six bilateral force channels from temporally ordered visual features.The model used four encoder layers, eight attention heads, and a regression head mapping 256-dimensional features to six outputs.
- 300-frame windows used 150-frame training strides with overlap, whereas validation and testing used non-overlapping 300-frame windows with padding for shorter trials.Padded frames were excluded from loss computation and self-attention.
- RMSE and MAE were evaluated separately for six bilateral force channels using valid non-padded test frames, with RMSE and MAE log-transformed for model assumptions.The LMM analyses emphasized RMSE and PE for effects involving camera view condition or ROI strategy.
- Four ROI strategies compared participant-only or participant-plus-object bounding boxes and corresponding pixel-level segmentation masks.Strategies 1–2 used bounding boxes, while Strategies 3–4 refined those representations with SAM-generated masks.