Source-linked AI summary

MotionQ: Operator-Conditioned Motion Quotients for Cross-Observation WiFi Gesture Recognition

Xiang Zhang, Huan Yan, Geying Yang, Jianchun Liu, Tao Liu, Zhi Liu, Meng Li

arXiv:2609.11818v1cs.HC

TL;DR

WiFi gesture recognition must generalize when geometry changes the physical observation operator and changes which motion cues are observable. MotionQ learns geometry-conditioned two-support motion quotients and enforces task sufficiency through single-link-retention interventions instead of representation matching. It achieves strong generalization across unseen layouts, extrapolative orientations, environments, and device configurations.

  • Problem

    Changing orientations, links, and transceiver layouts alters the wireless observation operator, while invariant representations may discard task-relevant cues not jointly observable across operators.

  • Method

    MotionQ generates an operator-conditioned two-support motion measure, removes support-order ambiguity with a motion quotient, and uses single-link-retention interventions to enforce task-level equivalence.

  • Results

    MotionQ substantially outperforms existing WiFi sensing methods under unseen layouts, extrapolative orientations, and environments, averaging 89.2% across six extrapolation tasks.

  • Takeaways & Limitations

    Preserving operator-conditioned task information is more appropriate than enforcing universal representation invariance for cross-observation WiFi gesture recognition.

  • Takeaways & Limitations

    MotionQ targets short-duration limb gestures whose dynamics can be compactly represented by two coordinated motion components, rather than multiple weakly coupled body parts moving along several directions.

Abstract

from arXiv · show

WiFi gesture recognition is accurate in fixed deployments but often degrades when user orientation, available links, or transceiver placement changes. Unlike ordinary domain shifts, these changes alter the wireless observation operator, so the same motion is expected to produce different measurements. Existing methods nevertheless pursue domain-invariant features and largely overlook changing layouts and observation configurations. Yet changing the observation operator also changes which task-relevant motion cues are physically observable, rather than merely altering the appearance of a fixed set of cues. Under a local linearization of the WiFi forward process, we derive a common task-observability condition under which a strict common linear representation is recoverable from every geometry-induced operator while preserving the gesture task. When the condition fails, enforcing stronger alignment across additional heterogeneous source operators may discard task-relevant cues still observable under individual operators. We therefore present MotionQ, which generates an operator-conditioned two-support motion measure for each candidate geometry. A motion quotient removes only the arbitrary ordering of its unlabeled supports and is represented by permutation-invariant central moments. Rather than matching quotients across operators, single-link-retention interventions encourage each view to retain information sufficient for gesture recognition. Extensive evaluations show that MotionQ is robust to extrapolative observation operators.

1 Introduction

MotionQ treats changing orientations, links, and transceiver layouts as changes in the wireless observation operator, questioning whether strict domain invariance can preserve gesture-relevant information. It derives a task-observability condition, introduces operator-conditioned motion representations, and reports strong generalization under unseen layouts and orientations.

  • 1 Introduction: The introduction frames cross-observation recognition as a distinct generalization problem because different operators can expose different task-relevant motion cues.Existing domain adaptation and domain generalization approaches primarily pursue invariant features, while cross-layout sensing receives comparatively limited attention.
  • 1 Introduction: It derives a common task-observability condition describing when task-sufficient invariant representations remain feasible across WiFi operators.When the condition fails, alignment cannot recover information missing from an operator.
  • 1 Introduction: MotionQ conditions motion measures on candidate geometry, removes arbitrary support identities through a motion quotient, and enforces task equivalence through single-link-retention interventions rather than representation matching.The pipeline processes complete and single-link observations and requires them to predict the same gesture label.
  • 1 Introduction: MotionQ substantially outperforms WiFi sensing methods under unseen layouts, extrapolative orientations, and environments, averaging 89.2% across six extrapolation tasks.It exceeds WiGRUNT, UniFi, and GesFi by 20.1, 11.0, and 16.9 percentage points, respectively.
  • 1 Introduction: The paper formulates changes in orientation, available links, and transceiver layouts as changes in the wireless observation operator.These changes alter how motion is physically observed rather than merely changing measurement appearance.

2 Preliminary

WiFi links observe geometry-dependent projections of gesture motion, so changing layouts can alter which task-relevant cues remain observable. The analysis formalizes when strict common representations are recoverable and motivates operator-conditioned, task-sufficient representations.

  • 2.1 WiFi Gesture Sensing as an Observation Process: A WiFi link observes a geometry-dependent projection of motion, so changing user orientation, links, or transceiver placement changes the resulting measurement.The bistatic observation relation shows that the same physical gesture can produce different wireless observations under different geometries.
  • 2.3 Local Common Task Observability under WiFi Operators: The local observation model maps gesture state through an operator-specific Jacobian, while a task variable specifies which variations must remain available for recognition.The common-representation analysis separates operator-dependent observability from task-relevant information.
  • 2.3 Local Common Task Observability under WiFi Operators: Strict common linear recovery is possible exactly when a representation is recoverable from every operator while remaining sufficient for the gesture task.The theorem characterizes this condition under the local linear WiFi observation model.
  • 2.3 Local Common Task Observability under WiFi Operators: Source-domain invariance does not guarantee target observability, because an unseen geometry may make source-common task cues unobservable.The result is explicitly local and concerns exact linear common recoverability, not arbitrary nonlinear or distribution-level generalization.
  • 2.3 Local Common Task Observability under WiFi Operators: Adding heterogeneous operators can shrink the strictly common task-relevant subspace, whereas simultaneous complementary links can preserve or enlarge observable gesture information.This contrast explains why stronger cross-operator alignment may cause negative transfer while additional simultaneous observations can add information.
  • 2.3 Local Common Task Observability under WiFi Operators: MotionQ therefore conditions representations on the observation operator and constrains them through task sufficiency rather than numerical equality across operators.This principle is presented as the direct motivation for MotionQ.

3 System Overview

MotionQ uses geometry-conditioned operators to generate two-support motion measures, then applies single-link-retention training and permutation-invariant moments across observation views.

  • 3 System Overview: MotionQ extracts per-link evidence, uses latent geometry hypotheses for bistatic conditioning, and generates a two-support motion measure for each hypothesis.The system overview describes geometry-conditioned link fusion and motion-measure generation.
  • 3 System Overview: The same pipeline processes complete and single-link observations, using a smooth worst-suboperator objective to retain gesture-sufficient information without matching representations.The intervention compares full and individual-link views through task sufficiency rather than representation equality.
  • 3 System Overview: Permutation-invariant central moments represent the motion quotient while removing only the arbitrary ordering of its unlabeled supports.This final stage preserves the quotient representation while enforcing permutation invariance.

4 Method

MotionQ conditions motion representations on latent observation-operator hypotheses and geometry, then enforces task sufficiency without requiring representation invariance across observations. It uses compact, permutation-invariant two-support motion measures and single-link interventions to improve robustness to changing wireless operators.

  • Per-link evidence and fusion: Amplitude–phase and Doppler-spectrum observations are encoded per link, fused across active links, and decoded into operator-specific motion states.Shared encoders prevent dependence on receiver identity or link count, while the active-link set receives layout-relative coordinate encoding.
  • Bistatic operator conditioning: MotionQ maintains latent geometry hypotheses and analytically conditions link evidence on bistatic observation operators before generating motion measures.Hypotheses share a candidate position and vary in orientation; known transceiver coordinates determine geometry-dependent projections for each link.
  • Minimal two-support motion measure: MotionQ represents each candidate geometry with a minimal two-support velocity measure that captures aggregate motion, directional dispersion, and component imbalance.Two supports provide the first non-zero internal moments beyond the mean without introducing an unconstrained high-dimensional field.
  • Task-equivalent training: Single-link-retention interventions require complete and single-link observations to remain sufficient for the same gesture task without directly aligning their motion quotients.This task-level constraint permits operator-conditioned representations to differ while promoting robustness to broader observation-operator changes.
  • Motion quotient: A motion quotient removes arbitrary ordering of the two unlabeled supports, allowing equivalent parameterizations to share the same representation.The supports have no predefined identities or correspondence to particular body parts.

5 Evaluations

MotionQ is evaluated across changing layouts, orientations, environments, and link availability, with results showing strong cross-observation generalization and robustness to extrapolative operators.

  • Overall Cross-Observation Performance: 90.618% average accuracy on primary protocols W1–W6 exceeds CORAL, DANN, WiGRUNT, UniFi, and GesFi by 1.818, 2.368, 10.077, 3.896, and 5.923 percentage points.Table 1 reports the highest average accuracy for MotionQ across protocols covering multi-factor observation changes and sufficient-coverage settings.
  • Overall Cross-Observation Performance: MotionQ achieves 89.943%, 91.794%, and 93.742% under three-to-two, six-to-two, and six-to-three link reductions, respectively, ranking first in all three settings.These evaluations exhaustively cover 35 two- and three-link configurations, which were not explicitly enumerated during single-retained-link training.
  • Overall Cross-Observation Performance: Across all fourteen task-level metrics, MotionQ averages 88.044%, exceeding the strongest external baseline, DANN, by 2.461 percentage points.On the PerceptAlign stress test, MotionQ reaches 77.397%, within 0.696 percentage points of CORAL while exceeding the other listed baselines.
  • Source Coverage and Extrapolative Orientations: 89.2% average accuracy across six endpoint-extrapolation configurations outperforms WiGRUNT, UniFi, and GesFi by 20.1, 11.0, and 16.9 percentage points, respectively.The advantage over the strongest WiFi-specific baseline is 8.5 percentage points on extrapolative orientations versus 1.4 points on covered orientations.
  • Source Coverage and Extrapolative Orientations: Replacing bracketing source orientations with equally sized but distant orientations reduces MotionQ accuracy by 23.0% and 13.3% in W5 and W6, while adding a distant source can improve it by 3.5%.The results indicate that source diversity helps when it improves target-side operator coverage, whereas invariant alignment may suppress operator-specific cues.

6 Discussion

MotionQ’s discussion defines the scope of its two-support motion model, its physical-observability boundary, and its deployment assumptions. The authors identify extensions for richer activities, improved target-side coverage, and more general sensing geometries.

  • Scope of motion modeling: MotionQ’s two-support prior targets short limb gestures but cannot fully model multiple weakly coupled body parts moving across several directions.The authors therefore do not claim general full-body activity recognition; MotionQ remains competitive but does not outperform CORAL on average in the PerceptAlign stress test.
  • Scope of motion modeling: The authors propose adaptive or hierarchical motion quotients that add supports only when observations require them.This is intended to extend the approach toward complex activities and cross-layout human pose estimation.
  • Physical observability and interpretation: MotionQ cannot recover motion cues physically absent from the target operator, so generalization still degrades when source operators poorly cover the target.The learned supports represent effective latent velocity components rather than identified body parts or metrically recovered trajectories.
  • Deployment assumptions: MotionQ assumes known transceiver coordinates and a two-dimensional bistatic geometry, while moving transceivers and three-dimensional motion remain future work.The stated calibration assumption is practical for common indoor deployments with stationary transceivers.

7 Conclusion

The conclusion frames cross-observation WiFi recognition as a problem of changing physical observation operators rather than changing measurement appearance alone. It presents MotionQ as an operator-conditioned, task-level approach and reports strong generalization to unseen layouts and extrapolative orientations.

  • Changing transceiver geometry alters the physical observation operator in cross-observation WiFi gesture recognition.
  • MotionQ learns geometry-conditioned motion quotients while enforcing equivalence only at the recognition-task level.
  • Experiments demonstrate strong generalization to unseen layouts and extrapolative orientations.

A Detailed Evaluation Protocols

The evaluation protocols define source–target separation, multi-factor Widar3.0 shifts, and independently evaluated PerceptAlign configurations. They also specify how users, scenes, layouts, and activity labels are assigned across protocols.

  • Protocol setup: Target observations are excluded from training, model selection, and hyperparameter tuning, with each configuration repeated using three fixed random seeds.Multi-target protocols report equal-weight macro-averages across independently evaluated targets.
  • Widar3.0 protocols: Widar3.0 protocols vary user orientation, environment, user population, receiver layout, and target orientation across disjoint source and target user sets.For W1–W3, sources come from Scene 1 and testing uses Scenes 2 and 3 with different users.
  • PerceptAlign protocols: PerceptAlign uses Scene 2 as source and Scene 3 as target, shares user identities across collections, and therefore is not strict cross-user generalization.Scene 3 includes three receiver configurations, while four base activities are retained.
  • Label definition: Recognition uses base activity labels rather than direction-specific execution variants because related directions are treated as global rotations under the cross-observation formulation.

B Model Details

The model details describe active-layout coordinate encoding, hypothesis-specific geometry prediction, bistatic conditioning, and the detailed evaluation protocol tables. These components provide geometric inputs and parameterized operator representations for MotionQ.

  • Active-layout coordinate encoding: For an active link set, the model computes receiver centroid, layout center, layout scale, and a transmitter-to-receiver-centroid frame axis.
  • Active-layout coordinate encoding: The model encodes receiver positions relative to the layout and applies a fixed numerical coordinate scale.
  • Geometry prediction: For each geometry hypothesis, the operator network predicts shared and anchor-specific angular residuals plus a local position, with bounded residuals and zero-initialized output heads.Initialization makes hypothesis orientations coincide with their anchors and the initial position equal the active-link centroid.
  • Multiplicative bistatic conditioning: The model uses dimensionless bistatic coordinates to construct a directional basis for geometry-conditioned processing.
  • Evaluation protocols: The appendix includes detailed Widar3.0 leave-one-orientation, leave-one-position, and leave-one-environment protocols, alongside PerceptAlign stress-test protocols.

C Signal Preprocessing and Backbone Details

The preprocessing pipeline constructs amplitude–phase and Doppler-spectrum inputs, then encodes them separately at native resolutions before feature fusion.

  • CSI-ratio preprocessing: CSI-ratio preprocessing suppresses shared phase distortions, after which amplitude and phase are stacked across valid subcarriers.The resulting representation is the amplitude–phase input.
  • Doppler preprocessing: The pipeline derives a Doppler frequency spectrum with 121 bins and resamples each gesture to 224 time steps.With 30 subcarriers, the AP and DFS inputs are 60 × 224 and 121 × 224, respectively.
  • Signal encoders: Separate ImageNet-initialized ResNet-18 encoders process the AP and DFS branches at native resolutions, preserving subcarrier and Doppler-bin organization.The branch features are then concatenated and projected.

D Relation to Gesture Execution

The gesture visualizations are qualitative execution sketches, while the paper emphasizes temporal and morphological correspondence in effective motion components rather than metrically recovered hand paths.

  • Gesture execution: The six Widar3.0 gestures are performed around the body and are not confined to the two-dimensional bistatic model’s world-coordinate xy plane.The visualized paths are cumulative latent velocities expressed in a hypothesis-conditioned effective motion frame.
  • Interpretation of visualizations: Absolute orientation, scale, and shape of the effective motion frame need not match the execution sketches.The sketches therefore provide qualitative rather than metrically recovered hand positions.
  • Gesture-specific patterns: Temporal and morphological patterns distinguish gestures: reversals characterize Push–Pull, sustained turning Draw-O, alternating sharp turns Draw-Zigzag, smooth excursions Sweep and Slide, and direct convergence Clap.These characteristics are described as consistent with recurrent patterns in the corresponding figure.
  • Support correspondence: The two supports remain exchangeable effective motion components, making their correspondence temporal and morphological rather than tied to a fixed ordering.This preserves the support symmetry emphasized in the execution analysis.
Loading 2609.11818v1…