Source-linked AI summary
Fast Generative Grasping via Lie Group-Constrained MeanFlow
S. Talha Bukhari, Yi Wei, Ruiqi Ni, Zachary Kingston, Aniket Bera
TL;DR
Robotic grasp generation must represent multimodal grasp distributions, while iterative diffusion and flow-based sampling is too slow for time-critical operation. The paper introduces Lie-group-constrained MeanFlow on SO(3) × R3, combining algebraic semigroup consistency with a flow-matching endpoint anchor. On ACRONYM and real-world noisy observations, it achieves reliable few-step generation with comparable grasp performance and millisecond-scale inference.
Problem
Iterative diffusion and flow-based grasp generators require many integration steps, limiting their use for reactive closed-loop robotics.
Method
GraspMF applies endpoint-parameterized MeanFlow on G = SO(3) × R3 with an algebraic semigroup objective anchored by Riemannian Conditional Flow Matching.
Results
GraspMF matches state-of-the-art grasp generation on ACRONYM, with few-step sampling achieving highest ID SR of 87.40%, highest OOD SR of 71.73%, and 15.5 ms latency.
Takeaways & Limitations
The formulation supports single- and few-step grasp synthesis, stable performance across sampling budgets, and real-world grasping under observation noise.
Abstract
from arXiv · showhide
Grasp synthesis is a core task in robotic manipulation, for which the solution typically forms a multimodal distribution rather than a point estimate. Generative robotic grasping aims to learn this distribution with deep generative models such as diffusion and flow-based approaches. The iterative nature of such generative models makes them flexible and generalizable; however, multi-step sampling impedes the time-critical operation required in robotics. We devise an approach to fast generative grasping based on MeanFlow on the product Lie group $\mathcal{G} = \mathrm{SO}(3) \times \mathbb{R}^3$. The training objective couples a purely algebraic semigroup consistency condition with Riemannian Conditional Flow Matching on $\mathcal{G}$ that anchors the average velocity to the data distribution. The resulting Lie Group-constrained MeanFlow formulation samples reliable grasps in $\leq 5$ network evaluations, matching the grasp generation performance of state-of-the-art diffusion and flow-based models on the ACRONYM dataset at millisecond-scale inference latency (up to $39\times$ speed-up). We further demonstrate that the approach directly translates to real-world robotic grasping without additional training or domain adaptation, exhibiting robust grasp synthesis under observation noise.
I. INTRODUCTION
Robotic grasping requires modeling multimodal grasp distributions, but diffusion and flow-based generators often need many integration steps. The paper applies MeanFlow on a Lie-group pose space to obtain fast grasp synthesis while retaining reported grasp quality and robustness.
- Motivation: Stable grasps form multimodal distributions in SE(3), motivating object-conditioned generative grasp models over point estimates.Generative models can sample multiple feasible grasps from point-cloud observations and compose with downstream objectives such as collision avoidance and reachability.
- Motivation: Tens to hundreds of integration steps limit diffusion and Flow Matching samplers for reactive, closed-loop robotic execution.Iterative refinement helps transport samples onto the grasp data manifold, including implicit contact constraints.
- Approach: MeanFlow predicts average velocity over a time interval using a teacher-free, simulation-free objective aligned with the probability path.This differs from consistency and shortcut models that optimize self-referential targets rather than the underlying transport vector field.
- Approach: GraspMF applies MeanFlow to G = SO(3)×R3 and couples algebraic semigroup consistency with a Conditional Flow Matching anchor.The formulation targets fast grasp pose generation on the product Lie group.
- Results: GraspMF matches state-of-the-art diffusion and flow-based methods on ACRONYM while operating at millisecond-scale latency and remaining stable across sampling budgets.The method is also demonstrated in real-world grasping with observation noise and efficient inference.
- Related Work: Prior generative grasping methods include score-based diffusion, equivariant Flow Matching, informative-source stochastic interpolants, and geometry-supervised approaches.These methods address multimodality, transformation consistency, sampling budgets, or imperfect and partial observations.
III. PRELIMINARIES
The paper represents rigid-body grasp poses on the product Lie group G = SO(3) × R3, using manifold, Lie-group, and Riemannian flow concepts to define conditional transport from a prior to grasp data.
- A. Riemannian Manifolds and Lie Groups: A manifold locally resembles Euclidean space, while a smooth manifold additionally has a C∞ differentiable structure.A Riemannian manifold equips tangent spaces with a smoothly varying inner product.
- A. Riemannian Manifolds and Lie Groups: A Lie group is both a differentiable manifold and a group whose multiplication and inversion are smooth.Left action identifies tangent spaces with the Lie algebra through the differential of translation.
- A. Riemannian Manifolds and Lie Groups: Figure 2 encodes rotations as cones on a sphere, with cone color representing target density and the red cone marking the identity rotation.The figure compares Flow Matching and MeanFlow on a closed-curve target distribution in SO(3).
- A. Riemannian Manifolds and Lie Groups: SE(3) is represented for modeling as the diffeomorphic product group G = SO(3) × R3, which admits a bi-invariant metric unlike the semidirect-product formulation.The product metric decouples rotational and translational pose components.
- B. Conditional Flow Matching: Riemannian Conditional Flow Matching uses geodesic interpolants between prior and target poses and trains a tangent vector field to match marginal velocity.At inference, the field is integrated on G using the exponential map to sample the target distribution.
IV. RIEMANNIAN MEANFLOW ON G
Riemannian MeanFlow is formulated on G = SO(3) × R3 by predicting clean grasp endpoints and deriving average velocities and flow maps in closed form.
- IV. Riemannian MeanFlow on G: The endpoint parameterization predicts the clean grasp directly, then obtains the average velocity and flow map from that prediction in closed form.This combines endpoint-regression stability with the two-time structure of flow maps.
- IV. Riemannian MeanFlow on G: MeanFlow defines a chord-form average velocity in the Lie algebra whenever the flow-map chord lies within the injectivity radius.The flow map pushes the marginal density from time s to time t.
- IV. Riemannian MeanFlow on G: MeanFlow trains a network to estimate average velocity so sampling evaluates the flow map directly instead of integrating the instantaneous velocity field.The average velocity extends to the time diagonal as the instantaneous left-trivialized velocity.
A. Endpoint Parameterization and the Induced Flow Map
The endpoint model induces geodesic flow-map jumps from a current pose toward a predicted clean grasp, while semigroup consistency makes direct and composed jumps agree.
- A. Endpoint Parameterization and the Induced Flow Map: The endpoint network Xθ maps a state and time interval endpoint to a predicted clean grasp endpoint.The prediction is interpreted as transport at the interval’s average velocity over the remaining horizon.
- A. Endpoint Parameterization and the Induced Flow Map: The induced flow map moves along the minimizing geodesic by fraction (t−s)/(1−s), splitting into rotational and translational components under the product metric.At t = 1, the fraction is unity and the map returns the single-step grasp prediction.
- A. Endpoint Parameterization and the Induced Flow Map: At t = s, the average velocity reduces to the instantaneous left-trivialized velocity and is anchored to marginal velocity by Riemannian Conditional Flow Matching.At the full horizon, the induced map equals the endpoint prediction.
- B. Semigroup Consistency: The semigroup property requires a direct flow-map jump from s to t to equal composition through any intermediate time r.The identity is exact for the flow map and is imposed on the learned map.
- B. Semigroup Consistency: The semigroup objective is algebraic, using forward network evaluations with expG and logG rather than covariant or network derivatives.The paper reports greater stability than differential identities whose targets contain derivative-of-network curvature-dependent variance.
- B. Semigroup Consistency: On the rotational factor, interval splitting is exact only up to Baker–Campbell–Hausdorff bracket terms, which vanish when the chord directions commute.The direct-product structure separately splits rotational and translational identities.
C. Training Objective
The training objective combines a data-anchored Riemannian Conditional Flow Matching loss with an algebraic semigroup consistency loss on the Lie group G. The anchor prevents degenerate self-consistent solutions, while the semigroup term trains interval-average velocities from composed flow-map states.
- Flow-Matching anchor: The network predicts the endpoint grasp Xθ, from which the average velocity and flow map are derived in closed form.The endpoint parameterization places the network output in G.
- Combined objective: The semigroup identity alone admits a zero-displacement predictor, so an explicit data-distribution anchor is required.With the anchor, the exact average velocity is a global minimizer and idealized optima reproduce the exact flow map under regularity conditions.
- Flow-Matching anchor: The Flow Matching anchor regresses the predicted endpoint against the data grasp through the conditional flow velocity on G.At the time diagonal, the model average velocity is matched to the marginal velocity induced by the data distribution.
- Semigroup loss: The semigroup loss compares a single-step average velocity with the chord velocity produced by a composed two-step flow map.The composed intermediate states are treated as stop-gradient self-consistency targets.
- Combined objective: The full objective is a weighted sum of the Flow Matching and semigroup terms, reducing to plain Riemannian CFM endpoint regression when λsemi = 0.Training combines both losses on each mini-batch with positive weights.
D. Inference
Inference iterates the learned flow map over a partition of [0, 1], with one network evaluation per step. Endpoint-prediction chord jumps differ from fixed-step Lie–Euler updates by targeting interval-average displacement rather than instantaneous-velocity truncation.
- Sampling procedure: A grasp is generated by drawing H0 from the source distribution and iterating Φθ over a partition with T sampling steps.Each update maps Htk to Htk+1 using the learned flow map.
- Sampling procedure: At T = 1, one network call returns the grasp directly as H1 = Xθ(H0, 0, 1).Each step costs one network evaluation plus a closed-form exponential/logarithm pair.
- Flow-map composition: Whenever the semigroup identity holds, composing T chord-form jumps returns the same result as the corresponding long-interval flow map.The method uses jumps toward the estimated clean grasp rather than Euler steps of a learned vector field.
- Numerical interpretation: Lie–Euler incurs O(∆t2) truncation error at fixed step size, whereas MeanFlow targets an interval-average velocity whose displacement reproduces the exact flow map.The remaining error is attributed to approximation of the two-time average-velocity field rather than numerical truncation.
E. Implementation Details
Implementation uses a lightweight SE3Dif-derived network conditioned on object point clouds and trained with stabilized, annealed objectives. Experiments report grasp metrics and inference cost across sampling budgets, including GraspMF at T = 5 and T = 1.
- Network: The network encodes 1024-point object clouds with a VNN encoder and predicts a 12-D pose representation containing rotation and translation outputs.Pose inputs use rigidly transformed keypoints, while the two time arguments use random Fourier features.
- Evaluation: Table I reports SR and EMD for in-domain and out-of-domain ACRONYM objects, including GraspMF at T = 5 and T = 1.Comparison methods use the sampling budgets reported for their original formulations.
- Data: The source distribution combines uniform Haar measure on SO(3) with an isotropic Gaussian on R3, while targets are expert-annotated valid grasps.Point clouds and translations are recentered, and object-grasp pairs receive random SO(3) rotations.
- Objective and optimization: λsemi is annealed from 0 to 1 over the first 1000 epochs while λcfm remains fixed at 1.The anchor dominates early training before self-consistency is enforced.
- Evaluation: Table II reports NFEs and average wall-clock latency for 100 grasps per object on a single NVIDIA RTX 5080 GPU.The reported methods and sampling budgets follow the rows of Table I.
- Inference setup: Inference measurements use 1024 input points and 100 sampled grasps per object after a 10-iteration warm-up.Grasps are generated by iterating the flow map over a uniform partition.
V. EXPERIMENTS & RESULTS
The experiments evaluate GraspMF against state-of-the-art multimodal grasp-generation methods in simulation and in a physical robotic grasping scenario. They include quantitative and qualitative results and ablation studies.
- Experimental scope: The evaluation covers simulation and a physical robotic grasping scenario.The study compares multimodal grasp-generation methods and examines quantitative, qualitative, and ablation results.
A. Experimental Setup
The experiments compare generated grasps qualitatively across representative object geometries, using the established evaluation protocol and IsaacGym simulation.
- A. Experimental Setup: The experimental setup follows prior deep generative grasping evaluation protocols.The paper presents its experimental details as an extension of prior work.
- A. Experimental Setup: Figure 3 qualitatively compares generated grasps for Laptop, Pencil, and Car in IsaacGym.Dark cyan denotes successful grasps and dark purple denotes failures; GraspMF is shown at T = 5 and T = 1.
1) Dataset:
The study trains and evaluates grasp-generation methods on ACRONYM objects, comparing GraspMF with diffusion-, flow-, and other fast-generation baselines under matched sampling budgets.
- 1) Dataset:: The ACRONYM subset contains 416 object instances and approximately 780 K valid grasps across ten shape categories.The categories include Book, Bottle, Bowl, Cap, CellPhone, Cup, Hammer, Mug, Scissors, and Shampoo.
- 1) Dataset:: The dataset split uses 90% of instances from each category for training and the remaining 10% for evaluation.
- 1) Dataset:: The comparison includes SE3Dif as a diffusion baseline, EGF as a flow-based baseline, BRIDGER as a fast-generation method, and VSIGD as a noisy-observation reference.
- 1) Dataset:: All methods are configured at the sampling budgets reported by their authors for noise-to-data transport.The effect of sampling budget is evaluated separately.
3) Performance Metrics:
Performance is assessed by simulated grasp success, distribution coverage, and inference latency across sampling budgets, with GraspMF maintaining strong performance at low-step settings.
- 3) Performance Metrics:: SR counts grasps that retain the object after a 5 s shake following simulated open-grip, approach, close-grip, and lift execution.The metric is computed from 100 sampled grasps per object instance.
- 3) Performance Metrics:: EMD compares 100 sampled grasps with 100 ground-truth samples using rotational geodesic and weighted translational pose distances.λ = 0.1 equates 1 m of translation with 5.7° of rotation, and EMD serves as a proxy for distribution coverage.
- 3) Performance Metrics:: Inference latency is the averaged wall-clock time in milliseconds to generate 100 grasps in parallel on one NVIDIA RTX 5080 GPU.
- 3) Performance Metrics:: 87.40% ID SR and 71.73% OOD SR are achieved by GraspMF at T = 5, with OOD EMD 0.4191, ID EMD 0.3702, and 15.5 ms latency.These results are reported as the strongest or competitive benchmark outcomes across the two protocols.
- 3) Performance Metrics:: 70.15% OOD SR with standard deviation 1.63 is maintained by GraspMF across sampling budgets T ∈ {1, 2, 5, 10, 20, 40, 70, 100}.By contrast, baseline SR drops sharply at low budgets, while GraspMF has the lowest latency at every budget.
- 3) Performance Metrics:: The ablations remove signed-distance regression, annealing, SVD projection, or the semigroup loss to evaluate the contribution of GraspMF design choices at T = 5.The ablation table defines each variant and reports the corresponding metrics.
E. Robustness to Partial Observations
GraspMF is evaluated under incomplete single-view geometry and in real-world tabletop grasping, where it maintains efficient few-step inference and robust grasp synthesis. Partial-observation performance is affected by view-dependent geometry, reference grasps, and normalization.
- Evaluation protocol: The partial-observation protocol uses single-view raycasting, subsampling to N = 1024 points, and reference grasps restricted to contacts with sampled geometry.Training and evaluation average results over 3 random views per object.
- Partial-observation evaluation: 68.46% OOD SR is achieved by GraspMF at T = 5 under single-view partial observations, exceeding BRIDGER’s 65.38%.Results are averaged over 3 random views per object, with 100 grasps generated per view.
- Partial-observation evaluation: 86.08% ID SR is achieved by GraspMF at T = 5, trailing BRIDGER’s 88.84% while sampling 7.8× faster.The reported latencies are 15.5 ms for GraspMF and 120.6 ms for BRIDGER.
- Protocol limitations: Absolute performance decreases because occluded geometry is ambiguous, view-dependent reference sets reduce supervision and increase target variance, and observed-cloud centering biases normalization.These effects are intrinsic to the single-view protocol.
- Real-world demonstration: GraspMF transfers to tabletop grasping with a wrist-mounted RGB-D camera and operates in the few-step regime while performing consistently across household objects.The robot executes collision-free plans and verifies stability by vertically lifting the object; each object is tested over 10 trials.