每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

arXiv:2608.197592026-08-19

不学任何物体,机械手也能学会怎么抓东西

多指机械手要抓取物体,通常需要用特定物体的数据集训练,但这样很难适应从未见过的新物体。GOAG换了个思路,只学习抓手自身的表面形状和关节运动方式,训练阶段完全不接触任何物体数据,只在推理阶段才输入物体的形状信息。在MultiDex基准测试上,GOAG平均成功率达到86.93%,与专门针对该数据集训练的顶尖方法效果相当,而且在生成大量抓取动作时速度明显更快。

他们做了什么

  1. 现有的深度学习抓取规划器依赖特定物体的数据集训练,难以泛化到新物体
  2. GOAG完全基于抓手自身的表面几何和关节运动生成训练数据,训练阶段从不接触任何物体的几何信息
  3. 团队采用改编自人类抓握分类法的6种抓取类型,在抓手表面随机采样接触区域,生成了300万条训练样本,数据生成仅耗时约1个GPU小时,而此前的方法需要1400个GPU小时
  4. 推理阶段,一个条件变分自编码器(CVAE)接收新物体的形状信息,预测物体表面上应接触的位置,再经过力闭合检验和优化步骤得到最终的抓取姿态
  5. GOAG在MultiDex数据集上平均成功率达到86.93%,与专门训练的顶尖方法表现相当;在五个不同的抓取数据集测试中,仅用一个未经过重新训练的Shadow Hand模型就取得了第二高的平均成功率,并且用真实机械臂成功抓取了11个YCB物体
Fig. 1: GOAG Paradigm. A successful grasp on an object induces dual contact zones on both object and gripper, at the intersection of the two geometries. Our method is built on this key observation: these contact zones (𝒞(.)) are closely the same from either perspective. GOAG capitalizes on this by training exclusively on gripper geometry, allowing it to learn a robust and generalizable grasping strategy without ever being exposed to a grasp database with specific objects geometries.
Fig. 1: GOAG Paradigm. A successful grasp on an object induces dual contact zones on both object and gripper, at the intersection of the two geometries. Our method is built on this key observation: these contact zones (𝒞(.)) are closely the same from either perspective. GOAG capitalizes on this by training exclusively on gripper geometry, allowing it to learn a robust and generalizable grasping strategy without ever being exposed to a grasp database with specific objects geometries.
TABLE I: Comparison of Dexterous Grasp Planning Methods.
Grasp RepresentationGraspRepresentationGripper PoseGripperPoseGripper Joint ValuesGripperJoint ValuesForce ClosureForceClosureNon- PenetrationNon-PenetrationTraining SetTrainingSetWorking Reference Frame
Grasp
Representation
Gripper
Pose
Gripper
Joint Values
Force
Closure
Non-
Penetration
Training
Set
Working
Reference Frame
Optional Grasp
Preference Interface
DFC [14]DirectOptimizedOptimizedNoObject
UniGrasp [24]IntermediateIK SolvedIK SolvedObjects + GrippersObject
GeoMatch [2]IntermediateIK SolvedIK SolvedObjects + GrippersObject
GenDexGrasp [12]IntermediateOptimizedOptimizedObjects + GrippersObject
ManiFM [37]IntermediateOptimizedOptimizedObjects + GrippersObjectContact Region
DRO-Grasp [32]IntermediateOptimizedOptimizedObjects + GrippersObjectPalm Orientation
DexDiffuser [33]DirectLearnedLearnedA posterioriA posterioriObjects + GrippersObject
DexGrasp Anything [39]DirectLearnedLearnedObjects + GrippersObject
GOAG (Ours)IntermediateSampledOptimizedGripper OnlyGripperPalm Full Pose
Fig. 2: Grasp Taxonomy Adaptation and Contact Sampling. (Top) We adapt the human grasp taxonomy from [7] to the Allegro Hand geometry. For each grasp type (e.g., C6, F27), we define a corresponding admissible contact region (black points), distinguishing it from the non-contact surface (blue points). (Bottom) Data generation mechanism: We randomly sample specific contact points (red) strictly within the admissible black regions. This allows the model to learn structured, feasible contact distributions based solely on gripper kinematics, independent of any object.
Fig. 2: Grasp Taxonomy Adaptation and Contact Sampling. (Top) We adapt the human grasp taxonomy from [7] to the Allegro Hand geometry. For each grasp type (e.g., C6, F27), we define a corresponding admissible contact region (black points), distinguishing it from the non-contact surface (blue points). (Bottom) Data generation mechanism: We randomly sample specific contact points (red) strictly within the admissible black regions. This allows the model to learn structured, feasible contact distributions based solely on gripper kinematics, independent of any object.
TABLE II: In-depth grasp performance analysis on the Multidex [12] test set. We report our grasp results, on three classic dexterous grippers in terms of success rate, efficiency and diversity.
MethodData DrivenObject-Agnostic TrainingSuccess Rate (%) ↑Efficiency (sec. / grasps) ↓Diversity (avg.) ↑
BarrettAllegroShadowHandAvg.BarrettAllegroShadowHandT (m)R (rad)Q (rad)
DFC [14]83.1082.7172.1579.32>1800>1800>18000.06071.4240.3579
GenDexGrasp [12] (full)70.2671.4871.1570.969.7816.4514.650.05191.4160.2567
DRO-Grasp [32] (pretrain, w/o controller)78.3075.8063.3072.470.880.421.720.05461.5150.2892
GOAG (w/o FC)86.3091.2074.7084.070.090.130.150.04801.3960.3162
GOAG87.4093.2077.9086.930.180.190.200.04791.4010.3170
Fig. 3: Overview of GOAG. Geometrical graspability is learned in an object-agnostic manner by focusing on the gripper’s capabilities. Training: We sample gripper configurations Q to generate ℋ⁡(Q) and corresponding contact points (𝒞⁡(ℋ⁡(Q))). To ensure transferability, we use a Basis Point Set (BPS) encoding tied to the gripper’s workspace. A Conditional Variational Autoencoder (CVAE) is trained to reconstruct these contact distributions, while a Links Mapper (PointNet++) learns to associate contact points with specific gripper links. Inference: A novel object 𝒪, positioned at the inverse gripper pose [R,T]−1, is BPS-encoded. By sampling a latent variable z∈ℝψ, the CVAE Decoder generatively predicts diverse, plausible contact points 𝒞^​(𝒪). The Links Mapper then labels which gripper link should reach each point. Generation: Finally, a Force Closure check ensures the predicted contacts yield a stable grasp, and a Grasp Optimization step outputs the final, refined gripper configuration Q∗.
Fig. 3: Overview of GOAG. Geometrical graspability is learned in an object-agnostic manner by focusing on the gripper’s capabilities. Training: We sample gripper configurations Q to generate ℋ⁡(Q) and corresponding contact points (𝒞⁡(ℋ⁡(Q))). To ensure transferability, we use a Basis Point Set (BPS) encoding tied to the gripper’s workspace. A Conditional Variational Autoencoder (CVAE) is trained to reconstruct these contact distributions, while a Links Mapper (PointNet++) learns to associate contact points with specific gripper links. Inference: A novel object 𝒪, positioned at the inverse gripper pose [R,T]−1, is BPS-encoded. By sampling a latent variable z∈ℝψ, the CVAE Decoder generatively predicts diverse, plausible contact points 𝒞^​(𝒪). The Links Mapper then labels which gripper link should reach each point. Generation: Finally, a Force Closure check ensures the predicted contacts yield a stable grasp, and a Grasp Optimization step outputs the final, refined gripper configuration Q∗.
TABLE III: Grasp generalization assessment across multiple grasp datasets. Following [39] we evaluate our method performances with the Shadow hand on multiple grasp test sets. We report the success rates as defined in IV-C. It is worth noting that state-of-the-art methods are retrained for each dataset. GOAG has only been trained once on Shadow hand kinematics.
MethodPer-Dataset TrainingDexGraspNet ↑UniDexGrasp ↑MultiDex ↑RealDex ↑DexGRAB ↑Avg. (%)
UniDexGrasp [35]33.923.721.627.120.825.42
GraspTTA [10]18.621.030.313.314.419.52
SceneDiffuser [9]26.628.369.821.739.137.10
UGG [16]46.946.055.332.742.744.72
DGA [39]57.553.179.144.857.958.48
GOAG43.0749.5177.9037.3762.1353.97
Fig. 4: GOAG grasp results on Multidex [12] objects. Grasps are shown for the Barrett (green), Allegro (pink), and Shadow Hand (purple) grippers.
Fig. 4: GOAG grasp results on Multidex [12] objects. Grasps are shown for the Barrett (green), Allegro (pink), and Shadow Hand (purple) grippers.

为什么重要

仓库或家庭中的机器人会遇到种类繁多的物体,每次遇到新物体都重新收集数据训练的成本很高。GOAG只需针对抓手本身训练一次,就能大幅降低训练成本和时间,为机器人应对陌生物体提供了更实用的抓取方案。

Fig. 5: Real-world setup and results with Allegro hand on YCB [5] objects. First row presents the real robot grasps. Second row presents corresponding virtual grasps. Objects have been rotated around the z-axis for a better understanding of the grasp poses.
Fig. 5: Real-world setup and results with Allegro hand on YCB [5] objects. First row presents the real robot grasps. Second row presents corresponding virtual grasps. Objects have been rotated around the z-axis for a better understanding of the grasp poses.

本文术语

  • 抓手(Gripper) · 用于抓取物体的机械手部分
  • 条件变分自编码器(CVAE) · 一种能根据输入条件生成多样化结果的深度学习生成模型
  • 基点集编码(BPS) · 将三维点云形状转换为固定大小数值表示的编码方法
  • 力闭合(Force Closure) · 判断一组接触点能否稳定夹住物体而不打滑的物理条件
  • PointNet++ · 一种专门用于处理三维点云数据的深度学习网络结构

论文原文摘要(英文)

Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets. We introduce a fundamentally different approach, grounded in the observation that the gripper and the object share identical surface geometry at their mutual contact points. We propose GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation, a novel deep generative model that learns a compact latent representation of a specific gripper's contact surface distribution, enabling the efficient sampling of valid grasp configurations without relying on object-specific training data. We show that by introducing object features only at inference time, our model can effectively retrieve admissible contact areas that are compatible with the gripper's capabilities. We validate our approach through extensive experiments on established grasp protocols in both simulated and real-world scenarios, demonstrating its effectiveness with different grippers from the literature. Our method delivers state-of-the-art results on the objects from the MultiDex dataset, achieving an average success rate of 86.93%. It offers significantly faster processing when generating numerous grasps, while matching the performance of leading approaches specifically trained on this dataset. Unlike these methods, our approach does not rely on object-specific training data, highlighting the advantages of object-agnostic learning. It effectively addresses the generalization challenges faced by traditional data-driven grasp planners. Code and videos are available on our project website https://cea-list.github.io/goagweb/ .

作者 · Julien Merand

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Julien Merand et al., arXiv:2608.19759, CC BY 4.0