高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

融合全局视角校正与细粒度语义解耦的机器人指令抓取检测方法

刘进 刘志太 李梓函 孙彦景 缪燕子 袁宪锋

刘进, 刘志太, 李梓函, 孙彦景, 缪燕子, 袁宪锋. 融合全局视角校正与细粒度语义解耦的机器人指令抓取检测方法[J]. 电子与信息学报. doi: 10.11999/JEIT260442
引用本文: 刘进, 刘志太, 李梓函, 孙彦景, 缪燕子, 袁宪锋. 融合全局视角校正与细粒度语义解耦的机器人指令抓取检测方法[J]. 电子与信息学报. doi: 10.11999/JEIT260442
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442
Citation: LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442

融合全局视角校正与细粒度语义解耦的机器人指令抓取检测方法

doi: 10.11999/JEIT260442 cstr: 32379.14.JEIT260442
基金项目: 国家重点研发计划(2025YFB4712700),中国博士后科学基金资助项目(2025M781675)
详细信息
    作者简介:

    刘进:男,副教授,研究方向为机器人具身操作

    刘志太:男,副研究员,研究方向为机器人智能系统及智能显微操作

    李梓函:女,无,研究生,研究方向为六自由度物体姿态估计

    孙彦景:男,教授,研究方向为智能安全与应急通信

    缪燕子:女,教授,研究方向为机器视觉与自动驾驶

    袁宪锋:男,副教授,研究方向为智能机器人、具身智能及视觉语言导航

    通讯作者:

    袁宪锋 yuanxianfeng@sdu.edu.cn

  • 中图分类号: TP242.6

Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection

Funds: National Key R&D Program of China (Grant No. 2025YFB4712700), and the China Postdoctoral Science Foundation (Grant No. 2025M781675)
  • 摘要: 根据语言指令准确地抓取目标物体,是服务机器人实现自然人机交互的必要前提。现有研究多采用大规模数据驱动训练或层次化特征融合策略实现视觉信息与指令文本的模态对齐,普遍忽视了底层视觉特征中目标与环境的强耦合特性,导致其在跨视角和未知背景下的组合泛化能力显著下降。针对上述挑战,本文构建了一个双视角跨场景同步抓取检测和定位数据集,用于系统评估并针对性提升现有模型的组合泛化能力。在此基础上,提出了一种指令驱动下的同步抓取检测和物体定位网络(SGL-Net)。首先,引入了跨模态全局上下文调制模块(CGCMM),利用自然语言指令的语义先验,在视觉特征提取的早期阶段对网络视觉通道进行自适应视角校正与背景干扰抑制。其次,设计了词-像素跨模态交叉对齐模块(WPCAM),采用扁平化注意力机制实现细粒度语义解耦,用以提升网络对动态复杂场景的语义理解能力。最后,采用统一解码器架构解耦融合后的多模态特征,实现目标物体位置与最优抓取位姿的同步输出。充足的定量与可视化实验结果表明,所提网络在多种不同场景视角的组合情况下均取得了最优的性能,并具备在真实物理世界中部署应用的能力。
  • 图  1  本文所提网络架构图

    图  2  融合策略对比图

    图  3  数据集可视化案例

    图  4  不同案例抓取检测及定位可视化结果。箭头所示为不精确预测,红色框为物体预测位置,蓝色线上方为案例1,下方为案例2。

    图  5  位于相似背景时抓取检测及定位可视化结果。箭头所示为不精确预测,红色框为物体预测位置,蓝色线上方为案例1,下方为案例2。

    图  6  其他不同背景下的泛化实验可视化结果。红色框为预测的物体位置,蓝色线上方为案例1,下方为案例2。

    图  7  真机抓取实验可视化结果。红色框为预测的物体位置,蓝红框为预测的抓取位姿

    表  1  数据集介绍

    场景案例训练集合测试集合测试样本视角
    案例1359138092底部(Bottom)
    案例2376069084顶部(Top)
    下载: 导出CSV

    表  2  本文算法及对比算法在案例1测试集合上的实验结果(%)

    算法 抓取检测精度 物体定位精度
    $ 30_{0.3}^{{^{\circ}}} $ $ 30_{0.4}^{{^{\circ}}} $ $ {0.25}_{{{10}^{{^{\circ}}}}} $ $ {0.25}_{{{20}^{{^{\circ}}}}} $ GA$ (30_{0.25}^{{^{\circ}}}) $ P@0.5 P@0.7 MIoU OIoU
    GRCNN[24] 16.34 10.74 11.05 15.31 18.35 0.68 0 5.47 6.04
    GGCNN[25] 3.62 1.45 2.14 4.34 5.14 1.69 0.27 5.62 6.19
    GGCNN2[26] 20.74 11.70 10.52 18.24 23.22 0.87 0.14 5.66 6.62
    CLIPort[14] 68.56 58.69 56.43 66.46 70.09 41.09 11.30 39.67 40.21
    ELGNet[15] 65.67 55.72 52.46 63.37 67.45 15.89 6.52 31.21 30.82
    本文算法 73.23 61.58 59.82 71.30 74.63 56.07 21.77 47.55 45.23
    下载: 导出CSV

    表  3  本文算法及对比算法在案例2测试集合上的实验结果(%)

    算法抓取检测精度物体定位精度
    $ 30_{0.3}^{{^{\circ}}} $$ 30_{0.4}^{{^{\circ}}} $$ {0.25}_{{{10}^{{^{\circ}}}}} $$ {0.25}_{{{20}^{{^{\circ}}}}} $GA$ (30_{0.25}^{{^{\circ}}}) $P@0.5P@0.7MIoUOIoU
    GRCNN[24]16.2912.8810.7613.7817.730.850.315.725.87
    GGCNN[25]2.201.061.192.502.961.230.006.727.28
    GGCNN2[26]24.0420.1517.8122.6225.281.010.206.287.11
    CLIPort[14]73.8070.1560.3471.0474.3171.4739.0257.3754.21
    ELGNet[15]73.0865.2059.069.9774.2170.6633.4055.2150.63
    本文算法76.5771.7563.7273.4977.3476.2045.7959.7356.66
    下载: 导出CSV

    表  4  案例1消融实验结果

    CGCMMWPCAMGAP@0.5MIoUOIoU
    74.6356.0747.5545.23
    71.1845.7141.2842.11
    69.6441.4040.2937.92
    下载: 导出CSV

    表  5  案例2消融实验结果

    CGCMMWPCAMGAP@0.5MIoUOIoU
    77.3476.2059.7356.66
    74.0868.9755.8853.07
    71.9560.3447.3345.42
    下载: 导出CSV

    表  6  机器人抓取实验

    算法Top视角Bottom视角
    CLIPort70%(35/50)80%(40/50)
    ELGNet72%(36/50)74%(37/50)
    本文算法84%(42/50)92%(46/50)
    下载: 导出CSV
  • [1] 张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.

    ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.
    [2] 余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.

    YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.
    [3] LI Renjie, DONG Wei, SUN Jiarui, et al. A fast integrated gait, footstep, and motion planning framework for wheeled-legged robots[J]. Journal of Bionic Engineering, 2026, 23(2): 607–621. doi: 10.1007/s42235-025-00830-5.
    [4] LIU Zhitai, ZHONG Hanhai, LIN Xiaotian, et al. Integrated motion control framework of wheel-legged biped robot on rugged terrain[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7383–7394. doi: 10.1109/TMECH.2025.3627438.
    [5] 崔永成, 田国会, 周昭旭, 等. 智能空间下面向动作序列生成的服务机器人指令解析方法[J]. 机器人, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.

    CUI Yongcheng, TIAN Guohui, ZHOU Zhaoxu, et al. A service robot instruction parsing method for action sequence generation in intelligent space[J]. Robot, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.
    [6] LENZ I, LEE H, and SAXENA A. Deep learning for detecting robotic grasps[J]. The International Journal of Robotics Research, 2015, 34(4/5): 705–724. doi: 10.1177/0278364914549607.
    [7] KUMRA S and KANAN C. Robotic grasp detection using deep convolutional neural networks[C]. 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017: 769–776. doi: 10.1109/IROS.2017.8202237.
    [8] LAILI Yuanjun, CHEN Zelin, REN Lei, et al. Custom grasping: A region-based robotic grasping detection method in industrial cyber-physical systems[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(1): 88–100. doi: 10.1109/TASE.2021.3139610.
    [9] CHU F J, XU Ruinian, and VELA P A. Real-world multiobject, multigrasp detection[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3355–3362. doi: 10.1109/LRA.2018.2852777.
    [10] BOUSSELHAM W, PETERSEN F, FERRARI V, et al. Grounding everything: Emerging localization properties in vision-language transformers[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 3828–3837. doi: 10.1109/CVPR52733.2024.00367.
    [11] YIN Heng, REN Yuqiang, YAN Ke, et al. ROD-MLLM: Towards more reliable object detection in multimodal large language models[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2025: 14358–14368. doi: 10.1109/CVPR52734.2025.01339.
    [12] ZHOU Yi, SU Hang, WANG Tian, et al. Onet: Twin U-Net architecture for unsupervised binary semantic segmentation in radar and remote sensing images[J]. IEEE Transactions on Image Processing, 2025, 34: 2161–2172. doi: 10.1109/TIP.2025.3530816.
    [13] WANG Hao, HU Keyan, GUO Xin, et al. A gift from the integration of discriminative and diffusion-based generative learning: Boundary refinement remote sensing semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5892–5909. doi: 10.1109/TPAMI.2026.3654243.
    [14] SHRIDHAR M, MANUELLI L, and FOX D. CLIPort: What and where pathways for robotic manipulation[C]. Proceedings of the 5th Conference on Robot Learning, London, UK, 2021: 894–906.
    [15] LIU Jin, XIE Jialong, and XIAO Leibing. Hierarchical multi-modal fusion for language-conditioned robotic grasping detection in clutter[J]. IEEE Robotics and Automation Letters, 2024, 9(10): 8762–8769. doi: 10.1109/LRA.2024.3440833.
    [16] VAN VO T, VU M N, HUANG Baoru, et al. Language-driven grasp detection with mask-guided attention[C]. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 7492–7498. doi: 10.1109/IROS58592.2024.10802256.
    [17] XIE Jialong, LIU Jin, ZHU Zhenwei, et al. Infusing multisource heterogeneous knowledge for language-conditioned segmentation and grasping[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5029611. doi: 10.1109/TIM.2024.3446625.
    [18] XIE Jialong, ZHOU Fengyu, LIU Jin, et al. Semi-supervised language-conditioned grasping with curriculum-scheduled augmentation and geometric consistency[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 4021–4028. doi: 10.1109/LRA.2025.3547619.
    [19] 张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.

    ZHANG Mei, JIN Ye, ZHU Jinhui, et al. FSG: Feature-level semantic-aware guidance for multi-modal image fusion algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.
    [20] LI Jindong, LI Yongguang, FU Yali, et al. CLIP-powered domain generalization and domain adaptation: A comprehensive survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5405–5424. doi: 10.1109/TPAMI.2026.3651700.
    [21] ZHOU Zhenning, ZHU Xiaoxiao, and CAO Qixin. AAGDN: Attention-augmented grasp detection network based on coordinate attention and effective feature fusion method[J]. IEEE Robotics and Automation Letters, 2023, 8(6): 3462–3469. doi: 10.1109/LRA.2023.3268596.
    [22] WANG Dexin, LIU Chunsheng, CHANG Faliang, et al. High-performance pixel-level grasp detection based on adaptive grasping and grasp-aware network[J]. IEEE Transactions on Industrial Electronics, 2021, 69(11): 11611–11621. doi: 10.1109/TIE.2021.3120474.
    [23] REZATOFIGHI H, TSOI N, GWAK J, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 658–666. doi: 10.1109/CVPR.2019.00075.
    [24] KUMRA S, JOSHI S, and SAHIN F. Antipodal robotic grasping using generative residual convolutional neural network[C]. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, 2020: 9626–9633. doi: 10.1109/IROS45743.2020.9340777.
    [25] MORRISON D, CORKE P, and LEITNER J. Learning robust, real-time, reactive robotic grasping[J]. The International Journal of Robotics Research, 2020, 39(2/3): 183–201. doi: 10.1177/0278364919859066.
    [26] XIE Jialong, LIU Jin, HUANG Saike, et al. Listen, perceive, grasp: CLIP-driven attribute-aware network for language-conditioned visual segmentation and grasping[J]. IEEE Transactions on Automation Science and Engineering, 2025, 22: 9729–9740. doi: 10.1109/TASE.2024.3510777.
  • 加载中
图(7) / 表(6)
计量
  • 文章访问数:  67
  • HTML全文浏览量:  31
  • PDF下载量:  1
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-04-14
  • 修回日期:  2026-07-14
  • 录用日期:  2026-07-14
  • 网络出版日期:  2026-07-24

目录

    /

    返回文章
    返回