Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection
-
摘要: 根据语言指令准确地抓取目标物体,是服务机器人实现自然人机交互的必要前提。现有研究多采用大规模数据驱动训练或层次化特征融合策略实现视觉信息与指令文本的模态对齐,普遍忽视了底层视觉特征中目标与环境的强耦合特性,导致其在跨视角和未知背景下的组合泛化能力显著下降。针对上述挑战,该文构建了一个双视角跨场景同步抓取检测和定位数据集,用于系统评估并针对性提升现有模型的组合泛化能力。在此基础上,提出一种指令驱动下的同步抓取检测和物体定位网络(SGL-Net)。首先,引入跨模态全局上下文调制模块(CGCMM),利用自然语言指令的语义先验,在视觉特征提取的早期阶段对网络视觉通道进行自适应视角校正与背景干扰抑制。其次,设计了词-像素跨模态交叉对齐模块(WPCAM),采用扁平化注意力机制实现细粒度语义解耦,用以提升网络对动态复杂场景的语义理解能力。最后,采用统一解码器架构解耦融合后的多模态特征,实现目标物体位置与最优抓取位姿的同步输出。充足的定量与可视化实验结果表明,所提网络在多种不同场景视角的组合情况下均取得了最优的性能,并具备在真实物理世界中部署应用的能力。Abstract:
Objective Accurate grasp detection from language instructions is essential for service robots to achieve natural human-robot interaction. Existing methods primarily rely on large-scale data-driven training or hierarchical feature fusion to align visual perception with textual instructions. However, they generally overlook the strong coupling between target objects and background clutter in low-level visual features, leading to degraded compositional generalization under cross-view and unseen-scene conditions. To address this limitation, a dual-view cross-scene grasp detection and object localization dataset is constructed to systematically evaluate and improve the compositional generalization of existing models. Based on this benchmark, a Simultaneous Grasp detection and object Localization Network (SGL-Net) is proposed to jointly predict object locations and optimal grasp poses. The proposed framework enables service robots to manipulate objects according to natural language instructions in real-world dynamic environments, providing technical support for embodied intelligence. Methods The proposed SGL-Net is illustrated in Fig. 1. First, a Cross-modal Global Context Modulation Module (CGCMM) is proposed to exploit semantic priors from language instructions for adaptive viewpoint correction and background suppression during the early stage of visual feature extraction. Second, a Word-Pixel Cross-modal Alignment Module (WPCAM) is designed to achieve fine-grained semantic decoupling through a flattening-based cross-modal attention mechanism, thereby improving semantic understanding in complex dynamic scenes. Finally, a unified decoder jointly predicts object locations and optimal grasp poses from the fused multimodal features. Results and Discussions Extensive quantitative and qualitative experiments are conducted on the reconstructed dual-view cross-scene dataset containing bottom-view and top-view scenes and on a real-world robotic grasping platform. Comparative results demonstrate that SGL-Net consistently outperforms mainstream CNN-based and CLIP-based methods in both grasp detection and object localization ( Tables 2 and3 ). Ablation studies further verify the effectiveness of CGCMM and WPCAM in improving fine-grained semantic alignment and semantic decoupling (Tables 4 and5 ). Furthermore, qualitative results (Figs. 4 ~6 ) and real-world robotic experiments (Fig. 7 ) demonstrate that SGL-Net can be reliably deployed in complex physical environments. Overall, the proposed network exhibits strong generalization capability and excellent potential for practical robotic applications.Conclusions To improve cross-view and cross-scene generalization in language-conditioned robotic grasp detection, this paper constructs a dedicated validation dataset and proposes SGL-Net, which jointly performs grasp detection and object localization. By integrating CGCMM and WPCAM, the proposed network accurately localizes instruction-specified objects and predicts optimal grasp poses. Experimental results obtained on multiple benchmark scenarios and a real-world robotic platform demonstrate the superior performance and practical applicability of the proposed method. Future work will focus on integrating Large Multimodal Models (LMMs) and adapting the proposed framework through fine-tuning to further improve zero-shot robotic grasp detection. -
表 1 数据集介绍
场景案例 训练集合 测试集合 测试样本视角 案例1 35913 8092 底部(Bottom) 案例2 37606 9084 顶部(Top) 表 2 本文算法及对比算法在案例1测试集合上的实验结果(%)
算法 抓取检测精度 物体定位精度 $ 30_{0.3}^{{^{\circ}}} $ $ 30_{0.4}^{{^{\circ}}} $ $ {0.25}_{{{10}^{{^{\circ}}}}} $ $ {0.25}_{{{20}^{{^{\circ}}}}} $ GA$ (30_{0.25}^{{^{\circ}}}) $ P@0.5 P@0.7 MIoU OIoU GRCNN[24] 16.34 10.74 11.05 15.31 18.35 0.68 0 5.47 6.04 GGCNN[26] 3.62 1.45 2.14 4.34 5.14 1.69 0.27 5.62 6.19 GGCNN2[25] 20.74 11.70 10.52 18.24 23.22 0.87 0.14 5.66 6.62 CLIPort[14] 68.56 58.69 56.43 66.46 70.09 41.09 11.30 39.67 40.21 ELGNet[15] 65.67 55.72 52.46 63.37 67.45 15.89 6.52 31.21 30.82 本文 73.23 61.58 59.82 71.30 74.63 56.07 21.77 47.55 45.23 表 3 本文算法及对比算法在案例2测试集合上的实验结果(%)
算法 抓取检测精度 物体定位精度 $ 30_{0.3}^{{^{\circ}}} $ $ 30_{0.4}^{{^{\circ}}} $ $ {0.25}_{{{10}^{{^{\circ}}}}} $ $ {0.25}_{{{20}^{{^{\circ}}}}} $ GA$ (30_{0.25}^{{^{\circ}}}) $ P@0.5 P@0.7 MIoU OIoU GRCNN[24] 16.29 12.88 10.76 13.78 17.73 0.85 0.31 5.72 5.87 GGCNN[26] 2.20 1.06 1.19 2.50 2.96 1.23 0.00 6.72 7.28 GGCNN2[25] 24.04 20.15 17.81 22.62 25.28 1.01 0.20 6.28 7.11 CLIPort[14] 73.80 70.15 60.34 71.04 74.31 71.47 39.02 57.37 54.21 ELGNet[15] 73.08 65.20 59.0 69.97 74.21 70.66 33.40 55.21 50.63 本文 76.57 71.75 63.72 73.49 77.34 76.20 45.79 59.73 56.66 表 4 案例1消融实验结果(%)
CGCMM WPCAM GA P@0.5 MIoU OIoU √ √ 74.63 56.07 47.55 45.23 √ 71.18 45.71 41.28 42.11 69.64 41.40 40.29 37.92 表 5 案例2消融实验结果(%)
CGCMM WPCAM GA P@0.5 MIoU OIoU √ √ 77.34 76.20 59.73 56.66 √ 74.08 68.97 55.88 53.07 71.95 60.34 47.33 45.42 表 6 机器人抓取实验
算法 Top视角 Bottom视角 CLIPort 70%(35/50) 80%(40/50) ELGNet 72%(36/50) 74%(37/50) 本文算法 84%(42/50) 92%(46/50) -
[1] 张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193. [2] 余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301. [3] LI Renjie, DONG Wei, SUN Jiarui, et al. A fast integrated gait, footstep, and motion planning framework for wheeled-legged robots[J]. Journal of Bionic Engineering, 2026, 23(2): 607–621. doi: 10.1007/s42235-025-00830-5. [4] LIU Zhitai, ZHONG Hanhai, LIN Xiaotian, et al. Integrated motion control framework of wheel-legged biped robot on rugged terrain[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7383–7394. doi: 10.1109/TMECH.2025.3627438. [5] 崔永成, 田国会, 周昭旭, 等. 智能空间下面向动作序列生成的服务机器人指令解析方法[J]. 机器人, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.CUI Yongcheng, TIAN Guohui, ZHOU Zhaoxu, et al. A service robot instruction parsing method for action sequence generation in intelligent space[J]. Robot, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074. [6] LENZ I, LEE H, and SAXENA A. Deep learning for detecting robotic grasps[J]. The International Journal of Robotics Research, 2015, 34(4/5): 705–724. doi: 10.1177/0278364914549607. [7] KUMRA S and KANAN C. Robotic grasp detection using deep convolutional neural networks[C]. 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017: 769–776. doi: 10.1109/IROS.2017.8202237. [8] LAILI Yuanjun, CHEN Zelin, REN Lei, et al. Custom grasping: A region-based robotic grasping detection method in industrial cyber-physical systems[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(1): 88–100. doi: 10.1109/TASE.2021.3139610. [9] CHU F J, XU Ruinian, and VELA P A. Real-world multiobject, multigrasp detection[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3355–3362. doi: 10.1109/LRA.2018.2852777. [10] BOUSSELHAM W, PETERSEN F, FERRARI V, et al. Grounding everything: Emerging localization properties in vision-language transformers[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 3828–3837. doi: 10.1109/CVPR52733.2024.00367. [11] YIN Heng, REN Yuqiang, YAN Ke, et al. ROD-MLLM: Towards more reliable object detection in multimodal large language models[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2025: 14358–14368. doi: 10.1109/CVPR52734.2025.01339. [12] ZHOU Yi, SU Hang, WANG Tian, et al. Onet: Twin U-Net architecture for unsupervised binary semantic segmentation in radar and remote sensing images[J]. IEEE Transactions on Image Processing, 2025, 34: 2161–2172. doi: 10.1109/TIP.2025.3530816. [13] WANG Hao, HU Keyan, GUO Xin, et al. A gift from the integration of discriminative and diffusion-based generative learning: Boundary refinement remote sensing semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5892–5909. doi: 10.1109/TPAMI.2026.3654243. [14] SHRIDHAR M, MANUELLI L, and FOX D. CLIPort: What and where pathways for robotic manipulation[C]. Proceedings of the 5th Conference on Robot Learning, London, UK, 2021: 894–906. [15] LIU Jin, XIE Jialong, and XIAO Leibing. Hierarchical multi-modal fusion for language-conditioned robotic grasping detection in clutter[J]. IEEE Robotics and Automation Letters, 2024, 9(10): 8762–8769. doi: 10.1109/LRA.2024.3440833. [16] VAN VO T, VU M N, HUANG Baoru, et al. Language-driven grasp detection with mask-guided attention[C]. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 7492–7498. doi: 10.1109/IROS58592.2024.10802256. [17] XIE Jialong, LIU Jin, ZHU Zhenwei, et al. Infusing multisource heterogeneous knowledge for language-conditioned segmentation and grasping[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5029611. doi: 10.1109/TIM.2024.3446625. [18] XIE Jialong, ZHOU Fengyu, LIU Jin, et al. Semi-supervised language-conditioned grasping with curriculum-scheduled augmentation and geometric consistency[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 4021–4028. doi: 10.1109/LRA.2025.3547619. [19] 张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.ZHANG Mei, JIN Ye, ZHU Jinhui, et al. FSG: Feature-level semantic-aware guidance for multi-modal image fusion algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042. [20] LI Jindong, LI Yongguang, FU Yali, et al. CLIP-powered domain generalization and domain adaptation: A comprehensive survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5405–5424. doi: 10.1109/TPAMI.2026.3651700. [21] ZHOU Zhenning, ZHU Xiaoxiao, and CAO Qixin. AAGDN: Attention-augmented grasp detection network based on coordinate attention and effective feature fusion method[J]. IEEE Robotics and Automation Letters, 2023, 8(6): 3462–3469. doi: 10.1109/LRA.2023.3268596. [22] WANG Dexin, LIU Chunsheng, CHANG Faliang, et al. High-performance pixel-level grasp detection based on adaptive grasping and grasp-aware network[J]. IEEE Transactions on Industrial Electronics, 2021, 69(11): 11611–11621. doi: 10.1109/TIE.2021.3120474. [23] REZATOFIGHI H, TSOI N, GWAK J, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 658–666. doi: 10.1109/CVPR.2019.00075. [24] KUMRA S, JOSHI S, and SAHIN F. Antipodal robotic grasping using generative residual convolutional neural network[C]. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, 2020: 9626–9633. doi: 10.1109/IROS45743.2020.9340777. [25] XIE Jialong, LIU Jin, HUANG Saike, et al. Listen, perceive, grasp: CLIP-driven attribute-aware network for language-conditioned visual segmentation and grasping[J]. IEEE Transactions on Automation Science and Engineering, 2025, 22: 9729–9740. doi: 10.1109/TASE.2024.3510777. [26] MORRISON D, CORKE P, and LEITNER J. Learning robust, real-time, reactive robotic grasping[J]. The International Journal of Robotics Research, 2020, 39(2/3): 183–201. doi: 10.1177/0278364919859066. -
下载: