Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection
-
摘要: 根据语言指令准确地抓取目标物体,是服务机器人实现自然人机交互的必要前提。现有研究多采用大规模数据驱动训练或层次化特征融合策略实现视觉信息与指令文本的模态对齐,普遍忽视了底层视觉特征中目标与环境的强耦合特性,导致其在跨视角和未知背景下的组合泛化能力显著下降。针对上述挑战,本文构建了一个双视角跨场景同步抓取检测和定位数据集,用于系统评估并针对性提升现有模型的组合泛化能力。在此基础上,提出了一种指令驱动下的同步抓取检测和物体定位网络(SGL-Net)。首先,引入了跨模态全局上下文调制模块(CGCMM),利用自然语言指令的语义先验,在视觉特征提取的早期阶段对网络视觉通道进行自适应视角校正与背景干扰抑制。其次,设计了词-像素跨模态交叉对齐模块(WPCAM),采用扁平化注意力机制实现细粒度语义解耦,用以提升网络对动态复杂场景的语义理解能力。最后,采用统一解码器架构解耦融合后的多模态特征,实现目标物体位置与最优抓取位姿的同步输出。充足的定量与可视化实验结果表明,所提网络在多种不同场景视角的组合情况下均取得了最优的性能,并具备在真实物理世界中部署应用的能力。Abstract:
Objective Accurate grasping of target objects from language instructions is an essential prerequisite for service robots to achieve natural human-robot interaction. Existing research primarily relies on massive data-driven training or hierarchical feature fusion schemes for the cross-modal alignment between visual perception and textual instructions. However, these methods commonly overlook the heavy entanglement between target objects and background environments within low-level features, leading to a significant degradation in compositional generalization ability under cross-view and unseen background scenarios. To address these issues, this paper constructs a dual-view cross-scenario synchronous grasp detection and localization dataset, which is used to systematically evaluate and specifically improve the compositional generalization ability of existing models. Building upon this benchmark, we propose a simultaneous detection and localization network to simultaneously output the target location and the optimal grasping pose. Consequently, service robots equipped with the proposed network can perform object manipulation in real-world dynamic environments according to human instructions, providing theoretical foundations and technical support for embodied intelligence. Methods The structure of the proposed Simultaneous Grasp and Localization Network (SGL-Net) is illustrated ( Fig. 1 ). Firstly, a cross-modal global context modulation module (CGCMM) is proposed. The module utilizes instruction priors to perform adaptive viewpoint correction and background suppression on visual channels during the early stages of feature extraction. Secondly, a word-pixel cross-modal alignment module (WPCAM) is designed to achieve fine-grained semantic decoupling via a flattened attention mechanism, further enhancing the model's comprehension in dynamic scenes. Finally, a unified decoder is employed to simultaneously output the target location and the optimal grasping pose.Results and Discussions Extensive quantitative and qualitative experiments are conducted on a restructured dual-view dataset, including bottom view and top view, and a real-world physical robotic platform. Comparative results demonstrate that SGL-Net significantly outperforms mainstream CNN-based and CLIP-based baseline methods in both grasp detection and object localization metrics ( Table 2 andTable 3 ). Ablation studies explicitly validate the critical contributions of two designed modules in achieving fine-grained semantic alignment and scene decoupling (Table 4 andTable 5 ). Furthermore, qualitative visualizations (Fig. 4 ,Fig. 5 , andFig. 6 ) and real-world physical experiments (Fig. 7 ) indicate that SGL-Net exhibits strong viability for direct deployment in real-world physical environments. Overall, the network exhibits superior robust generalization capabilities and reliable practical deployment potential when confronting complex physical environments.Conclusions To address the challenge of insufficient cross-view and cross-domain generalization capability in robotic grasping operations, this paper constructs a corresponding validation dataset and proposes a simultaneous grasp detection and object localization network. By integrating a cross-modal global context modulation module and a word-pixel cross-modal alignment module, the proposed network achieves accurate localization of instruction-specified objects and prediction of optimal grasp poses. Experimental results from diverse testing datasets and real-world grasping experiments demonstrate the superior performance of the proposed network, highlighting its potential for deployment in practical scenarios. To further enhance robotic grasping capabilities in zero-shot scenarios, future work will focus on integrating Large Multimodal Models and realizing network adaptation through fine-tuning techniques. -
Key words:
- Grasp detection /
- Language-conditioned /
- Multi-modal fusion /
- Complex scenarios
-
表 1 数据集介绍
场景案例 训练集合 测试集合 测试样本视角 案例1 35913 8092 底部(Bottom) 案例2 37606 9084 顶部(Top) 表 2 本文算法及对比算法在案例1测试集合上的实验结果(%)
算法 抓取检测精度 物体定位精度 $ 30_{0.3}^{{^{\circ}}} $ $ 30_{0.4}^{{^{\circ}}} $ $ {0.25}_{{{10}^{{^{\circ}}}}} $ $ {0.25}_{{{20}^{{^{\circ}}}}} $ GA$ (30_{0.25}^{{^{\circ}}}) $ P@0.5 P@0.7 MIoU OIoU GRCNN[24] 16.34 10.74 11.05 15.31 18.35 0.68 0 5.47 6.04 GGCNN[25] 3.62 1.45 2.14 4.34 5.14 1.69 0.27 5.62 6.19 GGCNN2[26] 20.74 11.70 10.52 18.24 23.22 0.87 0.14 5.66 6.62 CLIPort[14] 68.56 58.69 56.43 66.46 70.09 41.09 11.30 39.67 40.21 ELGNet[15] 65.67 55.72 52.46 63.37 67.45 15.89 6.52 31.21 30.82 本文算法 73.23 61.58 59.82 71.30 74.63 56.07 21.77 47.55 45.23 表 3 本文算法及对比算法在案例2测试集合上的实验结果(%)
算法 抓取检测精度 物体定位精度 $ 30_{0.3}^{{^{\circ}}} $ $ 30_{0.4}^{{^{\circ}}} $ $ {0.25}_{{{10}^{{^{\circ}}}}} $ $ {0.25}_{{{20}^{{^{\circ}}}}} $ GA$ (30_{0.25}^{{^{\circ}}}) $ P@0.5 P@0.7 MIoU OIoU GRCNN[24] 16.29 12.88 10.76 13.78 17.73 0.85 0.31 5.72 5.87 GGCNN[25] 2.20 1.06 1.19 2.50 2.96 1.23 0.00 6.72 7.28 GGCNN2[26] 24.04 20.15 17.81 22.62 25.28 1.01 0.20 6.28 7.11 CLIPort[14] 73.80 70.15 60.34 71.04 74.31 71.47 39.02 57.37 54.21 ELGNet[15] 73.08 65.20 59.0 69.97 74.21 70.66 33.40 55.21 50.63 本文算法 76.57 71.75 63.72 73.49 77.34 76.20 45.79 59.73 56.66 表 4 案例1消融实验结果
CGCMM WPCAM GA P@0.5 MIoU OIoU √ √ 74.63 56.07 47.55 45.23 √ 71.18 45.71 41.28 42.11 69.64 41.40 40.29 37.92 表 5 案例2消融实验结果
CGCMM WPCAM GA P@0.5 MIoU OIoU √ √ 77.34 76.20 59.73 56.66 √ 74.08 68.97 55.88 53.07 71.95 60.34 47.33 45.42 表 6 机器人抓取实验
算法 Top视角 Bottom视角 CLIPort 70%(35/50) 80%(40/50) ELGNet 72%(36/50) 74%(37/50) 本文算法 84%(42/50) 92%(46/50) -
[1] 张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193. [2] 余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301. [3] LI Renjie, DONG Wei, SUN Jiarui, et al. A fast integrated gait, footstep, and motion planning framework for wheeled-legged robots[J]. Journal of Bionic Engineering, 2026, 23(2): 607–621. doi: 10.1007/s42235-025-00830-5. [4] LIU Zhitai, ZHONG Hanhai, LIN Xiaotian, et al. Integrated motion control framework of wheel-legged biped robot on rugged terrain[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7383–7394. doi: 10.1109/TMECH.2025.3627438. [5] 崔永成, 田国会, 周昭旭, 等. 智能空间下面向动作序列生成的服务机器人指令解析方法[J]. 机器人, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.CUI Yongcheng, TIAN Guohui, ZHOU Zhaoxu, et al. A service robot instruction parsing method for action sequence generation in intelligent space[J]. Robot, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074. [6] LENZ I, LEE H, and SAXENA A. Deep learning for detecting robotic grasps[J]. The International Journal of Robotics Research, 2015, 34(4/5): 705–724. doi: 10.1177/0278364914549607. [7] KUMRA S and KANAN C. Robotic grasp detection using deep convolutional neural networks[C]. 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017: 769–776. doi: 10.1109/IROS.2017.8202237. [8] LAILI Yuanjun, CHEN Zelin, REN Lei, et al. Custom grasping: A region-based robotic grasping detection method in industrial cyber-physical systems[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(1): 88–100. doi: 10.1109/TASE.2021.3139610. [9] CHU F J, XU Ruinian, and VELA P A. Real-world multiobject, multigrasp detection[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3355–3362. doi: 10.1109/LRA.2018.2852777. [10] BOUSSELHAM W, PETERSEN F, FERRARI V, et al. Grounding everything: Emerging localization properties in vision-language transformers[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 3828–3837. doi: 10.1109/CVPR52733.2024.00367. [11] YIN Heng, REN Yuqiang, YAN Ke, et al. ROD-MLLM: Towards more reliable object detection in multimodal large language models[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2025: 14358–14368. doi: 10.1109/CVPR52734.2025.01339. [12] ZHOU Yi, SU Hang, WANG Tian, et al. Onet: Twin U-Net architecture for unsupervised binary semantic segmentation in radar and remote sensing images[J]. IEEE Transactions on Image Processing, 2025, 34: 2161–2172. doi: 10.1109/TIP.2025.3530816. [13] WANG Hao, HU Keyan, GUO Xin, et al. A gift from the integration of discriminative and diffusion-based generative learning: Boundary refinement remote sensing semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5892–5909. doi: 10.1109/TPAMI.2026.3654243. [14] SHRIDHAR M, MANUELLI L, and FOX D. CLIPort: What and where pathways for robotic manipulation[C]. Proceedings of the 5th Conference on Robot Learning, London, UK, 2021: 894–906. [15] LIU Jin, XIE Jialong, and XIAO Leibing. Hierarchical multi-modal fusion for language-conditioned robotic grasping detection in clutter[J]. IEEE Robotics and Automation Letters, 2024, 9(10): 8762–8769. doi: 10.1109/LRA.2024.3440833. [16] VAN VO T, VU M N, HUANG Baoru, et al. Language-driven grasp detection with mask-guided attention[C]. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 7492–7498. doi: 10.1109/IROS58592.2024.10802256. [17] XIE Jialong, LIU Jin, ZHU Zhenwei, et al. Infusing multisource heterogeneous knowledge for language-conditioned segmentation and grasping[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5029611. doi: 10.1109/TIM.2024.3446625. [18] XIE Jialong, ZHOU Fengyu, LIU Jin, et al. Semi-supervised language-conditioned grasping with curriculum-scheduled augmentation and geometric consistency[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 4021–4028. doi: 10.1109/LRA.2025.3547619. [19] 张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.ZHANG Mei, JIN Ye, ZHU Jinhui, et al. FSG: Feature-level semantic-aware guidance for multi-modal image fusion algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042. [20] LI Jindong, LI Yongguang, FU Yali, et al. CLIP-powered domain generalization and domain adaptation: A comprehensive survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5405–5424. doi: 10.1109/TPAMI.2026.3651700. [21] ZHOU Zhenning, ZHU Xiaoxiao, and CAO Qixin. AAGDN: Attention-augmented grasp detection network based on coordinate attention and effective feature fusion method[J]. IEEE Robotics and Automation Letters, 2023, 8(6): 3462–3469. doi: 10.1109/LRA.2023.3268596. [22] WANG Dexin, LIU Chunsheng, CHANG Faliang, et al. High-performance pixel-level grasp detection based on adaptive grasping and grasp-aware network[J]. IEEE Transactions on Industrial Electronics, 2021, 69(11): 11611–11621. doi: 10.1109/TIE.2021.3120474. [23] REZATOFIGHI H, TSOI N, GWAK J, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 658–666. doi: 10.1109/CVPR.2019.00075. [24] KUMRA S, JOSHI S, and SAHIN F. Antipodal robotic grasping using generative residual convolutional neural network[C]. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, 2020: 9626–9633. doi: 10.1109/IROS45743.2020.9340777. [25] MORRISON D, CORKE P, and LEITNER J. Learning robust, real-time, reactive robotic grasping[J]. The International Journal of Robotics Research, 2020, 39(2/3): 183–201. doi: 10.1177/0278364919859066. [26] XIE Jialong, LIU Jin, HUANG Saike, et al. Listen, perceive, grasp: CLIP-driven attribute-aware network for language-conditioned visual segmentation and grasping[J]. IEEE Transactions on Automation Science and Engineering, 2025, 22: 9729–9740. doi: 10.1109/TASE.2024.3510777. -
下载: