Advanced Search
Turn off MathJax
Article Contents
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442
Citation: LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng. Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260442

Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection

doi: 10.11999/JEIT260442 cstr: 32379.14.JEIT260442
Funds:  The National Key R&D Program of China (2025YFB4712700), China Postdoctoral Science Foundation (2025M781675)
  • Received Date: 2026-04-15
  • Accepted Date: 2026-07-14
  • Rev Recd Date: 2026-07-13
  • Available Online: 2026-07-24
  •   Objective  Accurate grasp detection from language instructions is essential for service robots to achieve natural human-robot interaction. Existing methods primarily rely on large-scale data-driven training or hierarchical feature fusion to align visual perception with textual instructions. However, they generally overlook the strong coupling between target objects and background clutter in low-level visual features, leading to degraded compositional generalization under cross-view and unseen-scene conditions. To address this limitation, a dual-view cross-scene grasp detection and object localization dataset is constructed to systematically evaluate and improve the compositional generalization of existing models. Based on this benchmark, a Simultaneous Grasp detection and object Localization Network (SGL-Net) is proposed to jointly predict object locations and optimal grasp poses. The proposed framework enables service robots to manipulate objects according to natural language instructions in real-world dynamic environments, providing technical support for embodied intelligence.  Methods  The proposed SGL-Net is illustrated in Fig. 1. First, a Cross-modal Global Context Modulation Module (CGCMM) is proposed to exploit semantic priors from language instructions for adaptive viewpoint correction and background suppression during the early stage of visual feature extraction. Second, a Word-Pixel Cross-modal Alignment Module (WPCAM) is designed to achieve fine-grained semantic decoupling through a flattening-based cross-modal attention mechanism, thereby improving semantic understanding in complex dynamic scenes. Finally, a unified decoder jointly predicts object locations and optimal grasp poses from the fused multimodal features.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on the reconstructed dual-view cross-scene dataset containing bottom-view and top-view scenes and on a real-world robotic grasping platform. Comparative results demonstrate that SGL-Net consistently outperforms mainstream CNN-based and CLIP-based methods in both grasp detection and object localization (Tables 2 and 3). Ablation studies further verify the effectiveness of CGCMM and WPCAM in improving fine-grained semantic alignment and semantic decoupling (Tables 4 and 5). Furthermore, qualitative results (Figs. 46) and real-world robotic experiments (Fig. 7) demonstrate that SGL-Net can be reliably deployed in complex physical environments. Overall, the proposed network exhibits strong generalization capability and excellent potential for practical robotic applications.  Conclusions  To improve cross-view and cross-scene generalization in language-conditioned robotic grasp detection, this paper constructs a dedicated validation dataset and proposes SGL-Net, which jointly performs grasp detection and object localization. By integrating CGCMM and WPCAM, the proposed network accurately localizes instruction-specified objects and predicts optimal grasp poses. Experimental results obtained on multiple benchmark scenarios and a real-world robotic platform demonstrate the superior performance and practical applicability of the proposed method. Future work will focus on integrating Large Multimodal Models (LMMs) and adapting the proposed framework through fine-tuning to further improve zero-shot robotic grasp detection.
  • loading
  • [1]
    张春云, 孟昕曈, 陶陶, 等. 面向机器人螺栓装配的视觉感知与力控协同方法[J]. 电子与信息学报, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.

    ZHANG Chunyun, MENG Xintong, TAO Tao, et al. Vision-guided and force-controlled method for robotic screw assembly[J]. Journal of Electronics & Information Technology, 2026, 48(5): 2053–2065. doi: 10.11999/JEIT251193.
    [2]
    余浩扬, 李艳生, 肖凌励, 等. 面向动态环境的巡检机器人轻量级语义视觉SLAM框架[J]. 电子与信息学报, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.

    YU Haoyang, LI Yansheng, XIAO Lingli, et al. A lightweight semantic visual simultaneous localization and mapping framework for inspection robots in dynamic environments[J]. Journal of Electronics & Information Technology, 2025, 47(10): 3979–3992. doi: 10.11999/JEIT250301.
    [3]
    LI Renjie, DONG Wei, SUN Jiarui, et al. A fast integrated gait, footstep, and motion planning framework for wheeled-legged robots[J]. Journal of Bionic Engineering, 2026, 23(2): 607–621. doi: 10.1007/s42235-025-00830-5.
    [4]
    LIU Zhitai, ZHONG Hanhai, LIN Xiaotian, et al. Integrated motion control framework of wheel-legged biped robot on rugged terrain[J]. IEEE/ASME Transactions on Mechatronics, 2025, 30(6): 7383–7394. doi: 10.1109/TMECH.2025.3627438.
    [5]
    崔永成, 田国会, 周昭旭, 等. 智能空间下面向动作序列生成的服务机器人指令解析方法[J]. 机器人, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.

    CUI Yongcheng, TIAN Guohui, ZHOU Zhaoxu, et al. A service robot instruction parsing method for action sequence generation in intelligent space[J]. Robot, 2024, 46(1): 1–15. doi: 10.13973/j.cnki.robot.230074.
    [6]
    LENZ I, LEE H, and SAXENA A. Deep learning for detecting robotic grasps[J]. The International Journal of Robotics Research, 2015, 34(4/5): 705–724. doi: 10.1177/0278364914549607.
    [7]
    KUMRA S and KANAN C. Robotic grasp detection using deep convolutional neural networks[C]. 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vancouver, Canada, 2017: 769–776. doi: 10.1109/IROS.2017.8202237.
    [8]
    LAILI Yuanjun, CHEN Zelin, REN Lei, et al. Custom grasping: A region-based robotic grasping detection method in industrial cyber-physical systems[J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(1): 88–100. doi: 10.1109/TASE.2021.3139610.
    [9]
    CHU F J, XU Ruinian, and VELA P A. Real-world multiobject, multigrasp detection[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3355–3362. doi: 10.1109/LRA.2018.2852777.
    [10]
    BOUSSELHAM W, PETERSEN F, FERRARI V, et al. Grounding everything: Emerging localization properties in vision-language transformers[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 3828–3837. doi: 10.1109/CVPR52733.2024.00367.
    [11]
    YIN Heng, REN Yuqiang, YAN Ke, et al. ROD-MLLM: Towards more reliable object detection in multimodal large language models[C]. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, USA, 2025: 14358–14368. doi: 10.1109/CVPR52734.2025.01339.
    [12]
    ZHOU Yi, SU Hang, WANG Tian, et al. Onet: Twin U-Net architecture for unsupervised binary semantic segmentation in radar and remote sensing images[J]. IEEE Transactions on Image Processing, 2025, 34: 2161–2172. doi: 10.1109/TIP.2025.3530816.
    [13]
    WANG Hao, HU Keyan, GUO Xin, et al. A gift from the integration of discriminative and diffusion-based generative learning: Boundary refinement remote sensing semantic segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5892–5909. doi: 10.1109/TPAMI.2026.3654243.
    [14]
    SHRIDHAR M, MANUELLI L, and FOX D. CLIPort: What and where pathways for robotic manipulation[C]. Proceedings of the 5th Conference on Robot Learning, London, UK, 2021: 894–906.
    [15]
    LIU Jin, XIE Jialong, and XIAO Leibing. Hierarchical multi-modal fusion for language-conditioned robotic grasping detection in clutter[J]. IEEE Robotics and Automation Letters, 2024, 9(10): 8762–8769. doi: 10.1109/LRA.2024.3440833.
    [16]
    VAN VO T, VU M N, HUANG Baoru, et al. Language-driven grasp detection with mask-guided attention[C]. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems, Abu Dhabi, UAE, 2024: 7492–7498. doi: 10.1109/IROS58592.2024.10802256.
    [17]
    XIE Jialong, LIU Jin, ZHU Zhenwei, et al. Infusing multisource heterogeneous knowledge for language-conditioned segmentation and grasping[J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 5029611. doi: 10.1109/TIM.2024.3446625.
    [18]
    XIE Jialong, ZHOU Fengyu, LIU Jin, et al. Semi-supervised language-conditioned grasping with curriculum-scheduled augmentation and geometric consistency[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 4021–4028. doi: 10.1109/LRA.2025.3547619.
    [19]
    张梅, 金叶, 朱金辉, 等. 特征级语义感知引导的多模态图像融合算法[J]. 电子与信息学报, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.

    ZHANG Mei, JIN Ye, ZHU Jinhui, et al. FSG: Feature-level semantic-aware guidance for multi-modal image fusion algorithm[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2909–2918. doi: 10.11999/JEIT250042.
    [20]
    LI Jindong, LI Yongguang, FU Yali, et al. CLIP-powered domain generalization and domain adaptation: A comprehensive survey[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026, 48(5): 5405–5424. doi: 10.1109/TPAMI.2026.3651700.
    [21]
    ZHOU Zhenning, ZHU Xiaoxiao, and CAO Qixin. AAGDN: Attention-augmented grasp detection network based on coordinate attention and effective feature fusion method[J]. IEEE Robotics and Automation Letters, 2023, 8(6): 3462–3469. doi: 10.1109/LRA.2023.3268596.
    [22]
    WANG Dexin, LIU Chunsheng, CHANG Faliang, et al. High-performance pixel-level grasp detection based on adaptive grasping and grasp-aware network[J]. IEEE Transactions on Industrial Electronics, 2021, 69(11): 11611–11621. doi: 10.1109/TIE.2021.3120474.
    [23]
    REZATOFIGHI H, TSOI N, GWAK J, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 658–666. doi: 10.1109/CVPR.2019.00075.
    [24]
    KUMRA S, JOSHI S, and SAHIN F. Antipodal robotic grasping using generative residual convolutional neural network[C]. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, 2020: 9626–9633. doi: 10.1109/IROS45743.2020.9340777.
    [25]
    XIE Jialong, LIU Jin, HUANG Saike, et al. Listen, perceive, grasp: CLIP-driven attribute-aware network for language-conditioned visual segmentation and grasping[J]. IEEE Transactions on Automation Science and Engineering, 2025, 22: 9729–9740. doi: 10.1109/TASE.2024.3510777.
    [26]
    MORRISON D, CORKE P, and LEITNER J. Learning robust, real-time, reactive robotic grasping[J]. The International Journal of Robotics Research, 2020, 39(2/3): 183–201. doi: 10.1177/0278364919859066.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(7)  / Tables(6)

    Article Metrics

    Article views (221) PDF downloads(11) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return