Advanced Search
Turn off MathJax
Article Contents
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming. A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260512
Citation: CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming. A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260512

A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction

doi: 10.11999/JEIT260512 cstr: 32379.14.JEIT260512
Funds:  The Natural Science Foundation of Hebei Province (F2025203055)
  • Received Date: 2026-04-27
  • Accepted Date: 2026-07-01
  • Rev Recd Date: 2026-07-01
  • Available Online: 2026-07-23
  •   Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is essential for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often reduce the stability of target region segmentation, leading to centroid deviation, trajectory jitter, missed detections, and trajectory discontinuity. To address these challenges, a Spatial-Temporal Collaborative Optimization Method is proposed for stable and continuous grab trajectory extraction. While maintaining high inference speed, the proposed method improves both trajectory extraction accuracy and trajectory stability, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates Spatial Representation Enhancement (SRE) and Temporal CONSistency constraint (TCONS). First, CBAM, BiFPN-lite, and shallow feature aggregation are incorporated to improve target-background separability, enhance multi-scale feature representation, and preserve boundary details. TCONS is then imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss to suppress the temporal accumulation of local errors. In addition, a stage-wise training strategy with warm-up epochs and a joint optimization objective is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The proposed method achieves the best segmentation performance on the real portal crane grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also improves performance on DAVIS2016 while maintaining comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, and decreases the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further demonstrate that, under different occlusion ratios, the proposed method maintains good region integrity and continuous target extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementary effects of SRE and TCONS, whereas parameter analysis shows that a TCONS weight of 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  A Spatial-Temporal Collaborative Optimization Method is proposed to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves high segmentation accuracy and stable trajectory extraction on DAVIS2016 and the real portal crane grab dataset, while maintaining comparable segmentation performance on SegTrackV2. It also exhibits strong continuous target extraction capability under occlusion without significantly sacrificing inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further improve the robustness and applicability of the proposed method in real-world environments.
  • loading
  • [1]
    HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN[C]. The 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2980–2988. doi: 10.1109/ICCV.2017.322.
    [2]
    CAELLES S, MANINIS K K, PONT-TUSET J, et al. One-shot video object segmentation[C]. The 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 5320–5329. doi: 10.1109/CVPR.2017.565.
    [3]
    OH S W, LEE J Y, XU Ning, et al. Video object segmentation using space-time memory networks[C]. The 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), 2019: 9225–9234. doi: 10.1109/ICCV.2019.00932.
    [4]
    CHENG H K and SCHWING A G. XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model[C]. 17th European Conference on Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 640–658. doi: 10.1007/978-3-031-19815-1_37.
    [5]
    MONNIER Q, POULI T, and KPALMA K. Survey on fast dense video segmentation techniques[J]. Computer Vision and Image Understanding, 2024, 241: 103959. doi: 10.1016/j.cviu.2024.103959.
    [6]
    XU Guoping, UDUPA J K, YU Yajun, et al. Segment anything for video: A comprehensive review of video object segmentation and tracking from past to future[J]. Neurocomputing, 2026, 682: 133439. doi: 10.1016/j.neucom.2026.133439.
    [7]
    HOU Zhiqiang, LI Fucheng, DONG Jiale, et al. Video object segmentation based on dynamic perception update and feature fusion[J]. Image and Vision Computing, 2024, 150: 105156. doi: 10.1016/j.imavis.2024.105156.
    [8]
    WANG Jingxin, ZHANG Yunfeng, BAO Fangxun, et al. Video object segmentation by multi-scale attention using bidirectional strategy[J]. Image and Vision Computing, 2024, 148: 105136. doi: 10.1016/j.imavis.2024.105136.
    [9]
    KIM J, KIM J, and HONG S. G-TRACE: Grouped temporal recalibration for video object segmentation[J]. Image and Vision Computing, 2024, 147: 105050. doi: 10.1016/j.imavis.2024.105050.
    [10]
    HOU Zhiqiang, WANG Chenxu, MA Sugang, et al. Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation[J]. Knowledge-Based Systems, 2025, 308: 112759. doi: 10.1016/j.knosys.2024.112759.
    [11]
    LIU Yong, YU Ran, YIN Fei, et al. Learning high-quality dynamic memory for video object segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3452–3468. doi: 10.1109/TPAMI.2025.3532306.
    [12]
    SUN Maojin and SUN Minghui. STSim-Mamb: A spatiotemporal similarity learning framework for unsupervised video object segmentation[J]. Image and Vision Computing, 2026, 169: 105945. doi: 10.1016/j.imavis.2026.105945.
    [13]
    KAZEMI E S, TOUBAL I E, RAHMON G, et al. Domain generalization for multiple video object segmentation and tracking using transformers and smart memory[J]. International Journal of Computer Vision, 2026, 134(5): 206. doi: 10.1007/s11263-026-02742-1.
    [14]
    LU Hannan, TIAN Zhi, WEI Pengxu, et al. Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation[J]. Neurocomputing, 2024, 602: 127878. doi: 10.1016/j.neucom.2024.127878.
    [15]
    侯志强, 董佳乐, 马素刚, 等. 基于多尺度特征增强与全局-局部特征聚合的视频目标分割算法[J]. 电子与信息学报, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.

    HOU Zhiqiang, DONG Jiale, MA Sugang, et al. Video object segmentation algorithm based on multi-scale feature enhancement and global-local feature aggregation[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.
    [16]
    陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.

    CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on Transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.
    [17]
    WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision – ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1.
    [18]
    LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. The 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 936–944. doi: 10.1109/CVPR.2017.106.
    [19]
    LIU Shu, QI Lu, QIN Haifang, et al. Path aggregation network for instance segmentation[C]. The 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 8759–8768. doi: 10.1109/CVPR.2018.00913.
    [20]
    丁建睿, 张听, 刘家栋, 等. 融合邻域注意力和状态空间模型的医学视频分割算法[J]. 电子与信息学报, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.

    DING Jianrui, ZHANG Ting, LIU Jiadong, et al. Medical video segmentation algorithm integrating neighborhood attention and state space model[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.
    [21]
    LI Jun, SUN Lijuan, REN Hengyi, et al. Learning effective feature representation for video object segmentation via memory[J]. Knowledge-Based Systems, 2024, 299: 112020. doi: 10.1016/j.knosys.2024.112020.
    [22]
    WANG Hui, ZHAO Yuqian, ZHANG Fan, et al. Multi-scale spatio-temporal memory network for semi-supervised video object segmentation[J]. Neurocomputing, 2025, 642: 130487. doi: 10.1016/j.neucom.2025.130487.
    [23]
    MIAO Bo, BENNAMOUN M, GAO Yongsheng, et al. Temporally consistent referring video object segmentation with hybrid memory[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(11): 11373–11385. doi: 10.1109/TCSVT.2024.3419119.
    [24]
    HOU Zhiqiang, CUI Hao, WANG Chenxu, et al. Frequency-aware fusion for improved video object segmentation[J]. Neurocomputing, 2025, 656: 131585. doi: 10.1016/j.neucom.2025.131585.
    [25]
    PERAZZI F, PONT-TUSET J, MCWILLIAMS B, et al. A benchmark dataset and evaluation methodology for video object segmentation[C]. The 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 724–732. doi: 10.1109/CVPR.2016.85.
    [26]
    LI Fuxin, KIM T, HUMAYUN A, et al. Video segmentation by tracking many figure-ground segments[C]. The 2013 IEEE International Conference on Computer Vision, Sydney, Australia, 2013: 2192–2199. doi: 10.1109/ICCV.2013.273.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(8)

    Article Metrics

    Article views (225) PDF downloads(7) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return