A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
-
摘要: 在港口散货装卸作业中,抓斗作为门座式起重机的关键执行部件,其连续轨迹提取易受尺度变化、局部遮挡、边界退化和帧间形态波动等因素影响。为此,该文提出一种面向抓斗轨迹稳定提取的空间与时序协同优化方法。该方法从空间表征增强与时序一致性约束两个层面展开:前者通过多尺度特征交互和边界细节增强,提高复杂背景与遮挡条件下抓斗区域的分割完整性;后者在分割表征层引入时序一致性约束,减弱局部误分、边界波动和短时漏检对轨迹连续性的影响,并通过联合优化目标实现端到端训练。实验结果表明,该方法在DAVIS2016, SegTrackV2和实际门座式起重机抓斗数据集上均表现出较优综合性能,其中在DAVIS2016上区域相似度J和轮廓精度F分别达到74.79%和74.70%,在实际抓斗数据集上平均绝对误差MAE和均方根误差RMSE较YOLOv8-seg分别下降约55.3%和52.6%,漏检率Miss rate由2.95%降至0.56%。遮挡实验进一步表明,该方法能够减弱轨迹抖动与短时中断,在复杂港口作业场景下具有较强的连续目标提取能力。同时,改进后模型推理速度保持在33fps以上,具备实时应用潜力。Abstract:
Objective In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is essential for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often reduce the stability of target region segmentation, leading to centroid deviation, trajectory jitter, missed detections, and trajectory discontinuity. To address these challenges, a Spatial-Temporal Collaborative Optimization Method is proposed for stable and continuous grab trajectory extraction. While maintaining high inference speed, the proposed method improves both trajectory extraction accuracy and trajectory stability, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios. Methods Built on YOLOv8-seg, the proposed framework integrates Spatial Representation Enhancement (SRE) and Temporal CONSistency constraint (TCONS). First, CBAM, BiFPN-lite, and shallow feature aggregation are incorporated to improve target-background separability, enhance multi-scale feature representation, and preserve boundary details. TCONS is then imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss to suppress the temporal accumulation of local errors. In addition, a stage-wise training strategy with warm-up epochs and a joint optimization objective is adopted to ensure stable convergence. Results and Discussions Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The proposed method achieves the best segmentation performance on the real portal crane grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also improves performance on DAVIS2016 while maintaining comparable performance with slight gains on SegTrackV2 ( Tables 1 and2 ,Fig. 2 ). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, and decreases the miss rate to 0.56% (Table 5 ). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3 ). Occlusion robustness experiments further demonstrate that, under different occlusion ratios, the proposed method maintains good region integrity and continuous target extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6 ,Figs. 4 and5 ). Ablation studies verify the complementary effects of SRE and TCONS, whereas parameter analysis shows that a TCONS weight of 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and8 ).Conclusions A Spatial-Temporal Collaborative Optimization Method is proposed to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves high segmentation accuracy and stable trajectory extraction on DAVIS2016 and the real portal crane grab dataset, while maintaining comparable segmentation performance on SegTrackV2. It also exhibits strong continuous target extraction capability under occlusion without significantly sacrificing inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further improve the robustness and applicability of the proposed method in real-world environments. -
表 1 抓斗数据集分割性能对比实验结果(%)
方法 J F Mask mAP50-95 YOLOv8-seg 89.45 97.83 88.24 YOLO11-seg 84.63 93.61 85.85 YOLO12-seg 84.84 94.14 85.73 YOLO26-seg 78.58 87.22 86.01 本文方法 90.05 98.56 88.54 表 2 公开数据集分割性能对比实验结果(%)
方法 DAVIS2016数据集 SegTrackV2数据集 J F Mask mAP50-95 J F Mask mAP50-95 YOLOv8-seg 69.38 69.05 47.54 47.53 53.44 82.23 YOLO11-seg 64.40 62.51 28.39 45.58 51.30 76.14 YOLO12-seg 69.85 68.07 31.16 45.16 50.11 76.37 YOLO26-seg 64.80 63.84 27.74 41.51 46.82 73.86 本文方法 74.79 74.70 51.21 47.65 53.48 82.10 表 3 模型复杂度比较
方法 Params(M) FPS 推理时间(ms) YOLOv8-seg 11.7905 35.80 27.93 YOLO11-seg 2.8428 42.78 23.38 YOLO12-seg 2.8210 37.09 26.96 YOLO26-seg 3.0531 35.91 27.84 本文方法 11.8378 33.24 30.08 表 4 DAVIS2016数据集轨迹稳定性对比实验结果
方法 MAE
(pixel)RMSE
(pixel)Mean jitter
(pixel)Miss rate
(%)YOLOv8-seg 38.43 56.58 37.59 7.87 YOLO11-seg 61.75 82.17 46.96 6.54 YOLO12-seg 37.38 55.03 31.93 7.73 YOLO26-seg 50.49 77.23 57.97 9.58 本文方法 32.11 49.26 31.04 4.05 表 5 抓斗数据集轨迹稳定性对比实验结果
方法 MAE
(pixel)RMSE
(pixel)Mean jitter
(pixel)Miss rate
(%)YOLOv8-seg 6.74 7.13 1.66 2.95 YOLO11-seg 7.20 7.88 3.82 13.74 YOLO12-seg 14.81 16.04 3.02 7.73 YOLO26-seg 9.50 11.42 8.99 18.19 本文方法 3.01 3.38 1.52 0.56 表 6 抓斗数据集遮挡鲁棒性总体对比实验结果
方法 J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) Max consecutive miss(frames) YOLOv8-seg 61.67 70.22 36.38 156.41 22.95 52 YOLO11-seg 43.80 51.34 48.26 133.65 37.00 85 YOLO12-seg 48.40 56.59 74.97 185.75 26.10 56 YOLO26-seg 57.10 64.87 32.19 102.42 28.02 87 本文方法 63.51 72.42 33.28 102.38 17.54 47 表 7 抓斗数据集消融实验结果
Baseline SRE TCONS J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) Mask mAP50-95(%) √ 89.45 97.83 6.74 7.13 2.95 88.24 √ √ 89.49 97.81 5.33 5.43 3.67 88.40 √ √ 89.39 97.78 5.33 5.76 2.89 88.43 √ √ √ 90.05 98.56 3.01 3.38 0.56 88.54 表 8 TCONS关键参数敏感性实验结果
参数 取值 J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) $ {\lambda }_{\text{tcons}} $ 0.2 89.53 97.80 3.02 3.10 1.11 0.3 90.05 98.56 3.01 3.38 0.56 0.4 89.91 98.07 4.17 4.35 0.67 $ \gamma $ 0.6 89.90 98.22 3.16 3.53 0.89 0.7 90.05 98.56 3.01 3.38 0.56 0.8 89.70 97.95 3.06 3.43 1.00 $ {E}_{\text{w}} $ 0 89.39 97.79 3.02 3.39 1.34 5 90.05 98.56 3.01 3.38 0.56 10 89.00 97.59 3.82 4.19 1.22 $ G $ 30 90.08 98.41 3.75 4.12 0.56 50 90.05 98.56 3.01 3.38 0.56 70 90.10 98.44 3.13 3.50 0.56 -
[1] HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN[C]. The 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2980–2988. doi: 10.1109/ICCV.2017.322. [2] CAELLES S, MANINIS K K, PONT-TUSET J, et al. One-shot video object segmentation[C]. The 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 5320–5329. doi: 10.1109/CVPR.2017.565. [3] OH S W, LEE J Y, XU Ning, et al. Video object segmentation using space-time memory networks[C]. The 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), 2019: 9225–9234. doi: 10.1109/ICCV.2019.00932. [4] CHENG H K and SCHWING A G. XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model[C]. 17th European Conference on Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 640–658. doi: 10.1007/978-3-031-19815-1_37. [5] MONNIER Q, POULI T, and KPALMA K. Survey on fast dense video segmentation techniques[J]. Computer Vision and Image Understanding, 2024, 241: 103959. doi: 10.1016/j.cviu.2024.103959. [6] XU Guoping, UDUPA J K, YU Yajun, et al. Segment anything for video: A comprehensive review of video object segmentation and tracking from past to future[J]. Neurocomputing, 2026, 682: 133439. doi: 10.1016/j.neucom.2026.133439. [7] HOU Zhiqiang, LI Fucheng, DONG Jiale, et al. Video object segmentation based on dynamic perception update and feature fusion[J]. Image and Vision Computing, 2024, 150: 105156. doi: 10.1016/j.imavis.2024.105156. [8] WANG Jingxin, ZHANG Yunfeng, BAO Fangxun, et al. Video object segmentation by multi-scale attention using bidirectional strategy[J]. Image and Vision Computing, 2024, 148: 105136. doi: 10.1016/j.imavis.2024.105136. [9] KIM J, KIM J, and HONG S. G-TRACE: Grouped temporal recalibration for video object segmentation[J]. Image and Vision Computing, 2024, 147: 105050. doi: 10.1016/j.imavis.2024.105050. [10] HOU Zhiqiang, WANG Chenxu, MA Sugang, et al. Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation[J]. Knowledge-Based Systems, 2025, 308: 112759. doi: 10.1016/j.knosys.2024.112759. [11] LIU Yong, YU Ran, YIN Fei, et al. Learning high-quality dynamic memory for video object segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3452–3468. doi: 10.1109/TPAMI.2025.3532306. [12] SUN Maojin and SUN Minghui. STSim-Mamb: A spatiotemporal similarity learning framework for unsupervised video object segmentation[J]. Image and Vision Computing, 2026, 169: 105945. doi: 10.1016/j.imavis.2026.105945. [13] KAZEMI E S, TOUBAL I E, RAHMON G, et al. Domain generalization for multiple video object segmentation and tracking using transformers and smart memory[J]. International Journal of Computer Vision, 2026, 134(5): 206. doi: 10.1007/s11263-026-02742-1. [14] LU Hannan, TIAN Zhi, WEI Pengxu, et al. Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation[J]. Neurocomputing, 2024, 602: 127878. doi: 10.1016/j.neucom.2024.127878. [15] 侯志强, 董佳乐, 马素刚, 等. 基于多尺度特征增强与全局-局部特征聚合的视频目标分割算法[J]. 电子与信息学报, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.HOU Zhiqiang, DONG Jiale, MA Sugang, et al. Video object segmentation algorithm based on multi-scale feature enhancement and global-local feature aggregation[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394. [16] 陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on Transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735. [17] WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision – ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [18] LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. The 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 936–944. doi: 10.1109/CVPR.2017.106. [19] LIU Shu, QI Lu, QIN Haifang, et al. Path aggregation network for instance segmentation[C]. The 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 8759–8768. doi: 10.1109/CVPR.2018.00913. [20] 丁建睿, 张听, 刘家栋, 等. 融合邻域注意力和状态空间模型的医学视频分割算法[J]. 电子与信息学报, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.DING Jianrui, ZHANG Ting, LIU Jiadong, et al. Medical video segmentation algorithm integrating neighborhood attention and state space model[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755. [21] LI Jun, SUN Lijuan, REN Hengyi, et al. Learning effective feature representation for video object segmentation via memory[J]. Knowledge-Based Systems, 2024, 299: 112020. doi: 10.1016/j.knosys.2024.112020. [22] WANG Hui, ZHAO Yuqian, ZHANG Fan, et al. Multi-scale spatio-temporal memory network for semi-supervised video object segmentation[J]. Neurocomputing, 2025, 642: 130487. doi: 10.1016/j.neucom.2025.130487. [23] MIAO Bo, BENNAMOUN M, GAO Yongsheng, et al. Temporally consistent referring video object segmentation with hybrid memory[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(11): 11373–11385. doi: 10.1109/TCSVT.2024.3419119. [24] HOU Zhiqiang, CUI Hao, WANG Chenxu, et al. Frequency-aware fusion for improved video object segmentation[J]. Neurocomputing, 2025, 656: 131585. doi: 10.1016/j.neucom.2025.131585. [25] PERAZZI F, PONT-TUSET J, MCWILLIAMS B, et al. A benchmark dataset and evaluation methodology for video object segmentation[C]. The 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 724–732. doi: 10.1109/CVPR.2016.85. [26] LI Fuxin, KIM T, HUMAYUN A, et al. Video segmentation by tracking many figure-ground segments[C]. The 2013 IEEE International Conference on Computer Vision, Sydney, Australia, 2013: 2192–2199. doi: 10.1109/ICCV.2013.273. -
下载: