A Spatial-Temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
-
摘要: 在港口散货装卸作业中,抓斗作为门座式起重机的关键执行部件,其连续轨迹提取易受尺度变化、局部遮挡、边界退化和帧间形态波动等因素影响。为此,该文提出一种面向抓斗轨迹稳定提取的空间与时序协同优化方法。该方法从空间表征增强与时序一致性约束两个层面展开:前者通过多尺度特征交互和边界细节增强,提高复杂背景与遮挡条件下抓斗区域的分割完整性;后者在分割表征层引入时序一致性约束,减弱局部误分、边界波动和短时漏检对轨迹连续性的影响,并通过联合优化目标实现端到端训练。实验结果表明,该方法在DAVIS2016、SegTrackV2和实际门座式起重机抓斗数据集上均表现出较优综合性能,其中在DAVIS2016上区域相似度J和轮廓精度F分别达到74.79%和74.70%,在实际抓斗数据集上平均绝对误差MAE和均方根误差RMSE较YOLOv8-seg分别下降约55.3%和52.6%,漏检率Miss rate由2.95%降至0.56%。遮挡实验进一步表明,该方法能够减弱轨迹抖动与短时中断,在复杂港口作业场景下具有较强的连续目标提取能力。同时,改进后模型推理速度保持在33fps以上,具备实时应用潜力。Abstract:
Objective In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is of great importance for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often make target region extraction unstable, leading to centroid deviation, trajectory jitter, missed detections, and trajectory interruption. To address these issues, this paper proposes a spatial-temporal collaborative optimization method for stable and continuous grab trajectory extraction. While maintaining a relatively high inference speed, the proposed method effectively improves both the accuracy and continuity of trajectory extraction, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios. Methods Built on YOLOv8-seg, the proposed framework integrates spatial representation enhancement and temporal consistency constraint. First, CBAM, BiFPN-lite, and shallow feature aggregation are introduced to improve target-background separability, strengthen multi-scale representation, and preserve boundary details. Then, a temporal consistency constraint is imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss, thereby suppressing the temporal accumulation of local errors. In addition, a stage-wise training strategy with warmup epochs and a unified loss function is adopted to ensure stable convergence. Results and Discussions Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The results show that the proposed method achieves the best segmentation performance on the grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also shows clear improvement on DAVIS2016 and maintains comparable performance with slight gains on SegTrackV2 ( Tables 1 and2 ,Fig. 2 ). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, while lowering the miss rate to 0.56% (Table 5 ). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3 ). Occlusion robustness experiments further show that, under different occlusion ratios, the proposed method maintains good region integrity and continuous extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6 ,Figs. 4 and5 ). Ablation studies verify the complementarity of spatial representation enhancement and the temporal consistency constraint, while parameter analysis shows that setting the TCONS weight to 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and8 ).Conclusions This paper proposes a spatial-temporal collaborative optimization method to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves favorable segmentation accuracy and trajectory stability on DAVIS2016 and the real grab dataset, maintains comparable segmentation performance on SegTrackV2, shows strong continuous extraction capability under occlusion, and incurs no significant loss in inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further enhance the robustness and applicability of the proposed method in real-world environments. -
表 1 抓斗数据集分割性能对比实验结果(%)
方法 J F Mask mAP50-95 YOLOv8-seg 89.45 97.83 88.24 YOLO11-seg 84.63 93.61 85.85 YOLO12-seg 84.84 94.14 85.73 YOLO26-seg 78.58 87.22 86.01 本文方法 90.05 98.56 88.54 表 2 公开数据集分割性能对比实验结果(%)
方法 DAVIS2016数据集 SegTrackV2数据集 J F Mask mAP50-95 J F Mask mAP50-95 YOLOv8-seg 69.38 69.05 47.54 47.53 53.44 82.23 YOLO11-seg 64.40 62.51 28.39 45.58 51.30 76.14 YOLO12-seg 69.85 68.07 31.16 45.16 50.11 76.37 YOLO26-seg 64.80 63.84 27.74 41.51 46.82 73.86 本文方法 74.79 74.70 51.21 47.65 53.48 82.10 表 3 模型复杂度比较
方法 Params(M) FPS Inference Time(ms) YOLOv8-seg 11.7905 35.80 27.93 YOLO11-seg 2.8428 42.78 23.38 YOLO12-seg 2.8210 37.09 26.96 YOLO26-seg 3.0531 35.91 27.84 本文方法 11.8378 33.24 30.08 表 4 DAVIS2016数据集轨迹稳定性对比实验结果
方法 MAE
(pixel)RMSE
(pixel)Mean jitter
(pixel)Miss rate
(%)YOLOv8-seg 38.43 56.58 37.59 7.87 YOLO11-seg 61.75 82.17 46.96 6.54 YOLO12-seg 37.38 55.03 31.93 7.73 YOLO26-seg 50.49 77.23 57.97 9.58 本文方法 32.11 49.26 31.04 4.05 表 5 抓斗数据集轨迹稳定性对比实验结果
方法 MAE
(pixel)RMSE
(pixel)Mean jitter
(pixel)Miss rate
(%)YOLOv8-seg 6.74 7.13 1.66 2.95 YOLO11-seg 7.20 7.88 3.82 13.74 YOLO12-seg 14.81 16.04 3.02 7.73 YOLO26-seg 9.50 11.42 8.99 18.19 本文方法 3.01 3.38 1.52 0.56 表 6 抓斗数据集遮挡鲁棒性总体对比实验结果
方法 J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) Max consecutive miss (frames) YOLOv8-seg 61.67 70.22 36.38 156.41 22.95 52 YOLO11-seg 43.80 51.34 48.26 133.65 37.00 85 YOLO12-seg 48.40 56.59 74.97 185.75 26.10 56 YOLO26-seg 57.10 64.87 32.19 102.42 28.02 87 本文方法 63.51 72.42 33.28 102.38 17.54 47 表 7 抓斗数据集消融实验结果
Baseline SRE TCONS J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) Mask mAP50-95(%) √ 89.45 97.83 6.74 7.13 2.95 88.24 √ √ 89.49 97.81 5.33 5.43 3.67 88.40 √ √ 89.39 97.78 5.33 5.76 2.89 88.43 √ √ √ 90.05 98.56 3.01 3.38 0.56 88.54 表 8 TCONS关键参数敏感性实验结果
参数 取值 J(%) F(%) MAE(pixel) RMSE(pixel) Miss rate(%) $ {\lambda }_{\text{tcons}} $ 0.2 89.53 97.80 3.02 3.10 1.11 0.3 90.05 98.56 3.01 3.38 0.56 0.4 89.91 98.07 4.17 4.35 0.67 $ \gamma $ 0.6 89.90 98.22 3.16 3.53 0.89 0.7 90.05 98.56 3.01 3.38 0.56 0.8 89.70 97.95 3.06 3.43 1.00 $ {E}_{\text{w}} $ 0 89.39 97.79 3.02 3.39 1.34 5 90.05 98.56 3.01 3.38 0.56 10 89.00 97.59 3.82 4.19 1.22 $ G $ 30 90.08 98.41 3.75 4.12 0.56 50 90.05 98.56 3.01 3.38 0.56 70 90.10 98.44 3.13 3.50 0.56 -
[1] HE Kaiming, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN[C]. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2980–2988. doi: 10.1109/ICCV.2017.322. [2] CAELLES S, MANINIS K K, PONT-TUSET J, et al. One-shot video object segmentation[C]. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 5320–5329. doi: 10.1109/CVPR.2017.565. [3] OH S W, LEE J Y, XU Ning, et al. Video object segmentation using space-time memory networks[C]. Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea (South), 2019: 9225–9234. doi: 10.1109/ICCV.2019.00932. [4] CHENG H K and SCHWING A G. XMem: Long-term video object segmentation with an Atkinson-Shiffrin memory model[C]. 17th European Conference on Computer Vision – ECCV 2022, Tel Aviv, Israel, 2022: 640–658. doi: 10.1007/978-3-031-19815-1_37. [5] MONNIER Q, POULI T, and KPALMA K. Survey on fast dense video segmentation techniques[J]. Computer Vision and Image Understanding, 2024, 241: 103959. doi: 10.1016/j.cviu.2024.103959. [6] XU Guoping, UDUPA J K, YU Yajun, et al. Segment anything for video: A comprehensive review of video object segmentation and tracking from past to future[J]. Neurocomputing, 2026, 682: 133439. doi: 10.1016/j.neucom.2026.133439. [7] HOU Zhiqiang, LI Fucheng, DONG Jiale, et al. Video object segmentation based on dynamic perception update and feature fusion[J]. Image and Vision Computing, 2024, 150: 105156. doi: 10.1016/j.imavis.2024.105156. [8] WANG Jingxin, ZHANG Yunfeng, BAO Fangxun, et al. Video object segmentation by multi-scale attention using bidirectional strategy[J]. Image and Vision Computing, 2024, 148: 105136. doi: 10.1016/j.imavis.2024.105136. [9] KIM J, KIM J, and HONG S. G-TRACE: Grouped temporal recalibration for video object segmentation[J]. Image and Vision Computing, 2024, 147: 105050. doi: 10.1016/j.imavis.2024.105050. [10] HOU Zhiqiang, WANG Chenxu, MA Sugang, et al. Lightweight video object segmentation: Integrating online knowledge distillation for fast segmentation[J]. Knowledge-Based Systems, 2025, 308: 112759. doi: 10.1016/j.knosys.2024.112759. [11] LIU Yong, YU Ran, YIN Fei, et al. Learning high-quality dynamic memory for video object segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(5): 3452–3468. doi: 10.1109/TPAMI.2025.3532306. [12] SUN Maojin and SUN Minghui. STSim-Mamb: A spatiotemporal similarity learning framework for unsupervised video object segmentation[J]. Image and Vision Computing, 2026, 169: 105945. doi: 10.1016/j.imavis.2026.105945. [13] KAZEMI E S, TOUBAL I E, RAHMON G, et al. Domain generalization for multiple video object segmentation and tracking using transformers and smart memory[J]. International Journal of Computer Vision, 2026, 134(5): 206. doi: 10.1007/s11263-026-02742-1. [14] LU Hannan, TIAN Zhi, WEI Pengxu, et al. Integrating instance-level knowledge to see the unseen: A two-stream network for video object segmentation[J]. Neurocomputing, 2024, 602: 127878. doi: 10.1016/j.neucom.2024.127878. [15] 侯志强, 董佳乐, 马素刚, 等. 基于多尺度特征增强与全局-局部特征聚合的视频目标分割算法[J]. 电子与信息学报, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394.HOU Zhiqiang, DONG Jiale, MA Sugang, et al. Video object segmentation algorithm based on multi-scale feature enhancement and global-local feature aggregation[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4198–4207. doi: 10.11999/JEIT231394. [16] 陈雷, 杨吉斌, 曹铁勇, 等. 一种基于Transformer特征金字塔的自蒸馏目标分割方法[J]. 电子与信息学报, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735.CHEN Lei, YANG Jibin, CAO Tieyong, et al. A self-distillation object segmentation method based on Transformer feature pyramid[J]. Journal of Electronics & Information Technology, 2025, 47(2): 551–560. doi: 10.11999/JEIT240735. [17] WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision – ECCV 2018, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [18] LIN T Y, DOLLÁR P, GIRSHICK R, et al. Feature pyramid networks for object detection[C]. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, 2017: 936–944. doi: 10.1109/CVPR.2017.106. [19] LIU Shu, QI Lu, QIN Haifang, et al. Path aggregation network for instance segmentation[C]. Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 8759–8768. doi: 10.1109/CVPR.2018.00913. [20] 丁建睿, 张听, 刘家栋, 等. 融合邻域注意力和状态空间模型的医学视频分割算法[J]. 电子与信息学报, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755.DING Jianrui, ZHANG Ting, LIU Jiadong, et al. Medical video segmentation algorithm integrating neighborhood attention and state space model[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1582–1595. doi: 10.11999/JEIT240755. [21] LI Jun, SUN Lijuan, REN Hengyi, et al. Learning effective feature representation for video object segmentation via memory[J]. Knowledge-Based Systems, 2024, 299: 112020. doi: 10.1016/j.knosys.2024.112020. [22] WANG Hui, ZHAO Yuqian, ZHANG Fan, et al. Multi-scale spatio-temporal memory network for semi-supervised video object segmentation[J]. Neurocomputing, 2025, 642: 130487. doi: 10.1016/j.neucom.2025.130487. [23] MIAO Bo, BENNAMOUN M, GAO Yongsheng, et al. Temporally consistent referring video object segmentation with hybrid memory[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(11): 11373–11385. doi: 10.1109/TCSVT.2024.3419119. [24] HOU Zhiqiang, CUI Hao, WANG Chenxu, et al. Frequency-aware fusion for improved video object segmentation[J]. Neurocomputing, 2025, 656: 131585. doi: 10.1016/j.neucom.2025.131585. [25] PERAZZI F, PONT-TUSET J, MCWILLIAMS B, et al. A benchmark dataset and evaluation methodology for video object segmentation[C]. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 724–732. doi: 10.1109/CVPR.2016.85. [26] LI Fuxin, KIM T, HUMAYUN A, et al. Video segmentation by tracking many figure-ground segments[C]. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, Australia, 2013: 2192–2199. doi: 10.1109/ICCV.2013.273. -
下载: