UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos
-
摘要: 视频场景图生成是一项从视频中提取目标及其相互关系的高级视觉理解任务。目前相关研究多聚焦于自然场景,在遥感领域仍有待深入探索。无人机视频视角变化明显且目标尺度差异大,使得数据标注成本显著增加,缺乏大规模基准数据集是现阶段无人机视频场景图生成研究的主要障碍。该文构建了一个无人机视频关系理解数据集,在50个无人机视频中标注了12个目标类别,共计
876687 个目标框;并基于目标框标注了20个关系类别,共计459164 个关系三元组。同时,提出了一种时空超图增强关系理解方法,利用超图结构对无人机视频复杂目标关联进行建模,从而实现更有效的特征聚合,增强了无人机视频关系理解能力。为评估主流方法在该数据集上的性能表现,该文对目标检测任务及场景图生成的三个子任务进行了全面测试与对比。结果表明,所提方法在场景图生成三个子任务共18个精度指标中8个优于基线方法。Abstract:Objective With the rapid development and extensive application of unmanned aerial vehicle (UAV) observation platforms, remote sensing video data with high spatiotemporal resolution has witnessed explosive growth. Understanding dynamic relationships in UAV videos is recognized as a pressing research challenge in intelligent remote sensing analysis. Current research on video scene graph generation is mainly focused on natural videos, while studies on remote sensing videos remain in the exploratory stage, with a lack of benchmark datasets annotated with high-order dynamic relationships. Traditional visual models are difficult to be directly adapted to remote sensing scenes, which are characterized by dynamic view changes, extreme object scale variations, and dense target distributions. To address these limitations, a UAV video relationship dataset (UAVREL) with high-order dynamic relationship annotations is constructed, a hypergraph-enhanced transformer method tailored to the characteristics of UAV videos is proposed, and a unified benchmark evaluation system for video scene graph generation in the field of UAV remote sensing is established in this study. These efforts promote the transformation of intelligent remote sensing interpretation technology from static object cognition to global dynamic scene understanding. Methods This study is mainly composed of three core parts: dataset construction, model design, and comprehensive experimental validation. First, in the dataset construction stage, the Unmanned Aerial Vehicle Benchmark for Object Detection and Tracking (UAVDT) dataset is selected as the basic data source. A semi-automatic annotation strategy is proposed, integrating object tracking, multi-person collaboration, and consensus verification to mitigate subjective bias and improve efficiency. The relationships are divided into three levels according to the requirements of cross-frame inference, among which high-order relationships are annotated with high priority. Second, A Hypergraph-Enhanced Transformer model (STHG) is proposed in this work. It includes six core functional modules: basic feature extraction, pairwise context encoding, spatial-temporal dual hypergraph enhancement, multi-source feature fusion, global temporal modeling and relation prediction. On this basis, a multi-label classifier generates the final dynamic scene graphs. In the experimental design, to verify the effectiveness of the dataset and the model, an object detection task and three video scene graph generation subtasks, namely predicate classification (PreCls), scene graph classification (SGCls), and scene graph detection (SGDet), are conducted on the UAVREL dataset. Representative object detection models and classic scene graph generation models are selected as baselines. Mean Average Precision (mAP) and mAP@50 are adopted as evaluation metrics for detection, while Recall@K and meanRecall@K (mR@K) are used for scene graph generation to comprehensively evaluate the performance of the model in relation recognition and graph construction. Results and Discussions In the object detection task, the impact of model architectures on detection performance is systematically verified through comparative experiments on eight models with different architectures ( Table 2 ). The results show that the selection of model architecture is of crucial importance to the mean Average Precision (mAP), a core evaluation metric. Single-stage anchor-free models represented by VFNet and DDOD exhibit significant advantages in comprehensive detection performance. From the perspective of category characteristics, all models perform poorly in detecting small-scale and easily occluded target categories, which reflects the common technical challenge faced by current general object detectors in small target detection tasks. In the video scene graph generation task, five methods are tested on three subtasks respectively (Table 3 ,Table 4 ,Table 5 ). The STHG method proposed in this paper shows significant performance advantages in all three core tasks. Meanwhile, experimental data indicate that the value of the average recall metric is consistently significantly lower than that of the traditional recall metric. This phenomenon clearly shows that the dataset poses great modeling challenges in task scenarios with low-frequency object relationships, and implicitly reflects that relationship prediction in such complex scenarios remains a key challenge to be solved urgently in the field of drone video scene graph generation.Conclusions This paper focuses on the critical theme of understanding dynamic relationships in drone videos. It constructs the UAVREL benchmark dataset, providing data support with deeper semantic relationships for this field. And it proposes a Hypergraph-Enhanced Transformer approach for Remote Sensing Videos. Experimental results demonstrate that this approach achieves superior performance across multiple evaluation metrics, thus validating its practical applicability in remote sensing dynamic relationship prediction tasks. Through dataset construction and algorithmic innovation, this paper not only lays a solid data foundation for understanding drone video relationships, but also facilitates a leap from low-level semantic analysis to high-level dynamic semantic cognitive modeling in remote sensing video analysis. Future research will focus on the following two directions: Firstly, deepening the temporal dimension modeling of dynamic remote sensing scene graph generation to enhance the ability to capture long-term evolutionary events and complex relationships; secondly, expanding the scene coverage and diversity of relationship categories in the dataset, continuously improving algorithm benchmarks, and promoting technological iteration and industry application implementation. -
Key words:
- Video scene graph generation /
- Remote sensing /
- Relation comprehension
-
表 1 自然场景与遥感场景中数据集统计对比
数据集名称 视频 帧数 / 图片数 尺寸 目标
标注关系
标注目标
类别目标
数量关系
类别关系
数量年份 自
然
场
景Visual Phrase[14] × 2.8K / √ √ 8 3.3K 17 1.8K 2011 VRD[15] × 5K / √ √ 100 / 70 38.0K 2016 Visual Genome[16] × 108K 72~ 1280 √ √ 33877 3.8M 42374 2.3M 2017 VrR-VG[17] × 59K 72~ 1280 √ √ 1600 282.5K 117 203.4K 2019 VidVRD[18] √ 296.2K 1920× 1080 √ √ 35 / 132 55.6K 2017 VidOR[19] √ 55.4K 640×360 √ √ 80 38.6K 50 297.4K 2019 Action Genome[13] √ 234.3K 1280 ×720√ √ 35 476.2K 25 1.7M 2020 PVSG[20] √ 153K 1920× 1080 √ √ 126 / 57 / 2023 遥
感
场
景DOTA[21] × 2.8K 800~ 4000 √ × 15 188.3K / / 2018 DIOR-R[22] × 23.5K 800×800 √ × 20 192.5K / / 2022 MONET[23] × 53K 800×600 √ × 2 162K / / 2023 ReCon1M[24] × 21K 400~ 10000 √ √ 60 873.8K 59 1.1M 2024 UAVDT[25] √ 80K 1080 ×540√ × 3 841.5K / / 2018 UAVid[26] √ 0.3K 4096 ×2160 √ × 8 / / / 2020 Brutal Running[27] √ 1K 227×227 √ × 1 / / / 2021 UIT-ADrone[28] √ 206.2K 1920× 1080 √ × 8 69.5K / / 2023 AeroEye[29] √ 261.5K 3840 ×2160 √ √ 57 2.2M 384 43M 2024 UAVREL(本文) √ 40K 1024 ×540√ √ 12 884.8K 20 459.7K 2026 表 2 目标检测结果
目标类别 方法 a b c d e f g h 关口 59.4 53.8 55.5 57.3 57.1 54.4 57.8 58.5 非机动车 19.4 19.1 15.4 21.0 19.8 17.7 21.6 21.5 公交车 53.4 52.3 51.4 54.1 53.8 54.0 56.3 54.8 公交车站 45.7 41.9 42.7 45.2 45.2 47.8 45.2 46.5 汽车 50.8 51.3 51.9 55.9 54.6 53.4 57.9 58.1 路口 58.5 59.5 58.8 60.0 56.7 56.1 63.3 62.4 停车点 49.1 50.5 47.5 48.0 46.8 49.6 51.1 50.9 行人 12.2 9.7 9.7 15.7 18.7 13.7 17.2 18.0 道路 55.1 55.0 54.0 56.2 54.5 52.7 57.0 57.4 人行道 46.6 47.6 45.5 44.0 43.0 44.0 49.2 48.9 收费站 57.6 57.3 45.7 48.1 39.3 47.8 54.0 53.5 货车 44.0 44.8 44.4 47.1 36.4 36.5 48.6 47.6 均
值mAP(%) 46.0 45.4 43.5 46.1 43.8 44.0 48.3 48.2 mAP@50(%) 69.3 67.5 68.3 70.8 68.5 68.3 73.0 72.8 注:表中(a)Faster R-CNN、(b)Cascade R-CNN、(c)RetinaNet、(d)ATSS、(e)AutoAssign、(f)FCOS、(g)VFNet、(h)DDOD 表 3 UAVREL数据集的场景图生成方法实验结果R@K(%)
表 4 UAVREL数据集的场景图生成方法实验结果mR@K(%)
方法 PreCls SGCls SGDet IMP[38] Motifs[39] VCTree[40] STTran[41] IMP[38] Motifs[39] VCTree[40] STTran[41] IMP[38] Motifs[39] VCTree[40] STTran[41] 阻碍 52.4 66.8 74.9 65.1 48.7 56.4 58.6 61.6 21.1 63.3 29.5 37.9 并排行驶 67.7 67.7 61.3 68.3 46.8 60.8 61.3 71.7 01.6 61.8 16.7 30.3 驾驶跟随 58.7 60.9 53.2 76.2 49.0 64.9 55.9 65.4 10.5 66.9 9.2 17.4 直行 72.6 82.8 81.6 92.2 72.0 81.9 80.6 89.2 58.5 67.0 66.1 72.3 放行 58.1 91.9 64.5 90.8 48.8 48.8 47.4 49.7 20.9 42.1 37.6 43.8 超车 84.9 90.9 80.1 91.3 74.4 79.6 76.2 81.3 22.6 57.2 41.3 50.6 停车等待 60.1 75.1 81.4 82.4 62.2 74.3 76.3 77.2 43.2 70.8 58.6 62.1 骑行穿过 64.6 97.8 92.8 96.6 70.7 66.0 58.6 68.3 35.9 45.3 37.3 41.8 …… …… …… …… …… ……. …… …… …… …… …… …… …… 避让 4.2 41.7 29.2 33.3 18.8 33.3 27.1 12.5 0.0 0.0 0.0 0.0 汇入 0.0 20.0 10.0 10.0 0.0 5.0 10.0 0.0 0.0 5.0 5.0 0.0 mR@20 42.7 50.8 48.0 46.8 41.5 42.8 43.2 44.4 19.3 32.8 23.2 26.3 mR@50 54.2 68.6 65.6 67.2 51.1 62.0 60.0 62.3 25.1 45.5 36.8 39.4 mR@100 57.7 73.9 70.3 73.6 54.4 66.8 64.1 67.7 28.2 49.9 42.1 45.6 表 5 STHG场景图生成方法实验结果(%)
STHG PredCls SGCls SGDet K=20 K=50 K=100 K=20 K=50 K=100 K=20 K=50 K=100 R@K 78.4 84.9 87.1 74.8 81.9 82.3 53.5 59.1 60.7 mR@K 49.1 66.8 72.5 44.9 62.6 68.7 30.9 43.6 47.8 H@K 60.3 74.8 79.1 56.1 71.0 74.9 39.2 50.2 53.5 -
[1] LIU Fang, WANG Jiahao, JIAO Licheng, et al. Remote sensing video tracking: Current status, challenges, and future[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 14338–14367. doi: 10.1109/JSTARS.2025.3573572. [2] MA Bokun, MU Caihong, LIU Yi, et al. RoSENet: Rotation and similarity enhancement network for multimodal remote sensing image land cover classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5511318. doi: 10.1109/TGRS.2025.3561850. [3] 金晶, 王峰. 分布式多卫星协同遥感图像场景分类方法[J]. 电子与信息学报, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866.JIN Jing and WANG Feng. A distributed multi-satellite collaborative framework for remote sensing scene classification[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866. [4] GAO Feng, JIN Xuepeng, ZHOU Xiaowei, et al. MSFMamba: Multiscale feature fusion state space model for multisource remote sensing image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5504116. doi: 10.1109/TGRS.2025.3535622. [5] ZHANG Yin, YE Mu, ZHU Guiyi, et al. FFCA-YOLO for small object detection in remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5611215. doi: 10.1109/TGRS.2024.3363057. [6] XIAO Yao, XU Tingfa, YU Xin, et al. A lightweight fusion strategy with enhanced interlayer feature correlation for small object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4708011. doi: 10.1109/TGRS.2024.3457155. [7] CHEN Tianxiang, YE Zi, TAN Zhentao, et al. MiM-ISTD: Mamba-in-mamba for efficient infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007613. doi: 10.1109/TGRS.2024.3485721. [8] 姚婷婷, 肇恒鑫, 冯子豪, 等. 上下文感知多感受野融合网络的定向遥感目标检测[J]. 电子与信息学报, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560.YAO Tingting, ZHAO Hengxin, FENG Zihao, et al. A context-aware multiple receptive field fusion network for oriented object detection in remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560. [9] SUN Xian, WANG Peijin, YAN Zhiyuan, et al. FAIR1M: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 184: 116–130. doi: 10.1016/j.isprsjprs.2021.12.004. [10] 周国宇, 张菁, 闫伊, 等. 聚焦注意力与紧致特征融合Transformer的城市遥感影像语义分割[J]. 电子与信息学报, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812.ZHOU Guoyu, ZHANG Jing, YAN Yi, et al. A focused attention and feature compact fusion Transformer for semantic segmentation of urban remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812. [11] CHEN Keyan, LIU Chenyang, CHEN Hao, et al. RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4701117. doi: 10.1109/TGRS.2024.3356074. [12] 于国栋, 蒋一纯, 刘云清, 等. 一种空间语义联合感知的红外无人机目标跟踪方法[J]. 电子与信息学报, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613.YU Guodong, JIANG Yichun, LIU Yunqing, et al. A spatial-semantic combine perception for infrared UAV target tracking[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613. [13] JI Jingwei, KRISHNA R, LI Feifei, et al. Action genome: Actions as compositions of spatio-temporal scene graphs[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 10233–10244. doi: 10.1109/CVPR42600.2020.01025. [14] SADEGHI M A and FARHADI A. Recognition using visual phrases[C]. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recogniti, Colorado Springs, USA, 2011: 1745–1752. doi: 10.1109/CVPR.2011.5995711. [15] LU Cewu, KRISHNA R, BERNSTEIN M, et al. Visual relationship detection with language priors[C]. 14th European Conference Computer Vision-ECCV 2016, Amsterdam, The Netherlands, 2016: 852–869. doi: 10.1007/978-3-319-46448-0_51. [16] KRISHNA R, ZHU Yuke, GROTH O, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations[J]. International Journal of Computer Vision, 2017, 123(1): 32–73. doi: 10.1007/s11263-016-0981-7. [17] LIANG Yuanzhi, BAI Yalong, ZHANG Wei, et al. VrR-VG: Refocusing visually-relevant relationships[C]. 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 2019: 10402–10411. doi: 10.1109/ICCV.2019.01050. [18] SHANG Xindi, REN Tongwei, GUO Jingfan, et al. Video visual relation detection[C]. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, USA, 2017: 1300–1308. doi: 10.1145/3123266.3123380. [19] SHANG Xindi, DI Donglin, XIAO Junbin, et al. Annotating objects and relations in user-generated videos[C]. Proceedings of the 2019 on International Conference on Multimedia Retrieval, Ottawa, Canada, 2019: 279–287. [20] YANG Jingkang, PENG Wenxuan, LI Xiangtai, et al. Panoptic video scene graph generation[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 18675–18685. doi: 10.1109/CVPR52729.2023.01791. [21] XIA Guisong, BAI Xiang, DING Jian, et al. DOTA: A large-scale dataset for object detection in aerial images[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 3974–3983. doi: 10.1109/CVPR.2018.00418. [22] CHENG Gong, WANG Jiabao, LI Ke, et al. Anchor-free oriented proposal generator for object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5625411. doi: 10.1109/TGRS.2022.3183022. [23] RIZ L, CARAFFA A, BORTOLON M, et al. The MONET dataset: Multimodal drone thermal dataset recorded in rural scenarios[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, Canada, 2023: 2546–2554. doi: 10.1109/CVPRW59228.2023.00253. [24] YAN Qiwei, DENG Chubo, LIU Chenglong, et al. ReCon1M: A large-scale benchmark dataset for relation comprehension in remote sensing imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 4507022. doi: 10.1109/TGRS.2025.3589986. [25] DU Dawei, QI Yuankai, YU Hongyang, et al. The unmanned aerial vehicle benchmark: Object detection and tracking[C]. 15th European Conference Computer Vision-ECCV 2018, Munich, Germany, 2018: 370–386. doi: 10.1007/978-3-030-01249-6_23. [26] LYU Ye, VOSSELMAN G, XIA Guisong, et al. UAVid: A semantic segmentation dataset for UAV imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 165: 108–119. doi: 10.1016/j.isprsjprs.2020.05.009. [27] HAMDI S, BOUINDOUR S, SNOUSSI H, et al. End-to-end deep one-class learning for anomaly detection in UAV video stream[J]. Journal of Imaging, 2021, 7(5): 90. doi: 10.3390/jimaging7050090. [28] TRAN T M, VU T N, NGUYEN T V, et al. UIT-ADrone: A novel drone dataset for traffic anomaly detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16: 5590–5601. doi: 10.1109/JSTARS.2023.3285905. [29] NGUYEN T T, NGUYEN P, LI Xin, et al. CYCLO: Cyclic graph transformer approach to multi-object relationship modeling in aerial videos[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 2868. [30] REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137–1149. doi: 10.1109/tpami.2016.2577031. [31] CAI Zhaowei and VASCONCELOS N. Cascade R-CNN: Delving into high quality object detection[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 6154–6162. doi: 10.1109/CVPR.2018.00644. [32] LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]. 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2999–3007. doi: 10.1109/ICCV.2017.324. [33] ZHANG Shifeng, CHI Cheng, YAO Yongqiang, et al. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 9756–9765. doi: 10.1109/CVPR42600.2020.00978. [34] ZHU Benjin, WANG Jianfeng, JIANG Zhengkai, et al. AutoAssign: Differentiable label assignment for dense object detection[EB/OL]. https://arxiv.org/abs/2007.03496, 2020. [35] TIAN Zhi, SHEN Chunhua, CHEN Hao, et al. FCOS: A simple and strong anchor-free object detector[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(4): 1922–1933. doi: 10.1109/tpami.2020.3032166. [36] ZHANG Haoyang, WANG Ying, DAYOUB F, et al. VarifocalNET: An IoU-aware dense object detector[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 8501–8519. doi: 10.1109/CVPR46437.2021.00841. [37] CHEN Zehui, YANG Chenhongyi, LI Qiaofei, et al. Disentangle your dense object detector[C]. Proceedings of the 29th ACM International Conference on Multimedia, 2021: 4939–4948. doi: 10.1145/3474085.3475351. (查阅网上资料,未找到本条文献出版地信息,请确认). [38] TANG Kaihua, NIU Yulei, HUANG Jianqiang, et al. Unbiased scene graph generation from biased training[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 3713–3722. doi: 10.1109/CVPR42600.2020.00377. [39] ZELLERS R, YATSKAR M, THOMSON S, et al. Neural motifs: Scene graph parsing with global context[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 5831–5840. doi: 10.1109/CVPR.2018.00611. [40] TANG Kaihua, ZHANG Hanwang, WU Baoyuan, et al. Learning to compose dynamic tree structures for visual contexts[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 6612–6621. doi: 10.1109/CVPR.2019.00678. [41] CONG Yuren, LIAO Wentong, ACKERMANN H, et al. Spatial-temporal transformer for dynamic scene graph generation[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 16352–16362. doi: 10.1109/ICCV48922.2021.01606. -
下载: