Advanced Search
Turn off MathJax
Article Contents
LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian. UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260221
Citation: LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian. UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260221

UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos

doi: 10.11999/JEIT260221 cstr: 32379.14.JEIT260221
Funds:  National Natural Science Foundation of China (Grant Number: 42301437)
  • Received Date: 2026-03-01
  • Accepted Date: 2026-07-28
  • Rev Recd Date: 2026-07-28
  • Available Online: 2026-09-01
  •   Objective  With the rapid development and extensive application of unmanned aerial vehicle (UAV) observation platforms, remote sensing video data with high spatiotemporal resolution has witnessed explosive growth. Understanding dynamic relationships in UAV videos is recognized as a pressing research challenge in intelligent remote sensing analysis. Current research on video scene graph generation is mainly focused on natural videos, while studies on remote sensing videos remain in the exploratory stage, with a lack of benchmark datasets annotated with high-order dynamic relationships. Traditional visual models are difficult to be directly adapted to remote sensing scenes, which are characterized by dynamic view changes, extreme object scale variations, and dense target distributions. To address these limitations, a UAV video relationship dataset (UAVREL) with high-order dynamic relationship annotations is constructed, a hypergraph-enhanced transformer method tailored to the characteristics of UAV videos is proposed, and a unified benchmark evaluation system for video scene graph generation in the field of UAV remote sensing is established in this study. These efforts promote the transformation of intelligent remote sensing interpretation technology from static object cognition to global dynamic scene understanding.  Methods  This study is mainly composed of three core parts: dataset construction, model design, and comprehensive experimental validation. First, in the dataset construction stage, the Unmanned Aerial Vehicle Benchmark for Object Detection and Tracking (UAVDT) dataset is selected as the basic data source. A semi-automatic annotation strategy is proposed, integrating object tracking, multi-person collaboration, and consensus verification to mitigate subjective bias and improve efficiency. The relationships are divided into three levels according to the requirements of cross-frame inference, among which high-order relationships are annotated with high priority. Second, A Hypergraph-Enhanced Transformer model (STHG) is proposed in this work. It includes six core functional modules: basic feature extraction, pairwise context encoding, spatial-temporal dual hypergraph enhancement, multi-source feature fusion, global temporal modeling and relation prediction. On this basis, a multi-label classifier generates the final dynamic scene graphs. In the experimental design, to verify the effectiveness of the dataset and the model, an object detection task and three video scene graph generation subtasks, namely predicate classification (PreCls), scene graph classification (SGCls), and scene graph detection (SGDet), are conducted on the UAVREL dataset. Representative object detection models and classic scene graph generation models are selected as baselines. Mean Average Precision (mAP) and mAP@50 are adopted as evaluation metrics for detection, while Recall@K and meanRecall@K (mR@K) are used for scene graph generation to comprehensively evaluate the performance of the model in relation recognition and graph construction.  Results and Discussions  In the object detection task, the impact of model architectures on detection performance is systematically verified through comparative experiments on eight models with different architectures (Table 2). The results show that the selection of model architecture is of crucial importance to the mean Average Precision (mAP), a core evaluation metric. Single-stage anchor-free models represented by VFNet and DDOD exhibit significant advantages in comprehensive detection performance. From the perspective of category characteristics, all models perform poorly in detecting small-scale and easily occluded target categories, which reflects the common technical challenge faced by current general object detectors in small target detection tasks. In the video scene graph generation task, five methods are tested on three subtasks respectively (Table 3, Table 4, Table 5). The STHG method proposed in this paper shows significant performance advantages in all three core tasks. Meanwhile, experimental data indicate that the value of the average recall metric is consistently significantly lower than that of the traditional recall metric. This phenomenon clearly shows that the dataset poses great modeling challenges in task scenarios with low-frequency object relationships, and implicitly reflects that relationship prediction in such complex scenarios remains a key challenge to be solved urgently in the field of drone video scene graph generation.  Conclusions  This paper focuses on the critical theme of understanding dynamic relationships in drone videos. It constructs the UAVREL benchmark dataset, providing data support with deeper semantic relationships for this field. And it proposes a Hypergraph-Enhanced Transformer approach for Remote Sensing Videos. Experimental results demonstrate that this approach achieves superior performance across multiple evaluation metrics, thus validating its practical applicability in remote sensing dynamic relationship prediction tasks. Through dataset construction and algorithmic innovation, this paper not only lays a solid data foundation for understanding drone video relationships, but also facilitates a leap from low-level semantic analysis to high-level dynamic semantic cognitive modeling in remote sensing video analysis. Future research will focus on the following two directions: Firstly, deepening the temporal dimension modeling of dynamic remote sensing scene graph generation to enhance the ability to capture long-term evolutionary events and complex relationships; secondly, expanding the scene coverage and diversity of relationship categories in the dataset, continuously improving algorithm benchmarks, and promoting technological iteration and industry application implementation.
  • loading
  • [1]
    LIU Fang, WANG Jiahao, JIAO Licheng, et al. Remote sensing video tracking: Current status, challenges, and future[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 14338–14367. doi: 10.1109/JSTARS.2025.3573572.
    [2]
    MA Bokun, MU Caihong, LIU Yi, et al. RoSENet: Rotation and similarity enhancement network for multimodal remote sensing image land cover classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5511318. doi: 10.1109/TGRS.2025.3561850.
    [3]
    金晶, 王峰. 分布式多卫星协同遥感图像场景分类方法[J]. 电子与信息学报, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866.

    JIN Jing and WANG Feng. A distributed multi-satellite collaborative framework for remote sensing scene classification[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4677–4688. doi: 10.11999/JEIT250866.
    [4]
    GAO Feng, JIN Xuepeng, ZHOU Xiaowei, et al. MSFMamba: Multiscale feature fusion state space model for multisource remote sensing image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5504116. doi: 10.1109/TGRS.2025.3535622.
    [5]
    ZHANG Yin, YE Mu, ZHU Guiyi, et al. FFCA-YOLO for small object detection in remote sensing images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5611215. doi: 10.1109/TGRS.2024.3363057.
    [6]
    XIAO Yao, XU Tingfa, YU Xin, et al. A lightweight fusion strategy with enhanced interlayer feature correlation for small object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4708011. doi: 10.1109/TGRS.2024.3457155.
    [7]
    CHEN Tianxiang, YE Zi, TAN Zhentao, et al. MiM-ISTD: Mamba-in-mamba for efficient infrared small-target detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 5007613. doi: 10.1109/TGRS.2024.3485721.
    [8]
    姚婷婷, 肇恒鑫, 冯子豪, 等. 上下文感知多感受野融合网络的定向遥感目标检测[J]. 电子与信息学报, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560.

    YAO Tingting, ZHAO Hengxin, FENG Zihao, et al. A context-aware multiple receptive field fusion network for oriented object detection in remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(1): 233–243. doi: 10.11999/JEIT240560.
    [9]
    SUN Xian, WANG Peijin, YAN Zhiyuan, et al. FAIR1M: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 184: 116–130. doi: 10.1016/j.isprsjprs.2021.12.004.
    [10]
    周国宇, 张菁, 闫伊, 等. 聚焦注意力与紧致特征融合Transformer的城市遥感影像语义分割[J]. 电子与信息学报, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812.

    ZHOU Guoyu, ZHANG Jing, YAN Yi, et al. A focused attention and feature compact fusion Transformer for semantic segmentation of urban remote sensing images[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4790–4800. doi: 10.11999/JEIT250812.
    [11]
    CHEN Keyan, LIU Chenyang, CHEN Hao, et al. RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4701117. doi: 10.1109/TGRS.2024.3356074.
    [12]
    于国栋, 蒋一纯, 刘云清, 等. 一种空间语义联合感知的红外无人机目标跟踪方法[J]. 电子与信息学报, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613.

    YU Guodong, JIANG Yichun, LIU Yunqing, et al. A spatial-semantic combine perception for infrared UAV target tracking[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4242–4253. doi: 10.11999/JEIT250613.
    [13]
    JI Jingwei, KRISHNA R, LI Feifei, et al. Action genome: Actions as compositions of spatio-temporal scene graphs[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 10233–10244. doi: 10.1109/CVPR42600.2020.01025.
    [14]
    SADEGHI M A and FARHADI A. Recognition using visual phrases[C]. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recogniti, Colorado Springs, USA, 2011: 1745–1752. doi: 10.1109/CVPR.2011.5995711.
    [15]
    LU Cewu, KRISHNA R, BERNSTEIN M, et al. Visual relationship detection with language priors[C]. 14th European Conference Computer Vision-ECCV 2016, Amsterdam, The Netherlands, 2016: 852–869. doi: 10.1007/978-3-319-46448-0_51.
    [16]
    KRISHNA R, ZHU Yuke, GROTH O, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations[J]. International Journal of Computer Vision, 2017, 123(1): 32–73. doi: 10.1007/s11263-016-0981-7.
    [17]
    LIANG Yuanzhi, BAI Yalong, ZHANG Wei, et al. VrR-VG: Refocusing visually-relevant relationships[C]. 2019 IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 2019: 10402–10411. doi: 10.1109/ICCV.2019.01050.
    [18]
    SHANG Xindi, REN Tongwei, GUO Jingfan, et al. Video visual relation detection[C]. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, USA, 2017: 1300–1308. doi: 10.1145/3123266.3123380.
    [19]
    SHANG Xindi, DI Donglin, XIAO Junbin, et al. Annotating objects and relations in user-generated videos[C]. Proceedings of the 2019 on International Conference on Multimedia Retrieval, Ottawa, Canada, 2019: 279–287.
    [20]
    YANG Jingkang, PENG Wenxuan, LI Xiangtai, et al. Panoptic video scene graph generation[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 18675–18685. doi: 10.1109/CVPR52729.2023.01791.
    [21]
    XIA Guisong, BAI Xiang, DING Jian, et al. DOTA: A large-scale dataset for object detection in aerial images[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 3974–3983. doi: 10.1109/CVPR.2018.00418.
    [22]
    CHENG Gong, WANG Jiabao, LI Ke, et al. Anchor-free oriented proposal generator for object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5625411. doi: 10.1109/TGRS.2022.3183022.
    [23]
    RIZ L, CARAFFA A, BORTOLON M, et al. The MONET dataset: Multimodal drone thermal dataset recorded in rural scenarios[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, Canada, 2023: 2546–2554. doi: 10.1109/CVPRW59228.2023.00253.
    [24]
    YAN Qiwei, DENG Chubo, LIU Chenglong, et al. ReCon1M: A large-scale benchmark dataset for relation comprehension in remote sensing imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 4507022. doi: 10.1109/TGRS.2025.3589986.
    [25]
    DU Dawei, QI Yuankai, YU Hongyang, et al. The unmanned aerial vehicle benchmark: Object detection and tracking[C]. 15th European Conference Computer Vision-ECCV 2018, Munich, Germany, 2018: 370–386. doi: 10.1007/978-3-030-01249-6_23.
    [26]
    LYU Ye, VOSSELMAN G, XIA Guisong, et al. UAVid: A semantic segmentation dataset for UAV imagery[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 165: 108–119. doi: 10.1016/j.isprsjprs.2020.05.009.
    [27]
    HAMDI S, BOUINDOUR S, SNOUSSI H, et al. End-to-end deep one-class learning for anomaly detection in UAV video stream[J]. Journal of Imaging, 2021, 7(5): 90. doi: 10.3390/jimaging7050090.
    [28]
    TRAN T M, VU T N, NGUYEN T V, et al. UIT-ADrone: A novel drone dataset for traffic anomaly detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16: 5590–5601. doi: 10.1109/JSTARS.2023.3285905.
    [29]
    NGUYEN T T, NGUYEN P, LI Xin, et al. CYCLO: Cyclic graph transformer approach to multi-object relationship modeling in aerial videos[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 2868.
    [30]
    REN Shaoqing, HE Kaiming, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137–1149. doi: 10.1109/tpami.2016.2577031.
    [31]
    CAI Zhaowei and VASCONCELOS N. Cascade R-CNN: Delving into high quality object detection[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 6154–6162. doi: 10.1109/CVPR.2018.00644.
    [32]
    LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]. 2017 IEEE International Conference on Computer Vision, Venice, Italy, 2017: 2999–3007. doi: 10.1109/ICCV.2017.324.
    [33]
    ZHANG Shifeng, CHI Cheng, YAO Yongqiang, et al. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 9756–9765. doi: 10.1109/CVPR42600.2020.00978.
    [34]
    ZHU Benjin, WANG Jianfeng, JIANG Zhengkai, et al. AutoAssign: Differentiable label assignment for dense object detection[EB/OL]. https://arxiv.org/abs/2007.03496, 2020.
    [35]
    TIAN Zhi, SHEN Chunhua, CHEN Hao, et al. FCOS: A simple and strong anchor-free object detector[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(4): 1922–1933. doi: 10.1109/tpami.2020.3032166.
    [36]
    ZHANG Haoyang, WANG Ying, DAYOUB F, et al. VarifocalNET: An IoU-aware dense object detector[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 8501–8519. doi: 10.1109/CVPR46437.2021.00841.
    [37]
    CHEN Zehui, YANG Chenhongyi, LI Qiaofei, et al. Disentangle your dense object detector[C]. Proceedings of the 29th ACM International Conference on Multimedia, 2021: 4939–4948. doi: 10.1145/3474085.3475351. (查阅网上资料,未找到本条文献出版地信息,请确认).
    [38]
    TANG Kaihua, NIU Yulei, HUANG Jianqiang, et al. Unbiased scene graph generation from biased training[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 3713–3722. doi: 10.1109/CVPR42600.2020.00377.
    [39]
    ZELLERS R, YATSKAR M, THOMSON S, et al. Neural motifs: Scene graph parsing with global context[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 5831–5840. doi: 10.1109/CVPR.2018.00611.
    [40]
    TANG Kaihua, ZHANG Hanwang, WU Baoyuan, et al. Learning to compose dynamic tree structures for visual contexts[C]. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 6612–6621. doi: 10.1109/CVPR.2019.00678.
    [41]
    CONG Yuren, LIAO Wentong, ACKERMANN H, et al. Spatial-temporal transformer for dynamic scene graph generation[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 16352–16362. doi: 10.1109/ICCV48922.2021.01606.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(11)  / Tables(5)

    Article Metrics

    Article views (79) PDF downloads(4) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return