DroneRFc-MM: Anti-UAV Multimodal Detection Measured Dataset
-
摘要: 多模态信息融合技术可有效改善反无人机探测系统的泛化能力、鲁棒性与场景适配能力。针对现有反无人机探测数据集在模态种类、无人机型号和标注信息等方面的缺陷,该文公开了DroneRFc-MM多模态反无人机探测数据集。该数据集同步采集了云台相机、广角相机、射频天线、激光雷达、毫米波雷达和传声器阵列6类传感器数据,覆盖6种消费级无人机机型,并提供机型、定位、姿态和速度等细粒度标注,可支持目标探测、机型识别和运动方向推理等任务,并提供了易用的样本提取工具代码。最后,作为该数据集的使用示范,该文评估了Qwen系列最新模型在无人机飞行方向推理任务的表现。Abstract:
Objective: A comprehensive multimodal benchmark is developed for Anti-Unmanned Aerial Vehicle (UAV) detection in low-altitude urban environments. Existing datasets generally provide limited sensing modalities and UAV models, with relatively coarse annotations that constrain tasks requiring spatial, motion, and cross-modal information. DroneRFc-MM addresses these limitations by providing synchronized multimodal data, broader coverage of consumer-grade DJI UAV models, and fine-grained annotations for target detection, UAV model recognition, trajectory analysis, flight-direction reasoning, and multimodal fusion evaluation. Methods: DroneRFc-MM is synchronously collected using six heterogeneous sensor types: a Pan-Tilt-Zoom (PTZ) camera, a fisheye camera, a Radio Frequency (RF) antenna, LiDAR, millimeter-wave radar, and a microphone array. Data are acquired on an open rooftop at a university in Zhejiang Province, representing a typical urban low-altitude environment. The dataset contains recordings of six consumer-grade DJI UAV models. All devices are synchronized using a common network time reference, with inter-device timestamp discrepancies of approximately 0.3 s. The UAVs fly in “H”-shaped and vertical reciprocating trajectories at distances of 20–60 m from the sensor array. Fine-grained annotations, including UAV model, position, attitude, and velocity, are derived from flight logs. For the flight-direction reasoning task, approximately 5-s multimodal clips are generated, including camera videos, RF spectrogram videos, microphone audio, and coordinate-based text representations of radar point-cloud data. Zero-shot inference is conducted using Qwen 3.6-Plus and Qwen 3.5-Omni-Plus with unified prompts. Prediction accuracy and inference time are evaluated by comparing predicted directions with ground-truth directions calculated from UAV positioning data. Results and Discussions: The DroneRFc-MM dataset provides multimodal data from six sensor types and six consumer-grade DJI UAV models, together with fine-grained annotations and sample extraction tools. In the flight-direction reasoning task, the Qwen-series multimodal large language models (MLLMs) achieve accuracies ranging from 20% to 30% across the different input modalities. The inference time is also relatively long, with the mean response time exceeding 40 s for most sensor inputs. These results indicate that current general-purpose MLLMs can capture weak motion-related information from UAV videos, audio, RF spectrograms, and point-cloud data, but their accuracy and response speed remain insufficient for practical real-time Anti-UAV detection. Conclusions: DroneRFc-MM provides a multimodal benchmark for Anti-UAV detection, UAV model recognition, flight-direction reasoning, and multimodal model evaluation. The dataset integrates six sensor types, six consumer-grade DJI UAV models, and fine-grained annotations within a common measurement framework. The experimental results show that current general-purpose MLLMs remain limited in flight-direction reasoning and real-time inference in Anti-UAV scenarios. Domain-specific pre-training, supervised fine-tuning, knowledge augmentation, and lightweight inference are therefore needed to improve their practical utility. Future work will expand the dataset scale and application scenarios to support intelligent and efficient low-altitude airspace management systems. -
Key words:
- Anti-UAV detection /
- Multimodal dataset /
- Low-altitude economy
-
表 1 常用反无人机探测设备比较
设备 探测范围(km) 优势 劣势 可见光相机 0~6 检测算法成熟,结果直观 容易受光照、遮挡影响 红外成像仪 0~6 可夜间工作,可信度高 成本高,分辨率低 X 波段雷达 3~15 探测距离远,可探知方位和速度 无法探测静止目标,使用受管制,虚警多 射频天线 0.5~10 探测距离远,包含信息丰富 成本高,城市同频干扰多,需指定频段 表 2 本数据集和现有反无人机探测多模态数据集比较
数据集 采集设备 无人机机型 Anti-UAV 可见光相机、红外相机 DJI: Phantom 4, Spreading Wings S1000, Tello MMAUD 立体相机、激光雷达、毫米波雷达、传声器阵列 DJI: Mavic 2, Mavic 3, Phantom 4, Avata, M300 DroneRFc-MM 云台相机、广角相机、射频天线、激光雷达、毫米波雷达、传声器阵列 DJI: Mini 2 SE, Mini 3, Mavic Air 2S, Air 3, Mavic 3, Avata 表 3 采集设备列表
探测设备 品牌型号 主要参数规格 PTZ相机 海康威视DS-2DC6423IW 分辨率2 560×1 440,23倍光学变焦 广角相机 海康威视DS-2CD3346 分辨率2 560×1 440,视场角H180°V98° 全向射频天线 Ettus VERT2450 配合NI USRP-2955,采样率100 MHz,中心频率2.45 GHz,单次连续采样时间50 ms 激光雷达 速腾聚创EM4 线数512,点频2592 万点/s,视场角H120°V27° 毫米波雷达 Arbe Phoenix 探测波段77~81 GHz,点云模式,视场角H100°V30° 电容传声器 爱华AWA14411 4通道阵列,采样率25.6 kHz,动态范围6.5~137 dB 表 4 无人机型号与采集编号
无人机型号 机型编号 DJI Mavic 3 A1 DJI Avata B1 DJI Mini 2 SE C1 DJI Mini 3 E1 DJI Air 3 F1 DJI Air 2s G1 表 5 飞行方向推理任务提示词
种类 提示词 提示词模板
(“{}”内为可变内容)“{data_type} showing a drone. {extra_explanation} Please identify the moving direction of the drone. Respond with only one uppercase letter: F for Forward (approaching the {device}), B for Backward(separating from the {device}), U for Upward, D for Downward, L for Left, R for Right, S for Stopping (moving distance < 2 m).” 云台相机 {data_type}=”Here is a PTZ camera video”, {device}=”camera” 广角相机 {data_type}=”Here is a fisheye camera video”, {device}=”camera” 射频天线 {data_type}=”Here is an RF spectrum video”, {device}=”antenna” 激光雷达 {data_type}=”Here are lidar pointcloud data”, {device}=”lidar”,
{extra_explanation}=” The data are in compact CSV format: each row is one point with three comma-separated values (x,y,z). Frames are separated by a blank line.”毫米波雷达
(视频输入){data_type}=”Here are bird-view, front-view, and right-view videos of a millimeter wave radar pointcloud.”, {device}=”radar” 毫米波雷达
(文本输入){data_type}=”Here are millimeter wave radar pointcloud data”, {device}=”antenna”,
{extra_explanation}=” The data are in compact CSV format: each row is one point with three comma-separated values (x,y,z). Frames are separated by a blank line.”传声器 {data_type}=”Here are 4 audio recordings”, {device}=”microphone array”, {extra_explanation}=” The audio uploading sequence is arranged from left to right in accordance with the mounting positions of the microphone array.” 表 6 各模态无人机飞行方向推理准确率
采集设备 输入模态 总问答 正确回答 准确率(%) 激光雷达 点云位置文本 317 97 30.60 传声器 音频 476 137 28.78 毫米波雷达 点云位置文本 469 125 26.65 广角相机 视频 477 120 25.15 云台相机 视频 476 111 23.31 射频天线 时频谱视频 318 70 22.01 毫米波雷达 点云三视图视频 469 95 20.25 -
[1] DOLATA M and SCHWABE G. Moving beyond privacy and airspace safety: Guidelines for just drones in policing[J]. Government Information Quarterly, 2023, 40(4): 101874. doi: 10.1016/j.giq.2023.101874. [2] LAKHWANI T S, SINJANA Y, and KAPOOR A P. Dynamic medical logistics with drone-truck collaboration and pheromone decay in ant colony optimization[J]. International Journal of Systems Science: Operations & Logistics, 2026, 13(1): 2626506. doi: 10.1080/23302674.2026.2626506. [3] XIE Yingdong, GUO Yan, and GAO Jie. Object detection and UAV inspection for intelligent agriculture-husbandry application research[C]. The 2025 International Conference on Smart Agriculture and Artificial Intelligence, Xi’an, China, 2025: 223–228. doi: 10.1145/3767624.3767656. [4] 钱志鸿, 王义君. 低空经济赋能者: 智能无人机技术体系综述与展望[J]. 电子与信息学报, 2026, 48(1): 1–33. doi: 10.11999/JEIT251246.QIAN Zhihong and WANG Yijun. Intelligent unmanned aerial vehicles for low-altitude economy: A review of the technology framework and future prospects[J]. Journal of Electronics & Information Technology, 2026, 48(1): 1–33. doi: 10.11999/JEIT251246. [5] 王威, 佘丁辰, 王加琪, 等. 多模型融合的无人机异常航迹校正方法[J]. 电子与信息学报, 2025, 47(5): 1332–1344. doi: 10.11999/JEIT241026.WANG Wei, SHE Dingchen, WANG Jiaqi, et al. Multi-model fusion-based abnormal trajectory correction method for unmanned aerial vehicles[J]. Journal of Electronics & Information Technology, 2025, 47(5): 1332–1344. doi: 10.11999/JEIT241026. [6] LIU Ziyi, AN Pei, YANG You, et al. Vision-based drone detection in complex environments: A survey[J]. Drones, 2024, 8(11): 643. doi: 10.3390/drones8110643. [7] 俞宁宁, 毛盛健, 周成伟, 等. DroneRFa: 用于侦测低空无人机的大规模无人机射频信号数据集[J]. 电子与信息学报, 2024, 46(4): 1147–1156. doi: 10.11999/JEIT230570.YU Ningning, MAO Shengjian, ZHOU Chengwei, et al. DroneRFa: A large-scale dataset of drone radio frequency signals for detecting low-altitude drones[J]. Journal of Electronics & Information Technology, 2024, 46(4): 1147–1156. doi: 10.11999/JEIT230570. [8] 任俊宇, 俞宁宁, 周成伟, 等. DroneRFb-DIR: 用于非合作无人机个体识别的射频信号数据集[J]. 电子与信息学报, 2025, 47(3): 573–581. doi: 10.11999/JEIT240804.REN Junyu, YU Ningning, ZHOU Chengwei, et al. DroneRFb-DIR: An RF signal dataset for non-cooperative drone individual identification[J]. Journal of Electronics & Information Technology, 2025, 47(3): 573–581. doi: 10.11999/JEIT240804. [9] YU Ningning, WU Jiajun, ZHOU Chengwei, et al. Open set learning for RF-based drone recognition via signal semantics[J]. IEEE Transactions on Information Forensics and Security, 2024, 19: 9894–9909. doi: 10.1109/TIFS.2024.3463535. [10] YU Ningning, WU Jiajun, ZHOU Chengwei, et al. SMNet: Multi-drone detection and classification via monitoring flight control signals on spectrograms[J]. IEEE Transactions on Cognitive Communications and Networking, 2025, 11(5): 3245–3259. doi: 10.1109/TCCN.2025.3537100. [11] KÜMMRITZ S. The sound of surveillance: Enhancing machine learning-driven drone detection with advanced acoustic augmentation[J]. Drones, 2024, 8(3): 105. doi: 10.3390/drones8030105. [12] SUN Chunlin, MAO Xingpeng, TANG Zhibo, et al. Radar false alarm suppression based on target spatial temporal stationarity for UAV detecting[J]. Drones, 2024, 8(12): 699. doi: 10.3390/drones8120699. [13] JIANG Nan, WANG Kuiran, PENG Xiaoke, et al. Anti-UAV: A large-scale benchmark for vision-based UAV tracking[J]. IEEE Transactions on Multimedia, 2023, 25: 486–500. doi: 10.1109/TMM.2021.3128047. [14] YUAN Shenghai, YANG Yizhuo, NGUYEN T H, et al. MMAUD: A comprehensive multi-modal anti-UAV dataset for modern miniature drone threats[C]. The 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 2024: 2745–2751. doi: 10.1109/ICRA57147.2024.10610957. [15] SHI Rui, YU Xiaodong, WANG Shengming, et al. RFUAV: A benchmark dataset for unmanned aerial vehicle detection and identification[J/OL]. arXiv preprint arXiv: 2503.09033, 2025. doi: 10.48550/arXiv.2503.09033. [16] ZHU Chen, ZHAO Zhouxiang, SHAN Zejing, et al. Robust target detection of intelligent integrated optical camera and mmWave radar system[J]. Digital Signal Processing, 2024, 145: 104336. doi: 10.1016/j.dsp.2023.104336. [17] SAKELLARIOU N, LALAS A, VOTIS K, et al. Multi-sensor fusion for UAV classification based on feature maps of image and radar data[J/OL]. arXiv preprint arXiv: 2410.16089, 2024. doi: 10.48550/arXiv.2410.16089. [18] ZHAO W X, ZHOU Kun, LI Junyi, et al. A survey of large language models[J]. Frontiers of Computer Science, 2026, 20(12): 2012627. doi: 10.1007/s11704-026-60308-3. [19] LI Junnan, LI Dongxu, SAVARESE S, et al. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models[C]. The 40th International Conference on Machine Learning, Honolulu, USA, 2023: 814. doi: 10.5555/3618408.3619222. [20] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale[C]. The 9th International Conference on Learning Representations, 2021. [21] RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C/OL]. The 38th International Conference on Machine Learning, 2021. [22] XI Zhiheng, CHEN Wenxiang, GUO Xin, et al. The rise and potential of large language model based agents: A survey[J]. Science China Information Sciences, 2025, 68(2): 121101. doi: 10.1007/s11432-024-4222-0. -
下载: