Latest Articles

Articles in press have been peer-reviewed and accepted, which are not yet assigned to volumes/issues, but are citable by Digital Object Identifier (DOI).
Display Method:
Dynamic Frequency Guidance and Semantic Purification Network for Infrared Dim and Small Target Detection
SHEN Xiaoru, CHANG Xia, WEI Wenjie
Available online  , doi: 10.11999/JEIT260582
Abstract:
  Objective  Infrared dim and small target detection remains challenging because targets are small, have low contrast, and are easily obscured by complex background clutter. Although existing methods improve detection performance, several limitations remain, including feature loss during downsampling, insufficient use of physical priors, and contamination of semantic features by background noise. These limitations can result in missed detections and false alarms. To address these problems, a detection network combining dynamic frequency guidance and semantic purification is proposed. Deep semantic features are dynamically coupled with adaptive frequency-domain priors to strengthen target and edge representations, improve detection accuracy, reduce false alarms, and enhance robustness in complex scenes.  Methods  A framework consisting of a shared encoder and a dual-branch decoder is constructed for infrared dim and small target detection (Fig. 1). First, a Local Target Enhancement (LTA) module is used to preprocess the original infrared image, strengthen small-target boundaries, and suppress background noise. The Sobel operator is then applied to extract edge information and provide a structural prior for subsequent feature learning. During feature extraction, a Learnable Downsampling (LD) module replaces conventional max pooling. Stride-2 convolution adaptively compresses the feature maps while preserving target details and reducing information loss. Residual blocks and progressive multiscale feature fusion are also used to strengthen semantic representations at different levels. A Dynamic Frequency Guidance Module (DFGM) is then introduced to enhance high-frequency details and target-edge information. Deep semantic features are used to dynamically predict the cutoff radius of a high-pass filter. A content-adaptive high-pass filter mask is constructed in the frequency domain to suppress low-frequency components. The filtered features are subsequently transformed back to the spatial domain to generate a frequency guidance map. To suppress spurious background responses introduced during frequency-domain enhancement, a Semantic Purification Module (SPM) is further designed. The deepest semantic features are used to generate a spatial attention map for top-down unidirectional filtering of the frequency guidance features. Background interference is thereby suppressed without contaminating high-level semantic representations with low-level noise. Finally, the dual-branch decoder reconstructs the target region and edge structure, thereby improving segmentation consistency, target boundary localization, and target-shape recovery.  Results and Discussions  Experimental results on the SIRST-Aug and IRSTD-1k datasets show that the proposed method outperforms mainstream comparison methods. On SIRST-Aug (Table 1), the method achieves the best Iintersection over Union (IoU), normalized Intersection over Union (nIoU), area under the receiver operating characteristic curve (AUC), and probability of detection (Pd), reaching 76.66%, 73.20%, 94.45%, and 99.17%, respectively. The false-alarm rate (Fa) is maintained at 33.96 × 10–6. On IRSTD-1k (Table 1), the method achieves the best IoU, AUC, and Fa values of 67.17%, 89.52%, and 10.62 × 10–6, respectively, while obtaining an nIoU of 67.67% and a Pd of 92.59%. Qualitative comparisons further show that the proposed method provides more accurate target localization and more complete target-shape segmentation under complex background interference (Figs. 4 and 5). The model contains 2.94 M parameters, requires 14.68 G FLoating-point OPerations (FLOPs), and has an average inference time of 7.54 ms, providing a favorable balance between detection performance and computational cost (Table 2). Ablation experiments further confirm the contributions of the key modules (Table 3). Adding DFGM increases AUC from 93.17% to 94.60%. Introducing LD increases nIoU to 73.32%, and adding SPM further increases IoU to 76.66%. The loss-function ablation results also show that the combined mask and edge losses provide a better balance among segmentation accuracy, target boundary localization, and false-alarm suppression (Table 4). Overall, the proposed method provides strong detection performance and effective false-alarm suppression.  Conclusions  An infrared dim and small target detection network based on dynamic frequency guidance and semantic purification is proposed. Through the coordinated use of LD, DFGM, and SPM, target details are effectively preserved, high-frequency structural information is strengthened, and semantic representations are purified. Background interference is consequently suppressed, and target-boundary recovery is improved. Experiments on SIRST-Aug and IRSTD-1k demonstrate that the proposed method outperforms existing mainstream methods across multiple evaluation metrics and maintains strong robustness in complex scenes. Future work can incorporate richer physical priors and self-supervised learning strategies to further improve model generalization and robustness under challenging conditions.
Secrecy Performance Analysis of Multi-tag Bistatic Backscatter Communication Systems With Outdated CSI and Link Correlation
LIU Yingting, TANG Yong, LI Xingwang
Available online  , doi: 10.11999/JEIT260823
Abstract:
  Objective  Due to feedback delay, the Channel State Information (CSI) used during tag selection may become outdated before data transmission, causing a mismatch between the selected tag and the tag with the largest backscatter-link channel gain during transmission. Most existing studies of outdated CSI assume independent and identically distributed (i.i.d.) channels, which may not adequately reflect the heterogeneous characteristics of practical links caused by different propagation distances. In addition, when the eavesdropper is close to the destination, the legitimate and eavesdropping links may experience correlated fading. Neglecting these factors may cause theoretical results to deviate from actual system performance. Accordingly, under independent and non-identically distributed (i.n.i.d.) channel conditions, the secrecy performance of a multi-tag Bistatic Backscatter Communication (BBC) system is studied under the joint effects of outdated CSI and correlation between the legitimate and eavesdropping links.  Methods  Candidate tags are ranked according to their backscatter-link channel gains, and the tag with the largest backscatter-link channel gain is selected for transmission. An outdated CSI model is introduced to characterize the mismatch between the CSI used during tag selection and the CSI during data transmission. Because tag selection depends only on the backscatter-link CSI, CSI aging is modeled only for this link. A correlated Rayleigh fading model is used to characterize the statistical dependence between the legitimate and eavesdropping links. Under i.n.i.d. Rayleigh fading, order statistics are used to derive the Probability Density Function (PDF) of the selected tag’s outdated backscatter-link channel gain, whereas the joint PDF of the legitimate- and eavesdropping-link channel gains is derived by incorporating their correlated fading relationship. Based on these results, closed-form and high-transmit-power asymptotic expressions for the Secrecy Outage Probability (SOP) are derived. The secrecy performance is further analyzed in terms of the legitimate-to-eavesdropping channel-gain ratio.  Results and Discussions  Monte Carlo simulations validate the analytical and asymptotic results. The results show that outdated CSI significantly degrades secrecy performance because feedback delay causes the CSI used during tag selection to differ from the CSI during transmission, so the selected tag may no longer provide the largest backscatter-link channel gain. For a fixed legitimate-to-eavesdropping channel-gain ratio, a secrecy outage floor emerges at high transmit power (Fig. 2), indicating that increasing transmit power alone cannot eliminate this performance bottleneck. In contrast, increasing the legitimate-to-eavesdropping channel-gain ratio effectively mitigates the outage floor and yields a secrecy diversity order of 1 with respect to this ratio (Fig. 4). Under the considered system model, correlation between the legitimate and eavesdropping links also improves secrecy performance (Fig. 3) by reducing the probability that the legitimate link experiences severe fading while the eavesdropping link remains strong. Moreover, despite outdated CSI, the proposed tag selection scheme based on backscatter-link channel-gain ranking remains effective and consistently outperforms random tag selection (Fig. 3).  Conclusions  A secrecy-performance analysis framework is developed for multi-tag BBC systems with outdated CSI and correlated legitimate and eavesdropping links. Closed-form SOP and high-transmit-power asymptotic expressions characterize the effects of outdated CSI, link correlation, and the legitimate-to-eavesdropping channel-gain ratio. The results identify outdated CSI as a major source of secrecy degradation and indicate that low-latency feedback is beneficial. Increasing the legitimate-link gain advantage over the eavesdropping link and selecting tags according to backscatter-link channel-gain ranking effectively improve secrecy performance.
Resource Allocation and Node Deployment for Multi-UAV Integrated Localization and Communication with Differentiated Requirements
ZHAO Yicheng, WANG Hai, QIN Zhen, SUN Weihao
Available online  , doi: 10.11999/JEIT260819
Abstract:
  Objective  In areas with weak ground-network coverage or limited Global Navigation Satellite System (GNSS) availability, multiple dual-functional Unmanned Aerial Vehicles (UAVs) can provide uplink communication and cooperative localization services. Ground terminals have different requirements for data volume, minimum communication rate, localization-accuracy threshold, service priority, and communication and localization service weights. Maximizing communication rate, localization accuracy, or a weighted sum of the two cannot directly reflect these differentiated requirements. A service below its activation threshold may be unusable, whereas resources allocated after demand saturation provide little additional value. A quasi-static service period with prior terminal positions is therefore considered. A topology-dependent Value of Service (VoS) is formulated, and Physical Resource Block (PRB) scheduling and UAV positions are jointly optimized.  Methods  Data or Sounding Reference Signals (SRSs) are transmitted by terminals over assigned PRBs. Time Difference of Arrival (TDoA) measurements are formed by synchronized UAVs, and a probabilistic air-to-ground channel model determines communication access and valid localization anchors. Communication VoS combines a sigmoidal rate utility with data completeness. Position Error Bound (PEB), derived from the accumulated Fisher Information Matrix (FIM), is mapped to an exponential utility to quantify localization VoS. Terminal priorities and communication and localization service weights are used to aggregate the two service values. The resulting problem P0 couples binary scheduling, non-concave utilities, time-frequency resources, service relationships, localization geometry, and UAV positions. For a fixed topology, the average number of allocated PRBs and active slots are used to construct a continuous service-level resource profile. This profile approximates the original schedule but does not provide an upper bound. Bandwidth responses are obtained by deterministic one-dimensional branch-and-bound search with damped Newton refinement and shadow-price bisection. Slot responses are obtained by deterministic comparison of a finite set of service-boundary, integer-slot, and feasible-domain-boundary candidates. Adjacent-integer recovery and deterministic slot packing are then used to construct a feasible integer schedule under resource-capacity and terminal-power constraints. If packing fails, the upward-rounded component with the smallest unit VoS loss is rolled back, and the schedule is repacked. VoS is then recomputed from the recovered integer schedule. Thus, the reported schedules remain feasible for P0 without any claim of global optimality. The fixed-topology resource-allocation procedure produces a complete resource response, including the recovered feasible schedule and its VoS, which is subsequently used to drive deployment. Coverage-repair, localization-geometry-improvement, and resource-saving candidates are generated according to service deficits, service relationships, and resource consumption. Normalized proxy scores, per-UAV candidate truncation, and beam search reduce the number of complete resource-response evaluations. For each retained deployment, service relationships are rebuilt, including communication access and localization anchors, and the FIM and resource allocation are recomputed. A candidate is accepted only when the VoS gain obtained from its complete resource response exceeds the preset threshold. The resulting outer sequence is monotonic and bounded and converges to a stable point within the generated candidate set. Resource allocation and node deployment remain coupled throughout the procedure. For each retained topology, communication access, localization anchors, the FIM, and the feasible integer schedule are recomputed before VoS is evaluated. Proxy scores are used only to rank candidates, whereas final acceptance is always based on the complete resource response. The next topology is therefore not selected from distance or localization geometry alone. The accepted update reflects terminal demand, resource scarcity, and localization geometry under the same feasible scheduling constraints used in the objective. This design keeps the optimization objective consistent with the final deployment decision and avoids a geometry-only selection rule.  Results and Discussions  Simulations are conducted in a 700 m × 700 m area with four UAVs at a baseline deployment height of 150 m. Communication-dominant, localization-dominant, and balanced terminals are included in equal proportions. At each load, all methods share 30 independent scenarios and identical random seeds. The deployment-height experiment reports a 95% confidence interval based on 1,000 scenario-level bootstrap resamples. For fair comparisons, all resource-allocation methods use the same topology, resource pool, and random scenarios. All deployment methods use the same proposed joint resource-allocation response, are evaluated from a cold start, and are subject to the same cap on complete resource-response evaluations, with initialization and training costs included. VoS-driven allocation yields smaller communication-rate and localization-accuracy demand deviations than the communication-priority and average-utility metrics because service saturation redirects resources from overprovisioned requests to insufficiently served requests (Fig. 2). In the convergence test, three feasible initializations converge to similar system VoS values. Damping suppresses oscillations near transitions between non-concave segments, and the small-scale benchmark indicates limited empirical loss from the continuous response and integer recovery (Fig. 3). With increasing terminal load, the proposed joint allocator outperforms a Particle Swarm Optimization (PSO)-based slot-response variant, as well as non-joint, learning-based, weight-driven, and random allocation methods (Fig. 4). Communication service is more sensitive to terminal load because both communication rate and data completeness must be maintained, whereas localization service accumulates information across anchors and slots. Resource-pool tests further show that additional bandwidth provides little benefit when slots are scarce, whereas increasing SRS bandwidth more directly improves localization. The structured deployment search achieves higher VoS with shorter end-to-end deployment time than PSO, Differential Evolution (DE), and Bayesian Optimization (BO) by focusing complete resource-response evaluations on service-aware candidates (Fig. 5). VoS first increases and then decreases with UAV deployment height because improved line-of-sight probability and multi-anchor visibility compete with increased propagation distance and path loss. Horizontal deployment updates balance communication-link quality and localization geometry according to terminal requirements (Fig. 6).  Conclusions  The proposed framework evaluates communication and localization services according to demand satisfaction and coordinates time-frequency resource allocation with iterative UAV deployment updates through complete resource responses and structured candidate search. It improves demand matching, system VoS, deployment efficiency, and interpretability while preserving feasibility under the original scheduling constraints. Prior terminal positions and a quasi-static service period are assumed. Future work will address dynamic terminal movement and changing demands over longer service periods through dynamic resource allocation and continuous UAV trajectory optimization.
Impact of Wireless Priors on the Computation and Energy Cost of MU-MIMO Precoding Learning
CONG Pengyu, HAN Shengqian, DENG Mingyu, LIU Shengjie, YANG Chenyang, SHEN Songhui
Available online  , doi: 10.11999/JEIT260388
Abstract:
  Objective  This paper investigates downlink Multi-User Multi-Input Multi-Output (MU-MIMO) precoding policy learning from the perspectives of computational complexity and energy consumption. Traditional numerical optimization algorithms achieve high performance but exhibit rapidly increasing computational complexity as the numbers of base station antennas and served users increase, leading to high inference latency and energy consumption. In recent years, deep learning has been widely adopted to reduce online computational cost; however, existing evaluations generally rely on training or inference time and FLoating-point OPerations (FLOPs), without direct measurements of energy consumption and power. More importantly, the computational cost of a deep learning model is closely related to network architecture design, which should effectively exploit the wireless prior knowledge of the precoding policy. Therefore, this paper develops a network architecture that matches the multidimensional permutation properties of the precoding policy and systematically investigates how wireless priors affect computational complexity and energy consumption through comprehensive hardware-based measurements and simulations.  Methods  The MU-MIMO precoding policy is formulated as a mapping from multiuser channel information to the optimal precoding matrix under a transmit power constraint. The optimal policy satisfies multidimensional joint permutation equivariance and invariance with respect to user indices, receive antenna indices, and base station antenna indices. To exploit these wireless priors, an Attention-based Graph Neural Network (AGNN) is proposed based on a hypergraph structure, in which the update and aggregation operations satisfy the required equivariance and invariance properties. An attention mechanism is incorporated to model inter-user interference and improve generalization across different numbers of users. For broadband precoding, multi-subcarrier channel information is aggregated at the input layer to construct an expanded feature representation. To quantify computational energy cost, a hardware measurement platform is developed to collect energy consumption and power for the GPU, CPU, and DRAM during both training and inference. Simulations are conducted using 3GPP TR 38.901 Urban Macrocell (UMa) channel datasets with different antenna array sizes and bandwidth configurations. The proposed AGNN is compared with a traditional numerical optimization algorithm based on Zero-Forcing Block Diagonalization (ZFBD) with greedy user pairing and two Transformer-based architectures that satisfy only one-dimensional permutation equivariance.  Results and Discussions  Two major findings are obtained. First, incomplete exploitation of wireless priors results in inferior performance and higher computational cost. In the MU-MISO scenario, the Transformer-based architectures achieve lower spectral efficiency than the ZFBD+Greedy baseline while requiring substantially larger models and higher inference FLOPs than AGNN. By matching the multidimensional permutation properties of the precoding policy, AGNN achieves higher spectral efficiency while reducing inference FLOPs by approximately one order of magnitude. Hardware measurements further demonstrate that AGNN reduces inference energy consumption and power on both the CPU and GPU. Second, in small- and large-scale MU-MIMO scenarios, ZFBD+Greedy increases the system sum rate by 10.9×, whereas inference FLOPs, inference time, and inference energy increase by 79.0×, 21.3×, and 38.6×, respectively. In contrast, AGNN increases the system sum rate by 11.5×, while inference FLOPs increase by only 1.89×. Meanwhile, inference time and inference energy are reduced to 0.03× and 0.20×, respectively. These results demonstrate that exploiting the multidimensional permutation properties of the precoding policy provides an effective approach for reducing computational complexity, inference latency, and energy consumption in large-scale 6G MU-MIMO systems.  Conclusions  This paper investigates the effect of wireless priors on the computational complexity and energy consumption of MU-MIMO precoding learning. By exploiting the multidimensional joint permutation equivariance and invariance of the optimal precoding policy, an AGNN is developed that is well matched to these properties. A hardware-aware measurement platform is established to obtain direct measurements of energy consumption and power for the GPU, CPU, and DRAM during training and inference. Simulations based on 3GPP TR 38.901 UMa channel datasets demonstrate that Transformer-based architectures satisfying only one-dimensional permutation equivariance achieve lower spectral efficiency while incurring substantially higher computational and energy costs. In contrast, AGNN achieves higher spectral efficiency while substantially reducing inference FLOPs, inference time, inference energy consumption, and training complexity. As system size increases, traditional numerical optimization algorithms exhibit much faster growth in computational and energy costs than in system sum rate, whereas the proposed learning method based on wireless priors maintains low inference latency and energy consumption. Overall, exploiting the wireless prior knowledge of the MU-MIMO precoding policy in network architecture design provides an effective solution for computationally efficient and energy-efficient high-dimensional precoding optimization in future 6G networks.
Determination of Key Geometric Parameters for Spaceborne Dual-Beam Along-Track Interferometric SAR under Asymmetric Geometry
SHEN Qingyuan, ZHANG Yangyang, LAI Tao, ZHU Yuting, XIE Zhifeng, WANG Xiaoqing
Available online  , doi: 10.11999/JEIT260870
Abstract:
  Objective  Spaceborne Dual-Beam Along-Track Interferometric Synthetic Aperture Radar (DBATI-SAR) acquires fore- and aft-looking radial velocities to support two-dimensional ocean surface current retrieval. The stability of the inversion depends on the relative directions of the two Radar Line-Of-Sight (RLOS) projections on the target local tangent plane. Conventional flat-Earth geometry models generally assume symmetric squint angles or time offsets. However, orbital curvature, Earth curvature, Earth rotation, and local projection nonlinearity produce asymmetric spaceborne geometries. Therefore, symmetric squint angles do not necessarily guarantee orthogonal ground-projected RLOS directions. A method is developed to determine the fore- and aft-looking observation times, squint angles, and down-looking angles directly under the ground-projected RLOS orthogonality constraint.  Methods  A complete Earth-Centered Earth-Fixed (ECEF) geometry model is established from the satellite state and target position. The RLOS is projected onto the target local tangent plane, and a signed ground-projected RLOS angle is defined relative to the zero-squint reference direction. The squint and down-looking angles are calculated from the observation times rather than prescribed independently. A two-dimensional observation geometry matrix is constructed to relate the horizontal current components to the two radial velocities. Under an ideal symmetric geometry used for the analytical derivation, the singular values and condition number show that a one-sided ground-projected RLOS angle of \begin{document}$ {45}^{{^{\circ}}} $\end{document} provides the best inversion conditioning. A normalized geometric amplification factor is introduced to quantify the additional error amplification caused by nonorthogonal projections. Near the zero-squint reference point, a closed-form analytical leading term is derived to relate the observation-time offset to the ground-projected RLOS angle. This expression reveals the effects of the reference slant range, down-looking angle, equivalent along-track velocity, and second-order range-history curvature. A local polynomial inverse mapping is then constructed from complete ECEF forward-geometry samples. Target ground-projected RLOS angles of \begin{document}$ {-45}^{{^{\circ}}} $\end{document} and \begin{document}$ +{45}^{{^{\circ}}} $\end{document} are substituted into the inverse mapping to obtain the fore- and aft-looking observation times, after which the corresponding squint and down-looking angles are calculated.  Results and Discussions  Numerical experiments are conducted using 500 random orbit-parameter sets. Fifth- and sixth-order polynomials produce relatively large inverse-mapping errors, whereas a seventh-order model substantially improves the accuracy. With 20 samples, the mean ground-projected RLOS angle error is 0.0392°, and further increases in polynomial order or sample number provide limited improvement (Fig. 2, Table 2). The analytical leading term agrees well with the complete ECEF geometry model. For a representative case, the angular root-mean-square error and maximum deviation are approximately \begin{document}$ {0.11}^{\circ } $\end{document} and \begin{document}$ {0.31}^{\circ } $\end{document}, respectively. Across the 500 cases, the root-mean-square errors of the fore- and aft-looking observation times predicted by the analytical leading term are 0.95 s and 0.57 s, respectively. These results indicate that the analytical leading term captures the dominant observation-time scale, while the numerical inverse mapping accounts for asymmetric higher-order effects (Fig. 3). The conventional flat-Earth geometry model produces a squint-angle correction of up to approximately \begin{document}$ {3}^{\circ } $\end{document} and a ground-projected RLOS orthogonality error of approximately \begin{document}$ {6.5}^{\circ } $\end{document}. The proposed method reduces the mean ground-projected RLOS angle error to approximately \begin{document}$ {0.04}^{\circ } $\end{document} and decreases the normalized geometric amplification factor from approximately 1.12 to 1.000 7 (Fig. 4). The semimajor axis has the strongest effect on the aft-looking observation time, which varies from approximately 34 s to 76 s, while variations caused by other orbital parameters remain within 7 s (Fig. 5). Under combined squint- and down-looking-angle perturbations of up to 0.2°, most samples retain ground-projected RLOS angle errors below \begin{document}$ {1}^{\circ } $\end{document} (Fig. 6).  Conclusions  A method is proposed to determine key geometric parameters for spaceborne DBATI-SAR under asymmetric geometry. The analytical leading term explains the dominant observation-time scale, whereas the complete ECEF numerical inverse mapping accurately determines the fore- and aft-looking observation times and corresponding squint and down-looking angles. The proposed method provides substantially higher geometric accuracy than the conventional flat-Earth geometry approach. Practical mission design should also consider pulse repetition frequency, azimuth ambiguity, Doppler bandwidth, beam-steering range, and along-track interferometric coherence.
Adaptive Fusion Detection for Dual-radar with Non-identical Clutter Distributions
ZHOU Baoyi, YANG Yong, YANG boyu
Available online  , doi: 10.11999/JEIT260616
Abstract:
  Objective  Small sea-surface targets have low Radar Cross Sections (RCSs), and their echoes are easily masked by intense sea clutter, resulting in extremely low Signal-to-Clutter Ratios (SCRs). Single-radar systems are constrained by a single operating frequency band and fixed observation angles, which limits their detection performance. Multi-radar collaborative detection provides an effective approach to overcoming this limitation. Existing multi-radar fusion detection methods are developed for different information fusion levels. However, feature-level and signal-level fusion methods generally assume identical clutter distributions, whereas decision-level fusion is less dependent on the specific clutter distribution. In practical multi-radar detection, differences in radar parameters, including frequency band, range resolution, and grazing angle, result in non-identical statistical characteristics of sea clutter. Furthermore, target RCS varies with observation azimuth and operating frequency, resulting in different SCRs for the same target observed by different radars and causing model mismatch in conventional fusion detectors. To address the coexistence of non-identical clutter distributions and different SCRs, a Neyman-Pearson (NP) criterion-based dual-radar Adaptive Fusion Detection method, termed NP-AFD, is proposed for Rayleigh and Weibull sea clutter.  Methods  Amplitude distribution fitting is performed on measured S-band and X-band sea clutter data using five commonly used models: Rayleigh, lognormal, Weibull, Gamma, and K distributions. Fitting accuracy is evaluated using the Mean Square Error (MSE) of the Probability Density Function (PDF) and Complementary Cumulative Distribution Function (CCDF), as listed in Table 1. Based on the fitting results, local optimal test statistics are derived for Rayleigh and Weibull sea clutter. Because the two radar observations are independent, the joint likelihood ratio is given by the product of their individual likelihood ratios. The fusion test statistic is formulated as a weighted sum of the two local test statistics. The fusion weights are adaptively determined from the estimated SCRs, with larger weights assigned to radars with higher estimated SCRs. Closed-form analytical expressions relating the false alarm probability and detection probability to the decision threshold are also derived.  Results and Discussions  Monte Carlo simulations with 105 independent trials verify the derived closed-form expressions. Under identical SCRs for the two radars, the simulated detection curves of the individual radars and NP-AFD agree well with the theoretical curves (Fig. 3). Compared with decision-level OR and AND fusion, NP-AFD consistently achieves the highest detection probability, followed by OR fusion, whereas AND fusion performs worst. Performance gain analysis shows that NP-AFD maintains a positive performance gain over the better-performing single radar across all tested Weibull shape parameters. In contrast, OR fusion exhibits negative performance gain at low SCRs when the Weibull shape parameter is small, whereas AND fusion maintains a negative performance gain across the full SCR range (Fig. 4). The simulated false alarm probability is well controlled around the preset value of 10–3 (Fig. 5). Furthermore, a two-dimensional joint evaluation is performed by independently varying the SCRs of the two radars, and detection probability contours are plotted with the SCRs of the two radars as the coordinate axes (Fig. 6). The results show that NP-AFD requires lower SCR combinations than OR and AND fusion to achieve the same detection probability. Experiments with measured sea clutter data further demonstrate that NP-AFD achieves the highest detection probability among all compared methods (Figs. 7 and 8). Its detection performance agrees well with the theoretical results (Figs. 9 and 10).  Conclusions  The challenges posed by non-identical clutter distributions and different SCRs in multi-radar collaborative detection are addressed. Based on the Neyman-Pearson criterion, a dual-radar adaptive fusion detection method, NP-AFD, is derived for Rayleigh and Weibull sea clutter. Closed-form expressions are obtained for the fusion weights, decision threshold, and detection probability. Theoretical derivations, simulations, and experiments with measured sea clutter data consistently demonstrate that NP-AFD outperforms OR fusion and AND fusion under arbitrary SCR combinations, providing superior detection performance and robustness to different SCR combinations. The method, however, remains sensitive to non-uniform sea clutter, as reflected by degraded false alarm control in the presence of sea spikes.
Resource Allocation for Multi-UAV Relay Networks in 6G Semantic Communication
XIAO Liming, GUAN Zheng, LIU Jie, YU Jihong, CHEN Liyuan
Available online  , doi: 10.11999/JEIT260520
Abstract:
  Objective  Sixth-Generation (6G) mobile networks aim to achieve global seamless coverage through space-air-ground integrated architectures. In this context, Unmanned Aerial Vehicles (UAVs) serve as mobile aerial relay nodes to support massive ground-user access in complex environments. However, traditional data-oriented communication paradigms incur substantial bandwidth overhead, limiting their applicability in spectrum-constrained UAV networks. In addition, conventional centralized resource allocation methods are difficult to implement in real time because of highly dynamic network topologies and the strong coupling among multidimensional resources. To address these challenges, semantic communication has emerged as a communication paradigm that extracts semantic information at the transmitter and reconstructs it at the receiver, thereby reducing redundant data transmission. Existing semantic-driven resource allocation methods, however, primarily focus on static terrestrial networks or single-UAV scenarios and do not adequately address coverage limitations and co-channel interference in multi-UAV relay networks. Therefore, a joint resource optimization model and a distributed resource allocation framework are proposed to improve semantic transmission efficiency and long-term user fairness through the joint optimization of multidimensional resources, thereby supporting intelligent resource scheduling in future 6G integrated networks.  Methods  A joint resource optimization model is formulated for multi-UAV relay networks under the semantic communication paradigm (Fig. 1), jointly optimizing the number of semantic symbols, UAV trajectories, power control, and channel allocation. To evaluate semantic communication performance, a Semantic Communication Quality of Service (SC-QoS) metric is proposed by combining Semantic Quantization Efficiency (SQE) with normalized transmission delay. The optimization objective is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem that maximizes the weighted sum of system-wide SC-QoS and long-term user fairness measured by Jain’s fairness index. To solve this problem, a Two-Stage Hybrid Reinforcement Learning (TS-HRL) framework is proposed (Fig. 2). In the first stage, a Capacity-Aware K-means (CA-K-means) algorithm performs heuristic UAV pre-deployment. By introducing a dynamic distance compensation term based on residual capacity, edge users are guided toward lightly loaded UAVs, thereby achieving load balancing while preserving spatial proximity. In the second stage, the dynamic scheduling problem is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and solved using a Recurrent Independent Proximal Policy Optimization with Parameter Sharing (R-IPPO-PS) algorithm (Algorithm 1). The algorithm employs Long Short-Term Memory (LSTM) networks to aggregate historical observations and action trajectories, enabling the inference of hidden environmental states. Furthermore, the parameter-sharing mechanism improves training efficiency as the number of agents increases, while the association preference decoupling strategy transforms discrete channel allocation decisions into continuous association preference variables to facilitate policy optimization.  Results and Discussions  The proposed TS-HRL framework is evaluated in a dynamic environment with randomly moving ground users. Convergence analysis shows that the proposed method achieves a higher initial reward and converges with fewer iterations than the random deployment, memoryless resource allocation, and Enhanced Independent Soft Actor-Critic (EI-SAC) baselines (Fig. 3). By combining LSTM-based temporal modeling with CA-K-means pre-deployment, the proposed framework reduces ineffective exploration and improves convergence stability. Compared with conventional bit-based communication, the semantic communication framework increases Semantic Spectral Efficiency (S-SE) by 3.6-fold, reduces transmission delay by 92.1%, and improves fairness by 14.3% (Fig. 4). Within the semantic communication framework, compared with the memoryless and heuristic schemes, the proposed method improves S-SE by 92.7% and 51.9%, reduces transmission delay by 13.5% and 57.7%, and improves fairness by 17.3% and 39.7%, respectively. Although the EI-SAC baseline achieves a fairness index of 0.93, its S-SE remains relatively low. In contrast, the proposed TS-HRL framework maintains a high level of fairness while achieving an S-SE approximately 4.3 times that of EI-SAC. Compared with random deployment, the proposed method improves S-SE by 11.3% with only a 1.1% decrease in fairness, demonstrating a better balance between transmission efficiency and fairness. As the number of users increases from 10 to 40, most baseline methods exhibit decreases in S-SE and fairness because of intensified co-channel interference and spectrum limitations (Fig. 5). In contrast, the adaptive scheduling strategy and global fairness reward mechanism mitigate performance degradation and maintain the average transmission delay below 0.1 ms. These results demonstrate that the proposed method improves overall system performance while ensuring long-term user fairness.  Conclusions  This paper investigates joint resource allocation and trajectory optimization for dynamic 6G multi-UAV relay networks under the semantic communication paradigm. A joint optimization model that couples the number of semantic symbols, UAV trajectories, power control, and channel allocation is formulated to maximize SC-QoS and long-term user fairness. To address the high-dimensional coupling of the optimization problem, a TS-HRL framework integrating CA-K-means pre-deployment with the R-IPPO-PS algorithm is proposed for multidimensional resource scheduling under partial observability. Simulation results demonstrate the convergence, stability, and scalability of the proposed method under different user densities. Through distributed multi-agent cooperation among UAVs, the proposed method improves semantic transmission efficiency and reduces transmission delay while maintaining long-term service fairness for ground users. These findings provide an effective approach to intelligent resource allocation for UAV-assisted semantic communication in future space-air-ground integrated networks.
Space-Time-Coding Metasurface-Enabled Integrated Design of Radar Communication and Electromagnetic Stealth
ZHANG Ming, WANG Zhe, WANG Boya, YANG Lin, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260536
Abstract:
  Objective  To address the complexity and limited modulation capabilities of existing reconfigurable metasurfaces that integrate radiation and stealth functions, a space-time-coding metasurface is proposed for dynamic switching between beam scanning and Radar Cross Section (RCS) reduction. By periodically modulating the states of the meta-atoms in the time domain, differentiated phase distributions are generated for the incident wave and its harmonics. This enables integrated radiation and electromagnetic stealth without complex Transmit-Receive (T/R) components or multilayer structures. The proposed design provides a simple and low-cost approach for integrated radar communication and electromagnetic stealth.  Methods  The metasurface adopts a metal-dielectric-metal structure, with each meta-atom integrating a PIN diode for 1-bit reflection-phase modulation. Electromagnetic simulations are performed using CST Microwave Studio. At the center frequency of 10.0 GHz, the two diode states provide a reflection phase difference close to 180°, with a near-lossless co-polarized reflection amplitude. Binary Particle Swarm Optimization (BPSO) is used to optimize the space-time-coding sequences for beam scanning and RCS reduction. An 8 × 8 prototype is fabricated using Printed Circuit Board (PCB) technology. A Vector Network Analyzer (VNA) equipped with an S97082A option is used to measure the radiation and scattering characteristics of the prototype.  Results and Discussions  Simulations and measurements confirm that the proposed space-time-coding metasurface provides beam scanning and RCS reduction. In the radiation mode, the fundamental, +1st, +2nd, +3rd, –1st, –2nd, and –3rd harmonics are steered to +2°, +14°, +32°, +46°, –15°, –28°, and –44°, respectively. The average error between the measured and target beam directions is 1.3°. The average sidelobe levels of the ±1st, ±2nd, and fundamental harmonics are 10.59 dB lower than the corresponding main lobes. In the scattering mode, the peak echo gain at 10.0 GHz is reduced by 11.56 dB relative to a copper plate. Over 9.9~10.1 GHz, the echo gain is reduced by approximately 10 dB, except at 9.93 GHz and 10.04 GHz, where the reductions are 8.07 dB and 7.89 dB, respectively. The maximum reduction reaches 14.83 dB at 9.95 GHz. These results verify the integrated radiation and scattering-control capabilities of the proposed metasurface.  Conclusions  A space-time-coding metasurface is proposed for integrated radiation and electromagnetic stealth. By integrating PIN diodes into the meta-atoms, the reflection states are periodically modulated on a planar metasurface to generate differentiated equivalent phase distributions at different harmonic frequencies. This enables multi-angle beam scanning without T/R components and reduces the echo gain through spatial and temporal coding. The proposed design simplifies the system architecture and reduces hardware requirements. The demonstrated beam scanning and RCS reduction indicate its potential for integrated radar communication and electromagnetic stealth.
Communication Performance and Fault Degradation Analysis of Boundary-interface Configurations in 2.5D Chiplet Systems
HOU Shuaikang, LIU Qinrang, LV Ping, LIU Zhengyu, XU Yuhang, LI Peijie, GUO Wei
Available online  , doi: 10.11999/JEIT260633
Abstract:
  Objective  Chiplet-based integration provides an important approach for constructing large-scale heterogeneous systems. In a 2.5D chiplet system, inter-chiplet packets travel from a source node to a boundary interface in the source chiplet, traverse an interposer network, and enter the destination chiplet through a destination-side interface. Because packaging resources, micro-bump count, routing resources, and power and area budgets are limited, interfaces can be deployed only at a subset of boundary nodes. Their number and spatial distribution affect access distance, service-region formation, interface-load distribution, and traffic remapping after failures. Existing studies have addressed die-to-die standards, interposer architectures, placement, routing, simulation, and fault tolerance. However, the independent structural effects of boundary-interface configuration remain insufficiently characterized. This study investigates how interface count, location, service-region partitioning, and failures affect communication performance and fault-induced degradation in 2.5D chiplet–interposer networks.  Methods  A graph-level structural model is developed for a 2.5D chiplet–interposer network. Each chiplet is modeled as an n × n mesh, and selected boundary nodes serve as active interfaces. Each node is mapped to its nearest active interface, and nodes mapped to the same interface form a service region. Average interface-access distance, abstract end-to-end hop count, squared coefficient of variation of interface load, maximum load ratio, and fault-induced degradation are evaluated. Three chiplet scales, n = 4, 8, and 16, are considered, with corresponding interface budgets of k = 2, 4, 6; k = 4, 8, 12; and k = 8, 16, 24, respectively. Five layouts are analyzed: Uniform-corner, Balanced-edge, Two-edge-centered, Single-edge-clustered, and Greedy-DL. Uniform, Hotspot, Transpose, Tornado, and Neighbor traffic patterns are considered. For fault analysis, each active interface is removed in turn, and affected nodes are remapped to their nearest healthy interfaces. A cycle-accurate gem5 model with Ruby/GARNET, comprising four 4 × 4 mesh chiplets and a 4 × 4 interposer mesh, is constructed. Uniform and Hotspot traffic are evaluated at injection rates from 0.01 to 0.15 flit/node/cycle. The latency saturation injection rate is defined as the first sampled rate at which the average packet latency exceeds 50 cycle. This two-level evaluation separates structural effects from cycle-level network behavior and enables interface count, placement, traffic pattern, and fault location to be compared under consistent topology and routing assumptions.  Results and Discussions  Graph-level results show that increasing the number of boundary interfaces reduces average interface-access distance and abstract end-to-end hop count, whereas the marginal benefit decreases as boundary coverage becomes sufficient (Fig. 3). Under a fixed interface budget, different interface locations induce different service-region partitions and interface-load distributions (Table 2, Fig. 4). For an 8 × 8 chiplet with k = 8 under Uniform traffic, Greedy-DL achieves the smallest average interface-access distance of 3.125, whereas Balanced-edge provides a favorable compromise, with an access distance of 3.500 and a maximum load ratio of 1.250. Single-edge-clustered produces the largest access distance and abstract end-to-end hop count because its interfaces are concentrated along one boundary. Under Hotspot traffic, Two-edge-centered reduces the squared coefficient of variation of interface load to 0.016, but its access distance remains higher than those of Balanced-edge and Greedy-DL, indicating a trade-off between distance and load balance (Table 2, Fig. 5). Interface failures cause only moderate increases in average access distance but substantial load concentration at the remaining healthy interfaces. Under Hotspot traffic, the worst-case maximum-load-ratio degradation is 23.94% for Balanced-edge and 93.75% for both Two-edge-centered and Single-edge-clustered (Table 3), indicating that fault sensitivity is mainly associated with service-region migration and load reconcentration. The gem5 results further support the graph-level observations. When the interface count increases from k = 1 to k = 4, low-load latency decreases from 24.50/24.44 cycle to 18.05/18.09 cycle under Uniform and Hotspot traffic, respectively, whereas the latency saturation injection rate increases from 0.06 to 0.14 (Fig. 6). Across the k = 4 layouts, Balanced-edge achieves low latency and hop count, whereas Single-edge-clustered has the highest baseline communication cost. Under Hotspot traffic, Balanced-edge reduces low-load latency by about 14.3% and average hop count by about 18.7% compared with Single-edge-clustered. Greedy-DL provides short communication paths under low load, but its latency saturation injection rate is 0.13, lower than the 0.14 achieved by Balanced-edge and several other layouts. This result indicates that minimizing access distance alone does not guarantee better medium- and high-load performance (Fig. 7). For fault validation, all 16 single-interface failure scenarios are evaluated for each of the three layouts, yielding 48 scenarios in total. The results show increased latency and hop count, together with an earlier onset of congestion. The worst-case latency saturation injection rate decreases to 0.10~0.11, and Balanced-edge exhibits smaller average and worst-case degradation than the more concentrated layouts (Fig. 8, Table 5).  Conclusions  Boundary-interface configuration is a key structural parameter in 2.5D chiplet interconnect design. It affects interface-access distance, service-region formation, interface-load distribution, and fault-induced traffic remapping. Increasing the interface count improves communication efficiency, but the benefit diminishes as boundary coverage increases. Under the same interface budget, interface placement determines whether traffic remains balanced or becomes concentrated after mapping and remapping. The graph-level metrics support low-cost structural screening and mechanism analysis, whereas gem5/Ruby GARNET simulations provide cycle-accurate validation. Therefore, interface count, location, load balance, and fault degradation should be evaluated jointly in 2.5D chiplet–interposer network design.
Efficient Hyperdimensional Computing Accelerator Design for Chinese Text Classification
YU Tianyang, WU Bi, LIU Weiqiang
Available online  , doi: 10.11999/JEIT260556
Abstract:
  Objective  With the proliferation of edge computing in the Internet of Things (IoT), smart wearables, and offline terminals, low-latency and privacy-preserving Chinese text analysis on local devices has become a core requirement. Although neural network-based and Transformer-based language models achieve high accuracy, their large parameter sizes and computational demands make them difficult to deploy on power- and storage-constrained edge devices. HyperDimensional Computing (HDC), an emerging brain-inspired computing paradigm, represents text using 2 k~10 k-dimensional hypervectors and employs lightweight encoding and querying mechanisms instead of complex multilayer networks, providing a hardware-friendly approach for edge-side text processing. However, existing HDC studies have mainly focused on alphabetic writing systems such as English, in which a limited set of letters is sufficient for N-gram encoding. For Chinese, which contains thousands of commonly used characters, directly applying these methods would require a large number of base hypervectors, resulting in substantial storage overhead and weakening the lightweight advantage of HDC. This paper aims to address this limitation by proposing an efficient character encoding method and a dedicated HDC framework with hardware acceleration for Chinese text classification.  Methods  Based on the glyph structure of Chinese characters, a hyperdimensional encoding method based on character glyph structure is proposed. Specifically, the encoding process of the Wubi input method is reverse-engineered to decompose each Chinese character into an equivalent Wubi letter sequence for efficient hyperdimensional encoding. A retrieval library covering 3 500 commonly used Chinese characters, as defined in the General Standard Chinese Characters Table issued by the Ministry of Education of the People’s Republic of China, is constructed. For each character, the Wubi library is queried to obtain a letter sequence of length 3 or 4, with each letter mapped to a base hypervector. Cyclic shift and binding operations are then applied to the base hypervectors of consecutive letters. This produces a character-level hypervector that captures both letter identity and positional order information, thereby avoiding confusion caused by different letter permutations. Compared with directly assigning an independent hypervector to each Chinese character, the proposed method reduces the storage requirement for base hypervectors by more than 95% (Table 1). Building on this character encoding method, the HDChinese framework is developed to support inference, training, and retraining. During inference, the sentence-level hypervector obtained by bundling all character hypervectors queries the class hypervectors using cosine similarity. Because the hypervectors are binary, the inner-product operation is efficiently implemented using XNOR logic. During training, class hypervectors are generated by bundling sentence hypervectors belonging to the same class and applying sign-based binarization. During retraining, the nonbinary class hypervectors are iteratively updated using misclassified samples with a learning rate of 0.25, thereby reducing the effect of outlier hypervectors on the cluster center. Furthermore, a dedicated hardware accelerator architecture is designed for HDChinese. The architecture comprises three main modules: an encoding module, a querying module, and a class hypervector update module. These modules can be selectively activated to support inference, training, and retraining modes, respectively (Fig. 2). The accelerator is prototyped on an Ultra96v2 FPGA development board equipped with a Xilinx ZYNQ System on Chip (SoC).  Results and Discussions  Four open-source Chinese text classification datasets covering binary and multiclass tasks are used for evaluation (Table 2). The Wubi-based encoding method achieves higher accuracy than the Pinyin-based method on all four datasets and reduces the storage requirement by 31.46% (Table 4). Ignoring uncommon Chinese characters, which occur at an average frequency of less than 0.5%, has a negligible effect on accuracy (Table 4). A hypervector dimension of 2 k is selected because increasing the dimension beyond 2 k provides no significant accuracy improvement (Fig. 3). FPGA measurements show that HDChinese achieves a model size of approximately 16 kB, representing a reduction of more than 99% compared with k-Nearest Neighbor (kNN), Support Vector Machine (SVM), and random forest models (Table 7). Compared with SVM and random forest, the proposed accelerator reduces inference time by 6.67%~47.25% and total training time by 4.91%~77.36%. After retraining, the classification accuracy reaches levels comparable to those of the comparison methods (Fig. 4, Table 7). The kB-scale model also enables operation without external memory chips, whereas the comparison methods require more than 400 MB of runtime memory. At 100 MHz, the FPGA implementation consumes 0.278 W, providing the most favorable trade-off between throughput and power consumption (Table 6).  Conclusions  A hyperdimensional encoding method based on Chinese character glyph structure is proposed to address the incompatibility between existing HDC methods and Chinese text. The HDChinese framework and its dedicated hardware accelerator reduce model complexity by more than 99% while maintaining competitive classification accuracy. Training time is reduced by 4.91%~77.36%, and inference time is reduced by 6.67%~47.25%. The proposed approach achieves a favorable balance between accuracy and computational and storage overhead, providing a hardware solution for Chinese text analysis on edge devices.
SkipSync: Accelerating Instruction Sampling for LLM Workloads
CAI Luoshan, ZHOU Yaoyang, WANG Kaifan, LIU Tianyi, SUN Ninghui, BAO Yungang
Available online  , doi: 10.11999/JEIT260397
Abstract:
  Objective   The rapid evolution of Large Language Models (LLMs) has increased the demand for efficient design and evaluation of domain-specific accelerators. Sampling-based performance evaluation methods reduce the cost of cycle-accurate simulation but rely on Instruction Set Simulators (ISS) to execute complete workloads for profiling and checkpoint generation. In LLM inference scenarios, frequent changes in model architectures, inference frameworks, and accelerator instruction sets prevent checkpoint reuse and substantially increase the overhead of ISS-based functional simulation. This overhead has become a major bottleneck in the evaluation workflow. Existing ISS acceleration methods, such as hardware virtualization and Dynamic Binary Translation (DBT), either require target and host systems to share the same Instruction Set Architecture (ISA) or have high implementation complexity, making them difficult to apply to rapidly evolving LLM workloads. Therefore, a flexible and efficient ISS acceleration method is needed to support fast and accurate sampling-based performance evaluation of LLM accelerators.  Methods   An ISS acceleration method, SkipSync, is proposed for instruction sampling of LLM inference workloads. The key observation is that the control flow of most LLM operators is independent of runtime tensor values and is determined by static parameters. Based on this observation, the Skip mechanism is introduced to bypass the execution stage of accelerator instructions within selected operators while preserving instruction fetch and decode to maintain sampling accuracy. To ensure correct subsequent execution, a lightweight Host-Simulator Synchronization mechanism is further introduced to synchronize host-computed results back to the ISS. Lightweight synchronization primitives and custom instructions are designed to integrate SkipSync into existing sampling-based performance evaluation workflows with minimal implementation effort.  Results and Discussions   SkipSync is implemented on QEMU and supports RISC-V vector and matrix extensions as accelerator instruction sets. Experimental results show that SkipSync substantially reduces ISS functional simulation overhead while preserving sampling accuracy. Compared with the baseline, average speedups of 4.64× and 5.51× are achieved in the Prefill and Decode stages, respectively (Fig. 5), by eliminating the dominant execution cost of vector and matrix instructions, which accounts for 78.67% of runtime in the Prefill stage and 81.95% in the Decode stage (Table 1). SkipSync outperforms SIMD_DBT (Fig. 6) and provides better support for newly introduced accelerator instructions. The synchronization mechanism introduces limited overhead, with an average runtime increase of only 8.5% relative to ideal Skip-only execution (Fig. 8). Thus, synchronization does not negate the performance gains from the Skip mechanism. SkipSync also maintains high sampling accuracy, with an average sampling error of 2.55%, and the sampling results are nearly identical to those of baseline QEMU execution (Fig. 9).  Conclusions  An extensible ISS acceleration method, SkipSync, is presented for sampling-based performance evaluation of LLM inference workloads. By combining the Skip mechanism with Host-Simulator Synchronization, SkipSync effectively alleviates the ISS bottleneck in LLM sampling workflows while preserving execution correctness and sampling accuracy. Experimental results demonstrate substantial speedups, low synchronization overhead, and negligible effects on sampling accuracy. The proposed method provides a practical solution for efficient accelerator evaluation in rapidly evolving LLM inference systems. Future work will extend SkipSync to more complex LLM inference optimization scenarios and explore automatic identification of skippable operators and insertion of synchronization primitives to further improve usability.
Reynolds Decomposition Motion-Guided Texture Learning for Scarred Myocardium Phenotyping
RUAN Dongsheng, YANG Daiguo, ZHANG Xiaolin, MA Jianhua, JIANG Mingfeng, WANG Yaming
Available online  , doi: 10.11999/JEIT260330
Abstract:
  Objective  Scarred myocardium is a key imaging marker of multiple cardiovascular diseases, and accurate phenotyping is clinically valuable for treatment planning and prognosis. Cine-MRI noninvasively provides both cardiac motion and anatomical texture information. However, existing classification methods have two major limitations. First, motion representation is often overly discretized and insensitive to subtle abnormalities. Second, multimodal fusion commonly relies on simple feature concatenation, which limits the ability to capture the deep pathological association between motion impairment and texture alteration. A motion-guided texture learning framework is therefore developed to improve the accuracy of non-invasive scarred myocardium classification.  Methods  A Motion-Guided Texture fusion Network (MGTNet) is proposed to enable deep interaction between motion and texture features. First, inspired by Reynolds decomposition in fluid dynamics, a Reynolds Decomposition Motion Network (RDMNet) is designed within a diffeomorphic registration framework to decompose the myocardial motion field into a regular mean motion component and an abnormal pulsatile motion component. This decomposition provides a refined motion representation that is sensitive to subtle abnormalities. Second, a Motion-Guided Cross-Attention (MGCA) module is designed, in which refined motion features serve as query features to dynamically modulate texture features and enhance the perception of suspected lesion regions. In addition, an inter-frame motion interaction module is used to capture temporal dependencies across cardiac frames, while an intra-frame texture extraction module learns multi-scale spatial texture patterns within each frame. Through a serial pipeline of motion field estimation, motion-guided texture enhancement, and feature fusion, end-to-end scarred myocardium classification is achieved.  Results and Discussions  Experiments on the CMRD and ACDC datasets show that MGTNet consistently outperforms existing single-modality and multimodal methods (Tables 1 and 2). It achieves accuracies of 95.6% on CMRD and 97.3% on ACDC, with AUC values of 96.6% and 96.0%, respectively. Compared with the baseline MTNet, MGTNet improves accuracy by up to 1.4 percentage points and the F1-score by up to 2.3 percentage points. Further comparisons with different motion field estimation methods (Tables 3 and 4) show that RDMNet provides more discriminative motion priors and achieves the best overall performance on both datasets. These results indicate that fine-grained motion modeling and motion-guided texture enhancement effectively capture complementary pathological information from Cine-MRI. The ROC curves of the nine methods on both datasets further support the superior classification performance of the proposed method.  Conclusions  A scarred myocardium classification method based on motion-guided texture learning is presented. Reynolds decomposition is incorporated into motion field estimation to separate regular mean motion from abnormal pulsatile motion, and cross-attention is used to guide texture extraction with motion priors. This strategy addresses the limitations of discrete motion representation and shallow multimodal fusion. The results confirm that deep motion-texture interaction improves the accuracy and robustness of non-invasive scarred myocardium classification and provides an effective approach for Cine-MRI-based myocardial phenotyping.
Survey of Satellite Covert Communications: Status, Key Technologies, and Future Challenges
DENG Hao, SUN Weiyuan, ZHU Zhengyu, PAN Gaofeng, SUN Gangcan
Available online  , doi: 10.11999/JEIT260177
Abstract:
  Objective  This survey systematically reviews the theoretical foundations, key technologies, and future challenges of Satellite Covert Communications. Based on the classical Alice-Bob-Willie model, the effects of satellite channel characteristics on covert communication capacity are analyzed to establish the theoretical basis. The network architecture of Satellite Covert Communications under the space-air-ground three-layer framework (Fig. 1) is summarized. Core technologies and optimization methods, including signal camouflage coding, beamforming, spectrum agility, Quantum Key Distribution (QKD), and Artificial Intelligence (AI)-assisted techniques, are systematically reviewed. Major security threats and corresponding multi-layer defense strategies, including Physical-Layer Security (PLS) and intelligent collaborative defense, are also summarized. This survey provides a theoretical foundation and technical guidance for developing highly secure and intelligent Satellite Covert Communications systems.  Significance   The significance of this survey lies in its systematic integration of the theoretical framework and key technologies for Satellite Covert Communications. To address the threats of detection, interference, and eavesdropping in open satellite communication environments, representative space-air-ground integrated architectures reported in the literature are reviewed. These architectures overcome the limitation of conventional encryption techniques, which protect only information content, by reducing statistical distinguishability at the physical layer to achieve a low probability of detection. The constraints imposed by satellite channels on covert communication capacity are clarified through the modified Square Root Law. Enhancement strategies based on adaptive coding, beamforming, spectrum agility, and related techniques are reviewed to establish a comprehensive technical framework for Satellite Covert Communications. These advances provide theoretical support and technical guidance for constructing highly survivable and secure space-air-ground integrated communication networks, with important applications in national defense, emergency communications, and Sixth-Generation (6G) Non-Terrestrial Networks (NTNs).  Progress   Existing studies demonstrate that Doppler spread has a dual effect on covert communication capacity. It increases the missed detection probability while introducing signal distortion, making adaptive coding necessary to maintain reliable transmission. The differences in detection capability among terrestrial, aerial, and orbital wardens (Willie) are quantified (Table 1), providing a theoretical basis for hierarchical defense design. Existing studies have also proposed multi-level covertness enhancement strategies. AI-assisted dynamic camouflage combined with sparse coding exploits background noise, inter-satellite links, and dynamic beamforming to improve covert throughput. At the network level, cooperative Unmanned Aerial Vehicle (UAV)-assisted transmission and dynamic spectrum coordination are identified as representative enhancement approaches (Fig. 3). Furthermore, a hierarchical defense framework is summarized from representative studies. This framework combines Reconfigurable Intelligent Surface (RIS)-assisted signal control, Stackelberg game-based incentives for cooperative jamming, XOR network coding, and federated learning for cross-domain threat feature sharing. These advances provide effective solutions for improving the covertness and security of Satellite Covert Communications.  Conclusions  This survey systematically reviews the theoretical foundations and recent advances in Satellite Covert Communications. The integration of multi-layer satellite constellations, dynamic aerial relay platforms, and software-defined networks supported by Quantum Key Distribution (QKD) enables resilient global covert communication. Extending the Alice-Bob-Willie model to practical satellite channels with non-ideal propagation characteristics provides guidance for covert throughput optimization and secure transmission. AI-assisted coding and waveform design further enable adaptive transmission strategies that respond to dynamic channel conditions. Future research should focus on robust covert transmission in dynamic Low Earth Orbit (LEO) environments, scalable constellation management, and the deep integration of AI and quantum technologies into 6G NTNs. The convergence of programmable satellites, Reconfigurable Intelligent Surfaces (RIS), and advanced machine learning is expected to further improve secure space communications.  Prospects   Future research should focus on four major directions: robust covert transmission under non-ideal channels, AI-assisted intelligent decision-making, integrated 6G NTN networking, and the integration of quantum communication technologies (Fig. 4). High-precision Doppler compensation techniques should be developed to mitigate rapid channel variations in LEO satellite systems. Robust covert transmission schemes should also be developed for imperfect Channel State Information (CSI), with deep reinforcement learning providing real-time resource optimization. Future studies should strengthen the integration of AI and quantum technologies by combining cross-layer QKD with covert transmission protocols and exploiting Software-Defined Satellite (SDS) capabilities for adaptive strategy deployment. Additional opportunities include using RIS to enhance spatial-domain signal control and applying blockchain technology to address trust management in multi-node cooperative networks. Efficient lightweight onboard algorithms and coordinated international frameworks for spectrum and orbital resource management should also be developed to support future Satellite Covert Communications systems.
A Novel TDMOSFET and Its Neural Network Modeling for Ternary Logic Applications
LU Bin, LU Haoran, ZHAO Xiaohong, DI Jiayu, XING Linlin
Available online  , doi: 10.11999/JEIT260413
Abstract:
  Objective  Complementary Metal-Oxide-Semiconductor (CMOS) technology continues to advance toward smaller device dimensions and higher integration. As circuit integration increases, short-channel effects and other phenomena increase leakage current in MOSFET devices, resulting in higher static power consumption. With the rapid development of artificial intelligence, traditional binary logic faces limitations in processing and storing massive amounts of data. Ternary logic has therefore attracted increasing attention because it provides higher information density and lower system complexity than binary logic. However, current ternary logic circuits still face challenges, including the use of multiple components, passive elements, and poor compatibility with conventional CMOS processes.  Methods  A novel Tunneling-Drift-Diffusion Metal-Oxide-Semiconductor Field-Effect Transistor (TDMOSFET) that combines quantum tunneling and drift-diffusion mechanisms is proposed. The device exhibits a constant off-state current characteristic, making it suitable for ternary logic applications. Its operating principle is analyzed, and an Artificial Neural Network (ANN) is used to model its electrical characteristics. The ANN model is trained using Technology Computer-Aided Design (TCAD) simulation data to predict the current-voltage (I-V) and capacitance-voltage (C-V) characteristics. The trained ANN model is further converted into a Verilog-A model and integrated into HSPICE to simulate basic ternary logic circuits, including the Standard Ternary Inverter (STI), Negative Ternary Inverter (NTI), Positive Ternary Inverter (PTI), Ternary NOT-AND gate (T-NAND), and Ternary NOT-OR gate (T-NOR).  Results and Discussions  The trained ANN model accurately predicts the I-V and C-V characteristics of the TDMOSFET. Compared with the TCAD results, the maximum relative errors for the drain current (IDS), gate-drain capacitance (CGD), and gate-source capacitance (CGS) are 39.43%, 5.05%, and 14.19%, respectively, whereas the corresponding average relative errors are 0.46%, 0.69%, and 0.51%. The ANN model is successfully converted into a Verilog-A model and integrated into HSPICE for circuit-level simulation. The STI, NTI, PTI, T-NAND, and T-NOR circuits are successfully simulated. The TDMOSFET-based ternary logic circuits do not require passive elements and are compatible with conventional CMOS processes.  Conclusions  A novel TDMOSFET with dual conduction mechanisms is proposed. When the gate-source voltage is below the turning voltage (Vturn), band-to-band tunneling is the dominant conduction mechanism, and the device operates similarly to a reverse-biased tunneling diode. When the gate-source voltage exceeds Vturn, the drift-diffusion mechanism becomes dominant, and the device exhibits characteristics similar to those of a conventional MOSFET. The proposed device maintains compatibility with conventional CMOS processes, simplifying the manufacturing process and reducing cost and integration complexity. TDMOSFET-based ternary logic circuits realize ternary operation without increasing the number of transistors, using passive elements, or requiring multivalued supply voltages. The proposed TCAD simulation → ANN modeling → Verilog-A modeling → HSPICE simulation framework can also be applied to the study of other emerging semiconductor devices.
An Adaptive Kalman Speech Enhancement Method Driven by Burst Noise Suppression and Dual-time-scale Perception
CHEN Bo, ZHENG ZeRui, SUN Chao, WANG ZheMing, SHEN Ying
Available online  , doi: 10.11999/JEIT260636
Abstract:
  Objective  Traditional Auto-Regressive (AR) Kalman speech enhancement algorithms face three critical challenges under non-stationary noise: limited adaptability of noise model updates, model mismatch caused by burst noise, and a lack of environment-aware covariance adjustment. These limitations degrade enhancement performance and restrict their use in practical speech communication. An improved adaptive Kalman speech enhancement method is therefore proposed to improve speech quality and intelligibility under complex noise conditions.  Methods  First, an environmental deviation measure based on the Energy Entropy Ratio (EER) is constructed to quantify the statistical deviation between the current frame and the background environment. A dual-time-scale EER tracking mechanism is then established to capture instantaneous variations and stable background statistics. Their difference is used to generate an adaptive deviation intensity factor for the joint adjustment of the process-noise and observation-noise covariances. Second, a two-stage burst noise discrimination scheme is developed based on the characteristics of burst noise. Frame energy changes and the Spectral Flatness Measure (SFM) are jointly used for preliminary burst noise detection. A Speech Modeling Metric (SMM) based on the linear prediction residual variance ratio is further used to distinguish speech from weak-colored burst noise and prevent contamination of the speech AR model.  Results and Discussions  Experiments are conducted on the NOIZEUS database. The proposed method outperforms the conventional AR-based Kalman filter and the sensitivity-based Augmented Kalman Filter (AKF) in terms of Short-Time Objective Intelligibility (STOI), Perceptual Evaluation of Speech Quality (PESQ), and Segmental Signal-to-Noise Ratio (SegSNR). It shows improved adaptability and stability under non-stationary and burst noise conditions. Dual-time-scale environmental perception and burst noise discrimination facilitate rapid adaptation to environmental changes and reduce speech distortion.  Conclusions  The proposed burst noise suppression and dual-time-scale adaptive Kalman method addresses the limitations of conventional AR-based Kalman speech enhancement under complex noise conditions. EER-based tracking enables environment-aware covariance adjustment, whereas the two-stage discrimination scheme improves burst noise detection and reduces contamination of the speech AR model. The experimental results demonstrate the robust performance of the proposed method and support its application to speech enhancement under non-stationary noise.
Intelligent Detection of DSSS Signals Under False-Alarm Rate Constraints Based on a Noise Score-Pool Threshold Calibration Mechanism
ZHANG Tao, TANG Xiaomei, SUN Guangfu
Available online  , doi: 10.11999/JEIT260414
Abstract:
  Objective  To address the degradation in Global Navigation Satellite System (GNSS) signal detection performance under weak-signal conditions and the limited control of the probability of false alarm (Pfa) in existing Deep Learning (DL) models, a DSSS signal detection method with a prescribed Pfa is investigated. The method enables the DL detector to be evaluated within a Constant False Alarm Rate (CFAR) framework.  Methods  DSSS signal detection is formulated as a binary classification problem, and a DL-based detection framework is developed. An improved one-dimensional ResNet-18 is designed for I/Q sampled time-series data. The input convolution kernel is set to 1×7, and the initial maximum pooling layer is removed to preserve weak-signal temporal features. The first residual layer is also configured without downsampling. A noise score-pool-based threshold calibration mechanism is developed to impose Pfa constraints on the detection decision. A large number of pure-noise samples are processed by the trained network to obtain the empirical distribution of confidence scores for the signal-present class. The decision threshold is then calibrated according to the quantile corresponding to the preset Pfa. In addition, the effect of signal normalization on detection performance is evaluated. The proposed method is validated using a simulated GPS L1 C/A signal dataset under different Signal-to-Noise Ratio (SNR) and Pfa settings and in non-ideal colored-noise environments.  Results and Discussions  The proposed method achieves high detection performance under different Pfa constraints. At a Pfa of 0.01, the detection probability reaches 100% at an SNR of –8 dB. When the Pfa decreases from 0.01 to 0.001 and 0.000 1, the detection curve shifts toward higher SNRs, but the decrease in detection performance remains limited. The unnormalized preprocessing strategy consistently outperforms Root-Mean-Square (RMS) normalization, providing a performance gain of approximately 1 dB. Compared with the traditional autocorrelation detection method, the proposed DL-based detector provides a detection performance gain of approximately 3~4 dB across the tested Pfa settings. In colored-noise environments not used for network training, the proposed method maintains effective detection performance and demonstrates robustness to noise mismatch. The structural ablation results further show that removing the maximum pooling layer and retaining the temporal resolution of the first residual layer improve detection performance under low-SNR conditions.  Conclusions  A DL-based DSSS signal detection method with noise score-pool-based threshold calibration is proposed. The empirical distribution of pure-noise confidence scores is used to calibrate the decision threshold, thereby incorporating the DL detector into a CFAR-based detection framework. The improved one-dimensional ResNet-18 effectively extracts features from I/Q time-series data, whereas the unnormalized preprocessing strategy preserves useful signal-amplitude information. The proposed method improves detection sensitivity while maintaining effective Pfa control and exhibits robustness under non-ideal colored-noise conditions.
A WiFi Multi-link Collaborative Human Tracking Method for Smart Home
PAN Houcheng, CAI Yushuang, YAO Junmei, ZHANG Tingting
Available online  , doi: 10.11999/JEIT260267
Abstract:
  Objective  WiFi-based indoor passive human tracking has attracted increasing attention for smart home applications because Internet of Things (IoT) devices are widely interconnected through existing WiFi infrastructures. Existing approaches estimate the Angle of Arrival (AoA) or Time of Flight (ToF) of the target-reflected path to achieve tracking. However, these methods are fundamentally limited by the difficulty of separating weak target-reflected signals from dense multipath propagation when commercial WiFi devices provide only a small antenna array and limited bandwidth under existing communication protocols. Therefore, Doppler Frequency Shift (DFS)-based approaches that exploit multiple links have become more practical. Dead Reckoning (DR) is widely adopted in these methods, but dynamic link selection for signal fusion remains challenging. Furthermore, the performance of DR-based methods degrades substantially when device-location errors are present. Existing methods also generally employ a single-transmitter-multi-receiver architecture, in which Access Point (AP) broadcasts downlink WiFi signals to STAtions (STA). It complicates data aggregation and limits the use of multiple receive antennas of AP’s. To address the challenges of multi-link signal fusion and device-location uncertainty, a Particle Filter (PF)-based method is proposed. Furthermore, a multi-transmitter-single-receiver architecture is adopted, in which multiple STAs transmit uplink packets to a single AP. This architecture simplifies data aggregation while exploiting the AP’s multi-antenna capability.  Methods  In the IEEE 802.11 standard, the wireless channel is estimated at the receiver using pilot signals embedded in WiFi packets, and Channel State Information (CSI) is continuously obtained. Human motion perturbs multipath propagation and produces time-varying changes in CSI. Therefore, CSI serves as the primary sensing signal because it implicitly captures target motion. In practice, raw CSI is affected by amplitude and phase impairments. Accordingly, signal preprocessing is first performed before Doppler extraction. Subsequently, multi-link DFS measurements are fused using PF, in which the target state is sequentially propagated and updated according to likelihoods derived from DFS measurements. This probabilistic framework naturally enables dynamic multi-link fusion because unreliable links receive lower weights rather than being deterministically discarded. Moreover, the effects of device-location errors are incorporated into the measurement-noise model, improving robustness to device-location uncertainty. When prior knowledge of device locations is unavailable, device self-localization is first performed by jointly estimating the Line-of-Sight (LoS) AoA and ToF between the AP and each STA. Preliminary device self-localization experiments (Fig. 5) demonstrate a median STA localization error of approximately 0.56 m (Table 1), providing reliable initialization for the subsequent tracking algorithm.  Results and Discussions  A prototype system is implemented using a multi-transmitter-single-receiver architecture with commercial Intel AX200/AX201 WiFi cards. CSI is collected over a 20 MHz channel centered at 5.24 GHz using PicoScenes, with IEEE 802.11ac packets transmitted at 100 or 200 Hz. Each CSI sample contains measurements from 57 subcarriers. Meanwhile, Fine Time Measurement (FTM) is performed on Channel 11 at 2.4 GHz using iw and hostapd. In the prototype system, each STA transmitter is equipped with a single antenna, while the AP receiver employs two antennas. Experiments are conducted using one AP and three or four STAs (Fig. 7). WiTraj and PITrack, both based on DR, are used as baseline methods. When device self-localization is required, the proposed method achieves a median tracking error of 0.47 m, compared with 1.78 m for WiTraj and 1.77 m for PITrack (Fig. 9), representing an accuracy improvement of approximately 70% over conventional DR-based methods. Error-injection experiments with random device-location perturbations show that, unlike conventional DR-based methods, which are highly sensitive to device-location errors, the proposed method remains robust to device-location uncertainty (Fig. 10). Additionally, computational complexity analysis shows that, although execution time increases with the particle count (Table 3), tracking performance reaches a stable level beyond a moderate particle count (Fig. 11). Therefore, real-time operation can be achieved by selecting an appropriate particle count. Finally, experiments in complex environments demonstrate that the proposed method consistently achieves higher tracking accuracy than the baseline methods (Fig. 12).  Conclusions  A PF-based method is proposed for multi-link collaborative passive human tracking using commercial WiFi devices. By probabilistically fusing measurements from multiple links, robust tracking is achieved even when device-location errors are present. When device locations are unavailable, device self-localization is achieved by combining CSI-based estimation with FTM measurements, providing the geometric information required for tracking. A prototype system based on a multi-transmitter-single-receiver architecture is developed to simplify data aggregation while exploiting the AP’s multi-antenna capability. Experimental results demonstrate that the proposed method achieves sub-0.5 m median tracking error when device-location errors are present, representing an improvement of approximately 70% over conventional DR-based methods. The proposed method also exhibits strong robustness to device-location uncertainty and supports real-time implementation when an appropriate particle count is selected. Future work will extend the framework to multi-person tracking and evaluate its performance under more realistic deployment conditions.
Cross-Domain Collaborative Enhancement for Tiny Object Detection in Remote Sensing Images
ZHANG Tianyang, ZHANG Xiangrong, WANG Guanchun, TANG Xu
Available online  , doi: 10.11999/JEIT260317
Abstract:
  Objective  Deep learning has substantially advanced object detection in Remote Sensing Images (RSIs). However, because of imaging conditions and the inherently small size of many objects, a large proportion of targets in RSIs occupy fewer than 16 × 16 pixels. Therefore, current object detection methods achieve substantially lower detection accuracy for tiny objects than for normal-scale objects. This limitation primarily arises from two critical factors: insufficient positive sample assignment and weak feature representation. To address these challenges, a Cross-Domain Collaborative Enhancement Detector (CDCEDet) is proposed. CDCEDet jointly optimizes label assignment in the spatial domain and enhances feature representation in the frequency domain, thereby improving the accuracy and robustness of tiny object detection in RSIs.  Methods  The overall framework of CDCEDet is illustrated in Fig. 2 and consists of three major components. First, a Scale-Adaptive Anchor Generator (SAAG) is designed to dynamically generate anchors that match the scales of ground-truth (GT) objects, thereby effectively alleviating the scale mismatch between anchors and tiny objects that has been largely overlooked in previous studies. Compared with conventional uniformly distributed anchor generators, SAAG substantially increases the number of positive samples assigned to tiny objects, even under Intersection over Union (IoU)-based label assignment. Second, a Quantile-based Adaptive Label Assignment (QALA) mechanism is developed to replace the conventional fixed IoU threshold-based label assignment. QALA models the IoU distribution between each GT object and its matched anchors to generate an adaptive label assignment threshold for each object, thereby further increasing the number of positive samples assigned to tiny objects. Third, a Frequency-Adaptive Fusion (FAF) module is developed to enhance feature representation from a frequency-domain perspective. An adaptive high-pass filter is used to strengthen high-frequency details and compensate for information loss caused by channel compression, whereas an adaptive low-pass filter preserves semantic consistency during feature upsampling, thereby reducing semantic inconsistency within upsampled objects.  Results and Discussions  Extensive experiments are conducted on two public remote sensing tiny object detection datasets, AI-TODv2 and AI-TOD-R. The proposed method is compared with several state-of-the-art methods, including RFLA, DCNet, and DCFL. On the AI-TODv2 dataset (Table 1), CDCEDet improves AP50 and AP50–95 by 1.8% and 0.7%, respectively, compared with the best existing method. On the AI-TOD-R dataset (Table 2), AP50 and AP50–95 are improved by 2.6% and 0.7%, respectively. These results demonstrate that CDCEDet achieves superior detection performance and strong generalization capability for tiny object detection in RSIs. Ablation studies and parameter analyses of the proposed modules (Tables 36) further verify the effectiveness of each component and their complementary contributions. Qualitative results on both datasets (Fig. 3) show that the proposed method accurately detects tiny objects in both sparse and dense scenes. As illustrated in Fig. 4, SAAG generates scale-matched anchors for individual objects and assigns substantially more positive samples to tiny objects than the conventional uniformly distributed anchor generator. Furthermore, visual comparisons with RFLA and DCNet (Fig. 5) demonstrate that CDCEDet achieves higher detection accuracy while substantially reducing missed detections.  Conclusions  A CDCEDet is proposed to address insufficient positive sample assignment and weak feature representation in remote sensing tiny object detection. Specifically, SAAG dynamically generates anchors that match the scales of GT objects, substantially increasing the number of positive samples assigned to tiny objects. QALA further improves label assignment by modeling the IoU distribution between GT objects and their matched anchors to adaptively determine the label assignment threshold, thereby effectively reducing the scale bias introduced by fixed IoU thresholds. In addition, FAF enhances feature representation from a frequency-domain perspective through an adaptive high-pass filter and an adaptive low-pass filter. Experimental results on two benchmark datasets demonstrate the superior detection performance and strong generalization capability of CDCEDet. Future work will focus on improving model efficiency and real-time performance to facilitate practical deployment in remote sensing applications.
DroneRFc-MM: Anti-UAV Multimodal Detection Measured Dataset
YU Taosong, YANG Qianqian, HU Zhuo, LI Mingkai, WU Jiajun, SU Yufan, PAN Junyu, SHI Zhiguo, CHEN Jiming
Available online  , doi: 10.11999/JEIT260889
Abstract:
Objective: A comprehensive multimodal benchmark is developed for Anti-Unmanned Aerial Vehicle (UAV) detection in low-altitude urban environments. Existing datasets generally provide limited sensing modalities and UAV models, with relatively coarse annotations that constrain tasks requiring spatial, motion, and cross-modal information. DroneRFc-MM addresses these limitations by providing synchronized multimodal data, broader coverage of consumer-grade DJI UAV models, and fine-grained annotations for target detection, UAV model recognition, trajectory analysis, flight-direction reasoning, and multimodal fusion evaluation. Methods: DroneRFc-MM is synchronously collected using six heterogeneous sensor types: a Pan-Tilt-Zoom (PTZ) camera, a fisheye camera, a Radio Frequency (RF) antenna, LiDAR, millimeter-wave radar, and a microphone array. Data are acquired on an open rooftop at a university in Zhejiang Province, representing a typical urban low-altitude environment. The dataset contains recordings of six consumer-grade DJI UAV models. All devices are synchronized using a common network time reference, with inter-device timestamp discrepancies of approximately 0.3 s. The UAVs fly in “H”-shaped and vertical reciprocating trajectories at distances of 20–60 m from the sensor array. Fine-grained annotations, including UAV model, position, attitude, and velocity, are derived from flight logs. For the flight-direction reasoning task, approximately 5-s multimodal clips are generated, including camera videos, RF spectrogram videos, microphone audio, and coordinate-based text representations of radar point-cloud data. Zero-shot inference is conducted using Qwen 3.6-Plus and Qwen 3.5-Omni-Plus with unified prompts. Prediction accuracy and inference time are evaluated by comparing predicted directions with ground-truth directions calculated from UAV positioning data. Results and Discussions: The DroneRFc-MM dataset provides multimodal data from six sensor types and six consumer-grade DJI UAV models, together with fine-grained annotations and sample extraction tools. In the flight-direction reasoning task, the Qwen-series multimodal large language models (MLLMs) achieve accuracies ranging from 20% to 30% across the different input modalities. The inference time is also relatively long, with the mean response time exceeding 40 s for most sensor inputs. These results indicate that current general-purpose MLLMs can capture weak motion-related information from UAV videos, audio, RF spectrograms, and point-cloud data, but their accuracy and response speed remain insufficient for practical real-time Anti-UAV detection. Conclusions: DroneRFc-MM provides a multimodal benchmark for Anti-UAV detection, UAV model recognition, flight-direction reasoning, and multimodal model evaluation. The dataset integrates six sensor types, six consumer-grade DJI UAV models, and fine-grained annotations within a common measurement framework. The experimental results show that current general-purpose MLLMs remain limited in flight-direction reasoning and real-time inference in Anti-UAV scenarios. Domain-specific pre-training, supervised fine-tuning, knowledge augmentation, and lightweight inference are therefore needed to improve their practical utility. Future work will expand the dataset scale and application scenarios to support intelligent and efficient low-altitude airspace management systems.
Dynamic Data Mapping and Co-Optimization Method for TSVs in 3D-Integrated MoE Accelerators
YANG Jialin, XIA Chenjie, WU Huiming, LI Ningyuan, SONG Yuan, LIU Bo
Available online  , doi: 10.11999/JEIT260565
Abstract:
  Objective  The rapid progress of large-scale intelligent computing, especially Mixture-of-Experts (MoE), has positioned Three-Dimensional Integrated Circuit (3D IC) based on Through-Silicon Via (TSV) as a key solution to memory-wall bottlenecks via high bandwidth and density. As a core 3D IC technology, TSVs enable vertical inter-chip connections, reducing path length, parasitic delays, power, and boosting data rates. MoE-specific accelerators, characterized by high data density and strong fault tolerance, introduce new challenges and opportunities for TSV layout. These include aggravated signal integrity and reliability issues in dense arrays, and the inadequacy of static TSV allocation for dynamic, bursty MoE traffic. Conversely, their inherent fault tolerance permits optimization design spaces for employing fault-tolerance mechanisms. This paper exploits MoE dataflow characteristics and hardware fault tolerance to devise a data mapping strategy for high-density TSV arrays based on fault-tolerance mechanisms, targeting improved performance and reliability.  Methods  This paper investigates cluster partitioning schemes and data mapping strategies to enhance the reliability of TSV data transmission. To address the high complexity of global optimization in large-scale TSV arrays, a cluster size partitioning scheme is proposed. By structurally partitioning a large-scale TSV array into several small-scale TSV clusters, the global optimization problem is decomposed into local, scalable subproblems, thereby improving optimization efficiency and flexibility while ensuring optimization moderation. Through comprehensive consideration of multiple metrics and simulation-based evaluation, the cluster size is finally determined to be 6×6. In response to the varying dataflow characteristics and load distribution across different computational stages, this paper proposes a Phase- and Load-Aware Dynamic Data Mapping (PLDM) strategy. The strategy pre-partitions the TSV array into multiple fixed-size clusters and classifies them into critical clusters and general clusters based on metrics such as coupling strength, bandwidth, and latency. At runtime, the PLDM strategy dynamically adjusts data mapping according to the characteristics of different computational stages. Furthermore, this paper achieves a co-optimization design of PLDM with the encoding circuit. The load monitoring module and the error monitoring module share certain data buffers and control status registers, enabling hardware resource reuse. Meanwhile, the error monitoring results provide real-time feedback on the reliability level of each TSV cluster, based on which the mapping controller preferentially allocates data transmission to clusters with lighter loads and lower bit error rates. This approach realizes resource sharing and load balancing, thereby improving data transmission reliability and link utilization efficiency for high-density TSV arrays.  Results and Discussions  This paper analyzes the bandwidth utilization and load balancing performance of three mapping schemes: random, static, and dynamic. The results show that both the dynamic and random mapping schemes achieve average bandwidth utilization close to the theoretical maximum. However, the random mapping scheme maps approximately 37.52% of critical data into general clusters with relatively high bit error rates, thereby increasing unreliability. Compared with static mapping, the dynamic mapping scheme improves average bandwidth utilization from 0.7982 to 0.8984, a relative increase of about 12.6%, reduces inter-cluster load fluctuation by 54.6%, and correspondingly improves load balancing by a factor of 2.2 (Fig. 4). Compared with random mapping, the dynamic mapping scheme reduces inter-cluster load fluctuation by about 8.6%, and reduces the latency of critical data and non-critical data by 36.6% and 34.7%, respectively (Table 2). To further evaluate the optimization effects of the proposed PLDM strategy on metrics such as load balancing and bandwidth utilization, four comparative schemes are configured: (1) Baseline scheme; (2) Static mapping scheme; (3) Sparse TSV layout using TSV-Aware Adaptive Fault-Tolerant Coding (TSV-AFTC) and PLDM; (4) High-density TSV layout based on scheme (3). Taking the Qwen3-30B-A3B model as an example, the normalized loads of 16 clusters in the Multi-Head Attention (MHA) and Feed-Forward Network (FFN) stages are compared across the four schemes. The results indicate that the proposed dynamic data mapping scheme achieves balanced load distribution across clusters in both the MHA and FFN stages, ranging from 0.48 to 0.52, while ensuring that all critical data are mapped to critical clusters. The high-density TSV scheme further reduces the load per cluster to approximately 0.34–0.37, demonstrating that dynamic mapping can effectively suppress stage-wise hot spots and improve load balancing (Fig. 7). Subsequently, system-level fault injection is applied to the transmitted data to simulate data reliability under extreme conditions for different schemes. The results show that for the proposed scheme (sparse), the degradation in perplexity (PPL) compared to the ideal case is controlled within 0.02, while the average bandwidth utilization is improved by approximately 15% and cluster load balancing is enhanced by a factor of 3.4. Under the high-density scheme, the PPL increase is controlled within 0.05, the average bandwidth utilization reaches about 71.7%, and the cluster load balancing is improved by a factor of 2 (Table 4).  Conclusions  This paper investigates TSV data mapping for 3D MoE accelerators and proposes a PLDM strategy based on TSV-AFTC, which allocates data from different computational stages to reliable and lightly loaded TSV clusters according to cluster-level bit error rates and load conditions. Through circuit co-design, approximately 12% of hardware resources can be saved. Compared with static mapping, the proposed scheme improves average bandwidth utilization by about 12.6% and enhances cluster load balancing by a factor of 2.2. Under system-level fault injection, the scheme limits the degradation of model inference perplexity to within 0.02, while achieving approximately 15% improvement in bandwidth utilization and a 3.4× enhancement in cluster load balancing.
A Dual-Trellis Message-Passing Decoding for Non-Binary LDPC Codes
XX XX
Available online  , doi: 10.11999/JEIT260958
Abstract:
  Objective  Due to their capacity approaching performance, Low Density Parity-Check (LDPC) codes have been widely applied to wireless communication and data storage systems. Compared to their binary counterparts, Non-Binary LDPC (NB-LDPC) codes with short or moderate code lengths have been demonstrated to achieve superior error performance under non-binary Belief Propagation (BP) decoding. However, the computational complexity of Check Node (CN) update of the optimal BP decoding is too complex for practical applications. Recently, many works have been presented to perform updates of CNs based on truncated messages, rather than full-length reliability messages, to significantly reduce the computational complexity of CN updates. Most of them construct the trellis of a CN based on the truncated input vectors, called truncated-trellis, such that CN updates are efficiently processed in parallel based on the selected candidate paths. These paths generally contain only a small number of deviation nodes, and such deviation nodes usually have high reliability. However, the Variable Node (VN) update in most decoding algorithms based on CN truncated-trellis still sequentially processes each element in the input vectors of each VN by the elementary steps. When the CN update is simplified, the complexity of the VN update may primarily determine the overall computational complexity. To address the above issues, this paper proposes the Dual-Trellis Min-Sum (DTMS) decoding algorithm. By further introducing truncated-trellises for VNs and updating the output messages of CNs and VNs in parallel, respectively, it further improves the decoding efficiency, while maintaining the similar decoding performance.  Methods  The different contributions of nodes in the CN truncated-trellis of the Pruning path Min-Sum (PMS) decoding algorithm on the selected highly reliable candidate paths are first analyzed, and it reveals that the selected highly reliable paths are primarily determined by the deviation nodes from the first few rows of the trellis of a CN, especially the second row. Thereby, it is not critical to update and sort every element of each output vector of one VN during the VN update. Next, a new trellis of one VN is constructed, and highly reliable elements over this trellis shared by all the output vectors of this VN are searched using a row-wise pruning strategy, such that the conventional element-wise VN updating procedure is transformed into a trellis-based parallel updating process based on an extra column in the trellis. In this basis, the unequal protection for the reliability values of each VN output vector is conducted, e.g., only the first few elements in each output vector of VN are updated and arranged, and the rest elements of each output vector are directly set to a compensation value. As a result, the computational complexity required for less reliable elements during each VN update can be significantly reduced, while retaining the crucial messages.  Results and Discussions  Experimental results show that compared with the PMS decoding algorithm using the original VN updating procedure, the proposed DTMS decoding algorithm maintains almost the same Bit Error Rate (BER) performance and convergence speed for decoding NB-LDPC codes under different finite fields, code lengths, and code construction methods (Figs. 37). Meanwhile, the number of real-domain operations required for the proposed simplified VN updates is reduced by approximately 71.8% on average (Table 2). In addition, the error-correction performance and convergence speed of the proposed DTMS decoding algorithm are close to those of the sub-optimal BP decoding algorithms (Figs. 37) with relatively low computational complexity (Table 3). The average performance gap of the DTMS decoding algorithm from the optimal BP decoding algorithm is only about 0.11 dB (Figs. 37). Thus, optimizing the VN updating is an effective way to further reduce the decoding complexity of truncated-trellis-based message-passing decoding algorithms.  Conclusions  This paper proposes a DTMS decoding algorithm to reduce the computational complexity of VN update in truncated-trellis-based decoding algorithms for NB-LDPC codes. Based on the CN updating process of the PMS decoding algorithm, the proposed algorithm further constructs a truncated-trellis and introduces the unequal protection scheme for VN update, such that the output vectors of each VN can be efficiently updated in parallel. Experimental results show that, under the same CN trellis-based update, the proposed parallel VN updating method significantly reduces the computational complexity compared with the original VN updating method, while maintaining similar decoding performance. Moreover, the proposed DTMS decoding algorithm performs closely to the sub-optimal BP decoding algorithms with similar convergence speed and lower complexity. In future studies, it will be interesting to further exploit the adaptive pruning strategies for the VN parallel updates. Based on the distribution of field elements from different iterations, less reliable field elements can be adaptively eliminated to reduce the set of candidate field elements, which may further reduce the complexity of VN update with negligible performance loss.
Study on Deployment Optimization of Reconfigurable Intelligent Surface for Troposcatter Communications
ZHAO Ziyan, SONG Zhiqun, LIU Lizhe, LI Yong, LI Xingjian, WANG Bin
Available online  , doi: 10.11999/JEIT260922
Abstract:
  Objective   Troposcatter communication serves as a valuable complement to satellite communication and thus is still quite promising in scenarios such as military long-distance communication. However, when a troposcatter communication system is deployed in mountainous environments, it is often faced with a prevalent and challenging engineering problem known as the “Line-of-Sight (LoS) obstruction”. Traditional solutions to this issue are still confronted with engineering difficulties. Increasing the antenna elevation angle to cross obstacles makes the scattering angle increase sharply and consequently lead to transmission loss surging beyond acceptable link budget limits; alternatively, building tall towers to raise antenna height preserves low-angle transmission but introduces construction difficulties and sacrifices the advantage of terrain concealment. Reconfigurable Intelligent Surface (RIS) has emerged as a disruptive technology in wireless communications, with the capability of reconstructing the wireless environment and artificially altering channel characteristics. It has been successfully applied in various civilian mobile communication systems. Obviously, it also provides an alternative to address the problem of LoS obstruction in troposcatter communications. Unfortunately, it has never been reported that RIS had been applied in such scenarios. Herein, to solve the LoS obstruction problem in troposcatter communications, RIS is involved for the first time in this field, a conceptual architecture of RIS-assisted troposcatter communication is set up, and then the problem of optimal RIS deployment is systematically investigated.  Methods   Based on the proposed framework of RIS-assisted troposcatter communication system, the deployment optimization of RIS is addressed step by step:Firstly, a three-dimensional model of feasible deployment region is established under four types of practical engineering constraints, i.e., the intrinsic constraint of obstacle-crossing, the optimal constraint of engineering upper bound, the antenna radiation constraint of Fresnel near-field region, and the hardware constraint of RIS effective angle.Secondly, the optimization problem of RIS deployment is formulated as minimizing the comprehensive system gain loss. The overall loss consists of three major components, namely, troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles. The first two parts are easily computed according to corresponding engineering knowledge of troposcatter communication and typical antenna theory, respectively. Then to calculate the dynamic gain attenuation of Cassegrain antennas, a quantitative model is developed based on cantilever beam bending theory. The model quantifies the pointing errors of a Cassegrain antenna caused by dynamic over-compensation with its elevation angle adjustment, and then the gain loss is calculated with the assistance of Taylor radiation pattern.Thirdly, through theoretical analysis and numerical verification via sectional slicing heatmaps, a dimensionality reduction property of the objective function is observed and validated. Within the feasible region, the first-order partial derivative of the objective function with respect to deployment height is always negative, implying that the global optimal deployment position necessarily lies on the upper boundary surface of the feasible region. The dimensionality reduction property is rigorously validated through slice analysis across the entire feasible domain, with more than 46,100 verification points confirming that the optimal position always resides on the upper boundary surface. This finding reduces the intractable three-dimensional constrained optimization problem to a much simpler two-dimensional manifold optimization, which significantly reduces the computational complexity of the optimization.Finally, based on this dimensionality reduction property, an improved gradient descent algorithm with momentum and adaptive backtracking line search (IGD-M&ABLS) is put forward. The algorithm introduces momentum gradient updates to suppress zigzag oscillations and accelerate convergence; it also incorporates an adaptive backtracking line search strategy to dynamically adjust step sizes, balancing iterative stability with computational efficiency.  Results and Discussions   A series of simulation experiments are conducted under typical engineering parameters, i.e., a 3-meter aperture Cassegrain antenna; 5 GHz signal frequency; mountain heights of 150 m, 200 m and 250 m, representing medium-high hills, the dividing line between hills and mountains, and relatively mountainous terrain, respectively. The proposed IGD-M&ABLS algorithm is benchmarked against Grid Search (GS, a classic deterministic exhaustive-search method) and Particle Swarm Optimization (PSO, a typical efficient heuristic algorithm). The results demonstrate that IGD-M&ABLS consistently converges to the global optimal solution with less gain losses than both benchmarks. Specifically, for the typical case with a 200 m mountain height, in 50 independent runs, IGD-M&ABLS always achieves the best objective function value of 11.6094 dB, much more steadily than PSO does, and it also outperforms GS's 11.6127 dB. As far as time consumption is concerned, IGD-M&ABLS exhibits remarkable advantages of computational efficiency. Its average runtime is approximately 0.005 seconds, compared with 0.07 seconds of PSO and more than one hour of GS. This order-of-magnitude improvement in computational speed is consistent with the theoretical complexity analysis. IGD-M&ABLS optimizes two independent variables on a 2D manifold, whereas PSO and GS handle three independent variables in the 3D feasible domain. Robustness tests under varying terrain conditions (H = 150 m and H = 250 m) confirm that IGD-M&ABLS reliably obtains the best results across different scenarios. In all simulation tests, IGD-M&ABLS demonstrates excellent stability and reproducibility, producing consistent results across multiple independent runs, while PSO exhibits randomness-induced variations and GS remains limited by its discretization step size.  Conclusions   This paper pioneers the application of RIS technology in troposcatter communication, providing a new technical solution to address the LoS obstruction problem in mountainous environments. A conceptual framework of RIS-assisted troposcatter communication system is established, incorporating a three-dimensional feasible RIS deployment region model with four practical engineering constraints. Then the optimization problem of RIS deployment is formulated as minimizing the overall system gain loss including troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles, and the computation method of its third term is also developed for the first time based on cantilever beam bending theory. More interestingly, the objective function is found and verified with dimensionality reduction property, i.e., its minimum value always resides on the upper boundary manifold surface. That property effectively transforms the complex 3D optimization into an equivalent 2D manifold problem. Finally, a new algorithm IGD-M&ABLS is proposed by introducing momentum and adaptive backtracking line search into the traditional gradient descent framework. Simulation results show that compared with benchmarks, IGD-M&ABLS algorithm achieves the best deployment positions with order-of-magnitude faster computation, while maintaining excellent stability and reproducibility.
Non-Orthogonal PSWFs Signal Detection Method Based on Adaptive Temporal-Spatial Feature Fusion
CHEN Wenhua, MAO Zhongyang, LU Faping, SUN Ye, GAO Yixuan
Available online  , doi: 10.11999/JEIT260024
Abstract:
  Objective   To address the demands of B5G/6G systems for high spectral efficiency and transmission reliability, Prolate Spheroidal Wave Functions (PSWFs)-based non-orthogonal modulation has attracted extensive research interest because of its strong time-frequency energy concentration. However, severe mutual interference among multiplexed PSWF signals degrades the performance of conventional detection methods in complex channel environments. Existing methods are limited by ideal channel assumptions or single-modal feature extraction and therefore cannot fully exploit the temporal and spatial information of PSWF signals or adapt to dynamic interference. An Adaptive Temporal-Spatial Feature Fusion (ATSFF) architecture is proposed for accurate and robust detection of non-orthogonal PSWF signals.  Method   A dual-path parallel framework is constructed to extract complementary temporal and spatial features. A Gated Recurrent Unit (GRU) network extracts deep temporal features and captures long-term dependencies from one-dimensional received signals. In the other path, one-dimensional signals are transformed into two-dimensional representations using the Gramian Angular Difference Field (GADF), and hierarchical spatial features are extracted using ResNet50. An adaptive probability-weighted fusion mechanism dynamically adjusts the contributions of the two feature branches according to their prediction uncertainty, thereby integrating complementary temporal and spatial information and improving detection robustness.  Results and Discussion   Simulations on a 32-class non-orthogonal PSWF signal dataset (Fig. 2) show that the proposed ATSFF method outperforms coherent detection, cross-term detection, Approximate Message Passing-Interleave Division Multiple Access (AMP-IDMA), and Temporal Multiple Sparse Bayesian Learning-Least Squares (TMSBL-LS) over the full Signal-to-Noise Ratio (SNR) range. t-SNE visualization (Fig. 4) shows that the fused features achieve better inter-class separation and greater intra-class compactness. At a bit error rate of 4 × 10–5, the proposed method achieves a gain of approximately 0.2 dB over cross-term detection (Fig. 6). Although ATSFF has higher computational overhead and lower real-time performance than conventional methods, its single-sample inference cost remains fixed after the network architecture is established, and GPU-based batch processing is supported. The method is therefore suitable for communication scenarios with high detection-accuracy requirements.  Conclusions   An adaptive temporal-spatial feature fusion method is proposed for non-orthogonal PSWF signal detection under severe mutual interference. Dual-path feature extraction is achieved using GRU and ResNet50, and a prediction-uncertainty-based adaptive probability-weighted fusion mechanism is used to integrate complementary temporal and spatial features. The simulation results demonstrate improved detection accuracy and robustness under complex channel conditions. The proposed method provides a feasible approach for high-accuracy detection of non-orthogonal PSWF signals.
UAVREL: A Benchmark Dataset for Dynamic Relation Comprehension in UAV Videos
LIU Xiaorui, DENG Chubo, HOU Zhongyan, YAN Qiwei, LU Wanxuan, HOU Yingyan, YU Hongfeng, SUN Xian
Available online  , doi: 10.11999/JEIT260221
Abstract:
  Objective  With the rapid development and extensive application of unmanned aerial vehicle (UAV) observation platforms, remote sensing video data with high spatiotemporal resolution has witnessed explosive growth. Understanding dynamic relationships in UAV videos is recognized as a pressing research challenge in intelligent remote sensing analysis. Current research on video scene graph generation is mainly focused on natural videos, while studies on remote sensing videos remain in the exploratory stage, with a lack of benchmark datasets annotated with high-order dynamic relationships. Traditional visual models are difficult to be directly adapted to remote sensing scenes, which are characterized by dynamic view changes, extreme object scale variations, and dense target distributions. To address these limitations, a UAV video relationship dataset (UAVREL) with high-order dynamic relationship annotations is constructed, a hypergraph-enhanced transformer method tailored to the characteristics of UAV videos is proposed, and a unified benchmark evaluation system for video scene graph generation in the field of UAV remote sensing is established in this study. These efforts promote the transformation of intelligent remote sensing interpretation technology from static object cognition to global dynamic scene understanding.  Methods  This study is mainly composed of three core parts: dataset construction, model design, and comprehensive experimental validation. First, in the dataset construction stage, the Unmanned Aerial Vehicle Benchmark for Object Detection and Tracking (UAVDT) dataset is selected as the basic data source. A semi-automatic annotation strategy is proposed, integrating object tracking, multi-person collaboration, and consensus verification to mitigate subjective bias and improve efficiency. The relationships are divided into three levels according to the requirements of cross-frame inference, among which high-order relationships are annotated with high priority. Second, A Hypergraph-Enhanced Transformer model (STHG) is proposed in this work. It includes six core functional modules: basic feature extraction, pairwise context encoding, spatial-temporal dual hypergraph enhancement, multi-source feature fusion, global temporal modeling and relation prediction. On this basis, a multi-label classifier generates the final dynamic scene graphs. In the experimental design, to verify the effectiveness of the dataset and the model, an object detection task and three video scene graph generation subtasks, namely predicate classification (PreCls), scene graph classification (SGCls), and scene graph detection (SGDet), are conducted on the UAVREL dataset. Representative object detection models and classic scene graph generation models are selected as baselines. Mean Average Precision (mAP) and mAP@50 are adopted as evaluation metrics for detection, while Recall@K and meanRecall@K (mR@K) are used for scene graph generation to comprehensively evaluate the performance of the model in relation recognition and graph construction.  Results and Discussions  In the object detection task, the impact of model architectures on detection performance is systematically verified through comparative experiments on eight models with different architectures (Table 2). The results show that the selection of model architecture is of crucial importance to the mean Average Precision (mAP), a core evaluation metric. Single-stage anchor-free models represented by VFNet and DDOD exhibit significant advantages in comprehensive detection performance. From the perspective of category characteristics, all models perform poorly in detecting small-scale and easily occluded target categories, which reflects the common technical challenge faced by current general object detectors in small target detection tasks. In the video scene graph generation task, five methods are tested on three subtasks respectively (Table 3, Table 4, Table 5). The STHG method proposed in this paper shows significant performance advantages in all three core tasks. Meanwhile, experimental data indicate that the value of the average recall metric is consistently significantly lower than that of the traditional recall metric. This phenomenon clearly shows that the dataset poses great modeling challenges in task scenarios with low-frequency object relationships, and implicitly reflects that relationship prediction in such complex scenarios remains a key challenge to be solved urgently in the field of drone video scene graph generation.  Conclusions  This paper focuses on the critical theme of understanding dynamic relationships in drone videos. It constructs the UAVREL benchmark dataset, providing data support with deeper semantic relationships for this field. And it proposes a Hypergraph-Enhanced Transformer approach for Remote Sensing Videos. Experimental results demonstrate that this approach achieves superior performance across multiple evaluation metrics, thus validating its practical applicability in remote sensing dynamic relationship prediction tasks. Through dataset construction and algorithmic innovation, this paper not only lays a solid data foundation for understanding drone video relationships, but also facilitates a leap from low-level semantic analysis to high-level dynamic semantic cognitive modeling in remote sensing video analysis. Future research will focus on the following two directions: Firstly, deepening the temporal dimension modeling of dynamic remote sensing scene graph generation to enhance the ability to capture long-term evolutionary events and complex relationships; secondly, expanding the scene coverage and diversity of relationship categories in the dataset, continuously improving algorithm benchmarks, and promoting technological iteration and industry application implementation.
Random-Linear-Network-Coding-based Cooperative Reliable Transmission Protocol for Underwater Acoustic Communication Networks
ZHANG Zhilin, PU Zhanqing, ZHU Yunan, LI Xueying, TIAN Jie, HUANG Haining
Available online  , doi: 10.11999/JEIT260648
Abstract:
  Objective  Reliable data delivery in underwater acoustic communication networks is challenged by high packet error rates, long propagation delays, limited bandwidth, and topology variations. In single-source dual-destination multi-hop transmission, the same data generation must be reliably delivered to two destination nodes. Packet losses at individual hops can accumulate during multi-hop forwarding and joint recovery at the two destinations, further complicating reliable delivery. Existing reliability-enhancement mechanisms, including retransmission, redundant forwarding, forward error correction, and multipath redundant transmission, generally rely on predetermined forwarding structures or fixed redundancy configurations. They have limited capability to exploit complementary coded information distributed among multiple relay nodes, resulting in insufficient joint recovery capability and high redundant transmission overhead. To address these limitations, a Network-Coded Cooperative Reliable Transmission Protocol for underwater acoustic communication networks (NCCRTP) is proposed.  Methods  NCCRTP operates on a generation basis and employs Random Linear Network Coding (RLNC) over the Galois field \begin{document}$ \text{GF}({2}^{8}) $\end{document}. To reduce coding overhead, each packet carries a Code IDentifier (CodeID) rather than the complete coding vector. The corresponding coding vector is recovered from a shared coding-vector dictionary at the relay node. During hop-by-hop forwarding, NCCRTP generates forward candidate structures subject to a residual-hop decreasing constraint and adaptively selects among three transmission modes: SINGLE, COOP, and BRANCH. SINGLE maintains a shared forwarding process toward the two destinations. COOP enables two relay nodes to jointly utilize linearly independent coded packets received at different nodes. BRANCH divides the transmission into two branches toward the different destinations. For each candidate structure, NCCRTP estimates the link success rate, calculates the required transmission budget, and evaluates the two-hop structural utility. The forwarding mode is then selected according to the tradeoff between recovery capability and transmission overhead.  Results and Discussions  Simulation results show that NCCRTP achieves the highest Joint Packet Delivery Ratio (JPDR) under both regular and random topologies. In the controlled comparison with Cooperative Uncoded transmission (CU), Single-branch Uncoded transmission (SU), and Single-branch Coded transmission (SC), NCCRTP consistently outperforms schemes using only cooperative forwarding or only RLNC. This result indicates that the reliability gain is jointly provided by distributed relay cooperation and joint utilization of linearly independent coded packets (Fig. 5). As the packet error rate increases or the end-to-end transmission depth increases from 3 to 7 hops, NCCRTP maintains a higher JPDR, demonstrating stronger robustness under lossy multi-hop conditions (Figs. 5(a) and 5(b)). In random topologies, Vector-Based Forwarding (VBF) and Focused Beam Routing (FBR) are separately combined with packet REPlication (REP) or RLNC to form the VBF+REP, VBF+RLNC, FBR+REP, and FBR+RLNC schemes (Figs. 6 and 7). Under medium-to-high packet error rate or multi-hop transmission conditions, NCCRTP improves the JPDR by up to approximately 50%, while reducing the equivalent transmission overhead per successful joint delivery by up to approximately 40% (Fig. 6). These results indicate that NCCRTP improves dual-destination reliability through adaptive forwarding-structure selection, link-quality-based transmission-budget control, and joint utilization of linearly independent coded packets rather than simply increasing redundant transmissions.  Conclusions  The reliability and redundant transmission overhead challenges in single-source dual-destination underwater acoustic multi-hop transmission are addressed by designing a cooperative transmission structure that enables RLNC to exploit distributed reception and complementary coded information among relay nodes. The proposed NCCRTP protocol adaptively selects the SINGLE, COOP, and BRANCH transmission modes according to residual-hop constraints, link-quality-based transmission-budget control, and two-hop structural utility evaluation. A lightweight coding-vector representation based on CodeID is also adopted to reduce the header overhead associated with carrying complete coding vectors. The protocol is evaluated under both regular and random topologies. The results show that: (1) NCCRTP achieves the highest JPDR among all compared schemes, demonstrating stronger joint recovery capability at the two destination nodes; (2) under medium-to-high packet error rates or multi-hop transmission conditions, NCCRTP improves the JPDR by up to approximately 50%; and (3) the equivalent transmission overhead per successful joint delivery is reduced by up to approximately 40%, indicating that the reliability gain mainly comes from adaptive structure selection, transmission-budget control, and joint utilization of linearly independent coded packets rather than excessive redundant transmissions. Future work will extend NCCRTP to more complex multi-source, multi-destination, multi-hop scenarios and further investigate its implementation and performance under node mobility and realistic underwater acoustic channel dynamics.
LEO Satellite Multi-beam Multicast Precoding and User Grouping Joint Optimization Algorithm
GUO Lili, FENG Yimeng, YUAN Peihong, GAO Yue
Available online  , doi: 10.11999/JEIT260375
Abstract:
  Objective  In Sixth-Generation (6G) Low Earth Orbit (LEO) satellite communication systems, multicast precoding is adopted to mitigate severe inter-beam interference caused by Full Frequency Reuse (FFR). However, conventional precoding algorithms exhibit cubic computational complexity, limiting their applicability to massive Multiple-Input Multiple-Output (MIMO) systems. Existing user grouping methods also fail to satisfy the fixed group-size requirement specified by the DVB-S2X standard. To address these limitations, a joint optimization framework is proposed that combines a low-complexity unsupervised deep learning-based precoding model with improved user grouping algorithms to improve the system sum rate and fairness.  Methods  An unsupervised deep learning model based on a hybrid Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) architecture is proposed for precoding (Fig. 2). The Convolutional Neural Network (CNN) extracts spatial features from Channel State Information (CSI), while the Long Short-Term Memory (LSTM) network captures high-level feature correlations. The model is trained by directly maximizing the system sum rate while satisfying the Per-Antenna power Constraint (PAC), without requiring supervised labels. For user grouping, two algorithms compatible with the DVB-S2X standard are developed. First, the CK-means algorithm extends conventional K-means clustering to ensure an equal number of users in each group while preserving high intra-group channel similarity. Second, the Fairness-Aware MAUG (FA-MAUG) algorithm prioritizes users with poor channel conditions during grouping, thereby improving system robustness and fairness.  Results and Discussions  The intra-group similarity metric is used to evaluate user grouping performance. The results show that the CK-means algorithm achieves an average similarity approximately 0.1 higher than that of the MAUG algorithm and nearly 0.5 higher than that of random grouping across different group sizes (Fig. 3), resulting in improved beamforming gain. In terms of system sum rate, the proposed CNN-LSTM precoding combined with CK-means grouping consistently outperforms the conventional Minimum Mean Square Error (MMSE) algorithm under different Signal-to-Noise Ratios (SNRs) and total transmit power levels (Fig. 4 and Fig. 5). Under different SNR conditions, the proposed CNN-LSTM precoding scheme improves the average system sum rate by 48.59% compared with the MMSE algorithm, whereas CK-means grouping increases the average system sum rate by 30.12% relative to random grouping. The effects of the number of users per group, the number of groups, and the number of antennas are further evaluated (Fig. 6-Fig. 8), demonstrating that the proposed framework maintains superior performance across systems of different scales. Complexity analysis further shows that the proposed precoding method reduces the online computational complexity from the cubic complexity of the conventional MMSE algorithm to linear complexity with respect to the number of antennas, making it well suited for real-time deployment in large-scale LEO satellite communication systems.  Conclusions  A joint optimization framework is proposed for LEO satellite multicast communication systems to address the high computational complexity of precoding and the limited fairness of conventional user grouping methods. Simulation results demonstrate that the proposed framework substantially improves the system sum rate, achieving an average gain of 48.59% over the conventional MMSE algorithm under different SNR conditions while simultaneously reducing online computational complexity. The proposed framework provides an effective and scalable solution for multi-beam interference mitigation and resource optimization in future 6G LEO satellite communication systems.
Image Classification Network Based on Complementary Decay Learning
YUAN Heng, TIAN Wenyue, ZHANG Shengchong
Available online  , doi: 10.11999/JEIT260751
Abstract:
  Objective  Image classification depends on complete and discriminative feature representations. Existing convolutional neural networks usually enhance positive and high-amplitude responses through activation functions and attention mechanisms. However, excessive reliance on dominant positive responses may shift attention from the whole object to local salient regions, while negative and low-amplitude responses containing edge, texture, and foreground-background transition information are often weakened. To address this problem, a Complementary Decay Learning Network (CDLNet) is proposed to suppress high-response dominance and preserve complementary feature information.  Methods  Inspired by the signal attenuation mechanism of the biological visual system, CDLNet introduces complementary decay learning into a residual network. The Spatial Complementary Decay (SCD) module divides features into positive-response, negative-response, and global-response branches, and applies differentiated decay to preserve salient regions, boundary details, and contextual information (Fig.2). The Channel Complementary Decay (CCD) module attenuates high-response channels while retaining middle- and low-response channels, thereby reducing channel dominance and promoting cooperative channel representation (Fig.3). The Complementary Decay Attention (CDA) module integrates SCD and CCD in parallel and is embedded into ResNet residual blocks to jointly regulate spatial structures and channel semantics (Fig.4, Fig.6).  Results and Discussions  Experiments are conducted on CIFAR-10, CIFAR-100, SVHN, Imagenette, Imagewoof, and ImageNet datasets. CDLNet achieves classification accuracies of 96.68%, 81.81%, 97.38%, 92.45%, 85.46%, and 62.88%, respectively. Compared with ResNet-34 and representative classification networks, CDLNet obtains higher accuracy on multiple datasets (Table 5, Table 6). Ablation experiments demonstrate that removing either SCD or CCD reduces classification accuracy, indicating that spatial and channel complementary decay both contribute to feature regulation (Table 4, Table 5). Visualization results show that CDLNet can enhance target regions, preserve structural details, and suppress irrelevant background responses (Fig.1, Fig.12). Although CDA increases parameters and computation, the accuracy improvement shows a reasonable balance between performance and complexity (Table 3).  Conclusions  CDLNet introduces spatial and channel complementary decay mechanisms into residual networks. By suppressing excessive high-response dominance and preserving negative-response and low-amplitude information, the proposed method improves the completeness and balance of feature representations, alleviates attention drift, and enhances image classification performance. Future work will further optimize the decay strategy and model complexity to improve efficiency and generalization.
A Heterogeneous Multi-View Semantic Fusion Training Method for Text Classification in Low-Resource Scenarios
XU Sen, DENG Jubiao, XU Xiufang, YAO Shanliang, BIAN Xuesheng, BEN Xianye
Available online  , doi: 10.11999/JEIT260365
Abstract:
  Objective  The increasing demand for text classification in specialized domains, such as finance, healthcare, and social media, is often hindered by the scarcity of labeled data. Low-resource or data-scarce scenarios significantly limit the effectiveness of conventional supervised learning and pre-trained language models, such as BERT, due to insufficient task-specific supervision. Existing methods that rely on single-view weak supervision or clustering often generate noisy pseudo-labels and fail to fully exploit the discriminative potential of hard-to-classify samples. Therefore, it is essential to develop a robust training method that leverages unlabeled data, mitigates noise sensitivity, and enhances semantic representation in low-label conditions. The primary objective of this research is to propose a text classification training framework specifically designed for data-scarce scenarios.  Methods  To address these challenges, a Heterogeneous Multi-View Semantic Fusion (HMVSF) training framework is introduced. HMVSF consists of three main stages: (1) Dual Clustering Consistency Screening (DCCS): Text samples are first represented in the word-frequency view using TF-IDF vectors. Two heterogeneous clustering algorithms, sIB and an improved K-means initialized via a genetic algorithm, are applied to the same feature space. Samples that are consistently assigned to intersecting clusters are selected as high-confidence pseudo-labeled instances. These samples are subsequently used for a lightweight supervised update of the shared encoder, ensuring reliable initialization for subsequent training. (2) Latent Dirichlet Allocation (LDA) Guided Hard-Sample Retraining: The remaining hard examples, which are not selected in the first stage, typically exhibit low discriminability in the word-frequency view but carry global semantic and topic-level information. LDA is employed to model the topic distribution of these samples. The optimal number of topics is adaptively determined by a combination of semantic coherence and statistical fitting criteria. Hard and soft pseudo-labels are generated based on the document-topic distribution, and the model undergoes targeted retraining using a combined cross-entropy and symmetric KL divergence loss to integrate topic-level semantic cues. (3) Downstream Fine-Tuning: Finally, the encoder is fine-tuned on a small labeled set by replacing temporary heads with a task-specific classification head. Standard cross-entropy is employed as the objective. The two-stage intermediate training is independent of the pre-training initialization, allowing HMVSF to be applied to both general-purpose BERT and masked language model (MLM)-based initializations, making it widely applicable.  Results and Discussions  The effectiveness of HMVSF is evaluated on six publicly available datasets, including DBpedia, AGNews, ISEAR, SMS Spam, Subjectivity, and Polarity, which cover both topic-based and non-topic-based text classification tasks. Two initialization routes are tested: Route A (BERT/BERTIT:CLUST) and Route B (BERTIT:MLM/BERTIT:MLM+CLUST). Under extremely low label budgets (e.g., 64 labeled samples), CCS-BERTIT:CLUST and CCS-BERTIT:MLM+CLUST outperform baseline models and single-stage variants across datasets (Figs.23, Tables 23). For instance, in Route A, CCS-BERTIT:CLUST achieves up to 9% accuracy improvement on DBpedia and ISEAR compared with BERTIT:CLUST, while error rates decrease by 10–54% (Table 2). Normalized Mutual Information (NMI) and intra-cluster Euclidean distance analyses confirm that the CCS selection produces stable and compact clusters, ensuring reliable pseudo-labels (Tables 45). Ablation experiments show that the two-stage design combining CCS and LDA is superior to single-stage variants, demonstrating complementary advantages: CCS provides high-confidence core samples, and LDA captures latent semantic structures of hard examples (Table 7). Parameter sensitivity analysis reveals that the model is robust under a wide range of cluster numbers and LDA weight coefficients, achieving stable performance when the LDA weighting coefficient is set to 0.3 (Figs.45). HMVSF also maintains competitive performance as labeled data increase. Computational experiments on an NVIDIA A40 GPU demonstrate that the two-stage intermediate training introduces acceptable additional overhead while enhancing representation quality and overall accuracy.  Conclusions  This study presents a novel heterogeneous multi-view semantic fusion training method (HMVSF) for data-scarce text classification. By integrating cluster-consistent sample selection in the word-frequency view and LDA-guided hard-sample retraining, HMVSF effectively leverages unlabeled data, mitigates pseudo-label noise, and enriches semantic representations. Extensive experiments across multiple datasets confirm that HMVSF significantly improves classification accuracy under low-label conditions while remaining robust to pre-training initialization. The framework is generalizable to various pre-trained language models and represents an effective and practical solution for text classification in low-resource scenarios.
Bayesian-Optimized Neural Network Rapid Solver for HEMP Waveform Distribution
WANG Jinjin, ZHAO Mo, JING Jing, WANG Wenbing, LIU Zheng, LI Jinxi, WU Wei, LIU Tieming
Available online  , doi: 10.11999/JEIT260877
Abstract:
  Objective  High-altitude Electromagnetic Pulse (HEMP) has large area, strong fields. It is difficult for HEMP's damage of critical infrastructures which are crucial to country's electromagnetic security. Especially under E1 environment, the waveform distribution characteristic will have direct influence on the effectiveness of protection design. But for cases which need a lot of waveforms calculation like system level effects simulation and protection scheme optimization etc., traditional calculation method of waveform parameter in coverage region is time consuming, which can not realize the real-time calculation of massive ground waveforms on the field distribution. The rapidly and accurately calculation of other parameters of the waveform function is still an issue for the electromagnetic environment field, limiting the useful application of computed fast HEMP environment results for a particular type of explosion when using these results in another calculation. This approach solves a problem that conventional computations cannot be linked into software to perform on-line computation, which provides the basis for the simulation calculation as well as HEMP damage evaluation, enables a fast comparison between different large scale scenarios, enabling us to calculate waveform functions for hundreds of different cases within a few minutes and thus constitutes an efficient tool for rapid analyses as part of the HEMP robust design and evaluation.  Methods  In this paper, an approach based on Bayesian optimization for deciding how many nodes should be included within each hidden layer of a multilayer feed forward NN is proposed. Reducing searching times for hidden layer node and ensure model precision, in fast solution modeling of HEMP waveform function. And a neural network is firstly utilized for abstracting all of the physical calculation process as a function and builds up a multilayer feedforward neural network model between input and output. However, in such networks, the choice of hidden layer node count affects network performance. In order to achieve both high prediction accuracy and low model complexity, Bayesian optimization finds the best solution in few iterations. While building the surrogate model, an initial data set is formed, and a gaussian process is utilized for the surrogate model which represents the distribution of the objective function. Updating the acquisition function, the best fit parameters, and For the HEMP waveform function, models for Emax, \begin{document}$ \alpha $\end{document}, \begin{document}$ \beta $\end{document}, k, and t0 are sequentially established to predict waveform parameters under different conditions. By calculating the waveforms along the north south axis and incorporating the angle between the burst point projection and the observation point, the HEMP waveform parameters at any location are further derived and computed.  Results and Discussions  The above method uses both numerical calculation and Bayesian optimization based neural network algorithm, in order to build up an artificial intelligence model of predicting the ground HEMP waveform parameters, covering arbitrary height of bursts, yield of gammas, position in a certain range. Bayesian optimization neural network method reduces searching times on the number of nodes in hidden layers and improve the predicting precision. In Bayesian optimization process, the searching space of hidden layer as [8, 50] is defined to avoid model over fitting, and gives a best network structure of quite small number of hidden layer nodes. To control algorithm training time, the maximum number of Bayesian evaluations is set to 20, approximately half the search space. Through Bayesian optimization, the optimal hidden layer node counts for Emax, \begin{document}$ \alpha $\end{document}, \begin{document}$ \beta $\end{document}, k, and t0 are found to be 19, 16, 11, 13, and 10, respectively. The simulation results indicate that the error of waveform parameter predicted by this method compared to the calculation result of simulation is lower than 3.12%, which has a better performance in comparison with several other methods. It works best on each metric. Experimental comparisons show the stability and generalization ability of the proposed algorithm for predicting different parameters. Analysis shows that Bayesian optimization, using a probabilistic surrogate model and an acquisition function, defines a probabilistic map between the number of nodes and the error on the validation set, to guide the following sampling steps with uncertainty estimation on predictions, which reduces the original calculation time from hours to seconds and the time complexity from O(n5) to O(n2), supporting large scale real time computation for the parameters in a HEMP waveform, given different scenarios.  Conclusions  This paper proposes a Bayesian optimization based multilayer feedforward neural network method to model the simulation computation process of HEMP ground waveforms, enabling rapid calculation of standard waveform functions for all points within the ground field distribution of HEMP over a certain range. The method uses Bayesian optimization to optimize the number of hidden layer nodes in the multilayer feedforward neural network, reducing search iterations and improving model prediction accuracy. Compared with five other artificial intelligence methods, the FBEMP method performs best in terms of all metrics. By combining the neural network with numerical derivation, all waveform functions within the field distribution coverage area can be calculated for different burst heights, gamma yields, longitudes and latitudes. This approach lowers the order of operation count from O(n5) in conventional numerical computation to O(n2), and reduces the calculation time of waveform functions for a given field distribution from hour scale to second scale, with errors on each parameter less than 3.12%. It realizes real time calculation of waveform function in HEMP field distribution, and had become applied to large-scale real-time HEMP waveform function calculations, solving the longstanding technical challenge of time-consuming HEMP environment computations that previously prevented real-time calculation, thereby providing an online computational environmental foundation for digital simulation and assessment in HEMP experiments.
Design and Performance Evaluation of Ultrasonic Nebulization Glow Discharge Detector for High-sensitivity and Rapid Detection of Metal Elements in Liquids
DING Yu, XU Jianan, PANG Maoyuan, WANG Yuhang, LI Jinyi, YU Weiye, HE Yihua, LI Xiangchu, TAN Qiang, LIU Xinxin, ZHOU Wangping
Available online  , doi: 10.11999/JEIT260436
Abstract:
  Objective  Water is essential for all organisms and ecological systems, and the composition and content of dissolved metal elements, especially copper (Cu), sodium (Na), and potassium (K), are crucial for maintaining ecological balance and biological health. Cu is an essential human trace element that forms enzymes with functional proteins, participating in antioxidant, energy supply, and immune processes. However, excessive Cu in water—from industrial wastewater, feed additives, pipeline corrosion, and electronic waste leakage—harms human health, causing vomiting, hypotension, jaundice, and hemolytic anemia. Aquatic organisms and plants, lacking effective detoxification systems, are more vulnerable to Cu pollution, which damages roots, induces oxidative stress, and inhibits growth. Na and K are vital for nerve conduction, muscle movement, and fluid balance, but excessive intake endangers those with hypertension or renal/cardiac insufficiency, and their imbalance in irrigation water causes soil salinization. Conventional detection methods (ICP-MS/AES, AAS, AFS) have excellent sensitivity but are limited by large size, complex pretreatment, high cost, and professional operation. Atmospheric pressure glow discharge (APGD) shows potential for miniaturization, but existing APGD-based technologies (PN-APGD, SCGD) require expensive equipment or complex pretreatment. Thus, a highly sensitive, rapid detector for on-site real-time detection of Cu, Na, K without additional pretreatment or driving equipment is urgently needed.  Methods  A Ultrasonic Nebulization Glow Discharge (UNGD) detector was designed, consisting of an ultrasonic nebulization unit, a plasma excitation unit, and a spectral signal collection unit. The ultrasonic nebulization unit adopted a self-designed centrifuge tube-based diversion chamber with a detachable microporous atomizing sheet, argon inlet/outlet, and a space-constrained transfer tube to form a short gas path, reducing aerosol loss. The plasma excitation unit used 180° coaxial tungsten needle electrodes (cathode/anode) with a double-layer fixing sleeve, clamped on a 3D platform for precise spacing adjustment, powered by a high-voltage DC power supply with a 20 kΩ ballast resistor. The spectral unit included an optical fiber probe and a three-channel AvaSpec spectrometer for high-resolution signal collection. Performance was evaluated using Cu/Na/K mixed solutions: single-factor experiments optimized parameters; characteristic spectra for qualitative analysis; recovery rates for anti-interference assessment; standard curves, LOD, RSD, and CRM detection for quantitative verification; and comparison with similar technologies.  Results and Discussions  Qualitative analysis of a water sample (Cu: 5.28 mg/L, Na: 5.23 mg/L, K: 3.89 mg/L) showed OH (281.1~309.0 nm) and N2 (315.0~406.0 nm) molecular bands, with obvious Cu (324.7 nm), Na (589.0 nm), and K (766.5 nm) characteristic peaks (Fig. 2); 324.7 nm was selected as Cu’s analytical line. Parameter optimization determined optimal conditions: discharge current 38 mA (Fig. 3), argon flow rate 0.5 L/min (Fig. 4), electrode spacing 1 mm (Fig. 5), sampling distance 22 mm (Fig. 6). Under these conditions, Cu, Na, K showed good linearity (R2: 0.9942, 0.9968, 0.9973), with LODs of 141.29 μg/L, 9.54 μg/L, 12.05 μg/L, and RSDs of 6.8%, 6.1%, 5.5% (n=11) (Table 1). Anti-interference tests showed 90%~110% recovery rates with 500 mg/L interfering cations (Fig. 7). CRM detection showed 89%~110% recovery rates, consistent with standard values (Table 2).  Conclusions  The UNGD detector achieves accurate quantitative analysis of Cu, Na, K in liquids without additional pretreatment or driving equipment. Its detachable atomizing sheet and short gas path improve sampling efficiency, while coaxial electrodes concentrate excitation energy. With excellent sensitivity, precision, and anti-interference ability, it provides reliable technical support for on-site real-time monitoring of water metal elements and has potential for extending to other metal detections, contributing to water ecological protection and biological health.
Patch-Sinusoidally Modulated SSPPs Leaky-Wave Antenna and Its Random Forest-Assisted Optimization Design
TANG Luping, CHENG Yonghao, CHEN Yibo, LIAO Chen
Available online  , doi: 10.11999/JEIT260651
Abstract:
  Objective  Spoof surface plasmon polaritons (SSPPs) leaky-wave antennas feature low-profile configuration and inherent frequency-scanning capability, making them promising for modern radar, communication, and intelligent sensing systems. However, strong nonlinear coupling among geometric parameters makes traditional optimization computationally costly, as full-wave simulations require thousands of evaluations and often converge to suboptimal local solutions due to landscape complexity. The leakage dynamics in SSPPs—slow-wave propagation, spatial harmonic coupling, and leakage rate distribution—adds complexity beyond conventional designs. To address this, we propose a patch-sinusoidally modulated SSPPs leaky-wave antenna and a machine learning framework integrating a random forest surrogate with particle swarm optimization (PSO) for efficient high-dimensional global optimization with reduced cost.  Methods  Unlike conventional groove-depth modulation, our antenna maintains uniform groove depth and a complete metal ground. Patch arrays with sinusoidally varying widths are loaded on both sides of the transmission line, with envelope functions \begin{document}$ Y=A\sin (Tx) $\end{document} and \begin{document}$ Y=A\sin (Tx+\pi ) $\end{document}, enabling flexible leakage control. The antenna is fully described by a nine-dimensional continuous parameter vector: groove width g, depth s, period p, six transition lengths g1–g6, port dimensions l and w, and modulation A, T. Using Latin hypercube sampling, 450 parameter samples are generated to ensure uniform coverage. Full-wave frequency-domain simulations (COMSOL, 9 GHz, approximately 27 min each) extract gain, S11, S21, scanning angle, side lobe level (SLL), and total efficiency. Four regression models—MLP, SVR, random forest (RF), and GPR—are systematically trained. RF employs bootstrap resampling with hyperparameters optimized via random search cross-validation (trees: 500–1200, depth: 12–24, min samples per split: 2–5). The trained RF surrogate is then embedded into PSO (40 particles, 120 iterations, inertia 0.72, c1=c2=1.5) with a weighted fitness function (G:1.5, η:0.7, SLL:0.55, S11:0.25, S21:0.7, θscan:0.18).  Results and Discussions  RF achieves the highest average R2 of 0.9554 across six outputs, outperforming GPR (0.9489), SVR (0.9212), and MLP (0.8952). For key radiation indicators, RF attains gain MAE of 0.032 dBi, SLL MAE of 0.168 dB, and efficiency MAE of 0.015. Scatter plots of predicted versus simulated values cluster tightly around the diagonal, and residual histograms show means near zero with no systematic bias, confirming excellent prediction accuracy and generalization. After RF-PSO optimization, full-wave simulation confirms substantial improvements: gain rises from 13.84 dBi to 14.52 dBi, SLL drops from –17 dB to –19 dB,peak total efficiency increases from 84.9% to 92.9%, S11 improves from –23.26 dB to –27.92 dB (4.66 dB), and scanning range expands from 57.2° to 61.5°. The scanning angle versus frequency curve exhibits good linearity across the operating band, and the two-dimensional far-field patterns show improved symmetry. The decrease in S21 (from –3.76 dB to –5.50 dB) together with the gain/efficiency increase indicates that more energy is effectively converted into radiation rather than being dissipated or reflected. Sensitivity analysis with ±2% perturbations (50 samples) shows all coefficients of variation (CV) below 1.3%: gain CV 0.21% (<±0.1 dBi), SLL CV 1.26% (±0.4 dB), efficiency CV 0.88% (±0.01), S11 CV 0.97%, S21 CV 0.91%, scan CV 0.37%. These fluctuations are far smaller than optimization gains, confirming excellent robustness under typical fabrication tolerances. Comparison with recent leaky-wave antennas (both SSPP-based and SIW) demonstrates superior SLL (−19 dB), competitive efficiency (89% vs. 94.95% and90% in prior SSPP work), and scanning range (61.5°) outperforming most single-port SSPP antennas (e.g., 20°, 16°, 33°, 13°). The number of full-wave simulations is reduced by approximately 90% (450 training + 1 validation vs. 4,800 simulations for conventional PSO).  Conclusions  This paper proposes a patch-sinusoidally modulated SSPPs leaky-wave antenna and an RF-assisted PSO framework for synergistic optimization in nine dimensions. The RF surrogate achieves an average R2 of 0.9554. The optimized antenna shows significantly improved performance across all metrics: gain by 0.68 dBi, SLL by 2 dB, efficiency by 8%, S11 by 4.66 dB, and scanning range by 4.29°, while maintaining compact dimensions. Sensitivity analysis confirms robustness under typical fabrication tolerances. The proposed methodology reduces the number of full-wave simulations by approximately 90%. This work marks a methodological advancement by introducing machine learning surrogate modeling into SSPPs leaky-wave antenna design for efficient high-dimensional optimization. Future work includes fabrication, experimental validation, extension to millimeter-wave bands, and reconfigurable antenna designs.
Physical-layer Network Coding Aided Polar Slotted Random Access Algorithm
SHAO Caiping, QIU Yuping, XIE Zhaopeng, SONG Dan, CHEN Jian, CHEN Pingping
Available online  , doi: 10.11999/JEIT260867
Abstract:
  Objective  Massive Machine-Type Communications (mMTC) constitutes a fundamental pillar of 5G and emerging 6G wireless networks, dedicated to supporting massive connectivity for the Internet of Things (IoT). In grant-free random access scenarios, sporadic and uncoordinated transmissions by massive terminals inevitably induce severe packet collisions under heavy traffic loads. Traditional scheduled access protocols become inefficient due to prohibitive signaling overhead. Consequently, grant-free slotted ALOHA protocols based on Successive Interference Cancellation (SIC)—such as Contention Resolution Diversity Slotted ALOHA (CRDSA), Irregular Repetition Slotted ALOHA (IRSA), and Coded Slotted ALOHA (CSA)—have garnered widespread attention. However, these classical schemes fundamentally rely on the presence of collision-free (degree-1) slots to trigger and sustain the iterative graph-peeling decoding process. Under practical constraints of finite frame lengths and heavy traffic, collision-free slots are drastically depleted, causing severe decoding stalling and throughput degradation. Furthermore, pure SIC mechanisms are intrinsically vulnerable to error propagation. Although Polar Slotted ALOHA (PSA) introduces polarization transforms across time slots to enhance packet recovery, its collision resolution remains constrained by the initial SIC condition. To resolve these challenges, a joint decoding algorithm combining Polar Slotted ALOHA with Physical-Layer Network Coding (PSA-PNC) is proposed over the Slot Erasure Channel (SEC). The objective is to transform destructive multi-user packet collisions into algebraically solvable linear network coding equations, thereby eliminating the strict reliance on collision-free slots and significantly elevating concurrent multi-user detection capability and throughput performance under heavy traffic loads.  Methods  A joint physical-layer and MAC-layer random access framework is established over the Slot Erasure Channel (Fig. 1). At the transmitter side, active users independently select transmission time slots based on an irregular degree distribution polynomial without inter-user coordination or channel collision feedback. The sender remains blind to multi-user collision patterns in the channel. At the base station receiver, the superimposed signals across slots are equivalently modeled as a sparse global input matrix over a binary finite field, followed by packet-level polar encoding. To resolve dense collisions without relying on clean slots, a closed-loop iterative receiver architecture is developed (Fig. 2). In each iteration, channel observation sequences are initially processed by a packet-level Successive Cancellation (pSC) or packet-level Successive Cancellation List (pSCL) decoder to extract reliable equivalent combined packets from information slots. Instead of being discarded, the collided slots corresponding to these reliable packets are utilized to construct a local sparse linear Network Coding (NC) equation system. A Generalized Matrix Inversion (GMI) criterion is subsequently executed to analyze the column-rank characteristics of the access pattern matrix and achieve global algebraic multi-user decoupling. Solvable user packets are directly recovered without requiring full-rank matrix conditions or degree-1 slots. The algebraically decoupled user packets are then utilized as prior information to reconstruct physical-layer codewords and subtracted from the observation buffer via iterative SIC, continuously reducing the dimensionality of the unresolved collision space. High-reliability packets output by the pSCL decoding paths are leveraged in the iterative loop to effectively suppress error propagation and guarantee the linear independence of residual equations. Furthermore, the polarization evolution process of equivalent multi-user packets over the SEC is theoretically proved to be equivalent to that over a scalar Binary Erasure Channel (BEC), enabling rigorous calculation of frame error rate bounds via Bhattacharyya parameters (Fig. 3).  Results and Discussions  Extensive theoretical analyses and Monte Carlo simulations are conducted to evaluate the performance of the proposed PSA-PNC scheme over the Slot Erasure Channel. The theoretical polarization bounds of packet-level polar decoding are verified under various slot erasure probabilities (\begin{document}$ \epsilon \in \left\{0.1,0.2,0.3\right\} $\end{document}), exhibiting precise consistency with simulation curves and validating the polarization threshold effect (Fig. 3). In terms of system throughput, simulation results demonstrate that for a frame length of \begin{document}$ N=1024 $\end{document}, the proposed PSA-PNC scheme with pSCL (\begin{document}$ L=8 $\end{document}) achieves a peak normalized throughput of approximately 0.87 packets/slot at a normalized load of \begin{document}$ G\approx 0.90 $\end{document}, yielding an approximate 15% throughput improvement over baseline PSA and outperforming Coded Slotted ALOHA under identical finite-length configurations (Fig. 4(a)). In the low-load region, all evaluated schemes exhibit near-identical linear throughput growth due to the abundance of collision-free slots (Fig. 4(a)). When the slot erasure rate increases to \begin{document}$ \epsilon =0.35 $\end{document}, the number of recoverable reliable equivalent packets decreases, leading to insufficient NC equations and observable throughput degradation in high-load regions, which confirms the operational boundary of the algorithm (Fig. 4(a)). For a short frame length of \begin{document}$ N=64 $\end{document}, a consistent throughput gain ranging from 0.12 to 0.18 is maintained by PSA-PNC, demonstrating strong robustness against finite-length decoding stalling in short-packet scenarios (Fig. 4(b)). In terms of transmission reliability, under a target Packet Loss Rate (PLR) of \begin{document}$ {10}^{-2} $\end{document} at \begin{document}$ N=1024 $\end{document}, the supportable normalized load upper bound is extended from \begin{document}$ G\approx 0.75 $\end{document} in baseline PSA to \begin{document}$ G\approx 0.84 $\end{document} in PSA-PNC (Fig. 5(a)). Furthermore, steeper waterfall regions and significantly lower error floors are consistently maintained across various frame lengths from \begin{document}$ N=64 $\end{document} to \begin{document}$ N=1024 $\end{document} (Fig. 5(b)).  Conclusions  A joint decoding scheme combining Polar Slotted ALOHA with Physical-Layer Network Coding (PSA-PNC) is established to resolve the severe decoding stalling problem in grant-free random access. By constructing a closed-loop iterative receiver integrating packet-level polar decoding, GMI-based algebraic equation solving, and iterative SIC cancellation, destructive multi-user collisions are converted into solvable linear equations. The dependence on collision-free slots is effectively eliminated, and the supportable load threshold, peak normalized throughput, and packet recovery reliability are substantially enhanced under heavy traffic loads. Future research will be directed toward non-ideal multipath fading channels, such as Rayleigh fading, and the design of low-complexity sparse receiver architectures for practical massive access implementations.
A knowledge distillation framework for hypergraph neural networks with rapid inference capabilities
LI Junzheng, YU Hongtao, HUANG Ruiyang, JIANG Haocong, LIU Shuo, YANG Suchang
Available online  , doi: 10.11999/JEIT260694
Abstract:
  Objective  Hypergraph Neural Networks (HGNNs) have gained widespread attention for their strong ability to model high-order correlations among entities, but their computational complexity and memory consumption grow exponentially as the hypergraph scale expands, severely restricting their deployment in large-scale industrial scenarios. Existing knowledge distillation methods that distill HGNNs into Multi-Layer Perceptrons (MLPs) are troubled by poor interpretability, low accuracy, severe information loss caused by Softmax-based soft labels, and neglect of node reliability heterogeneity. To address these critical challenges, this paper proposes a novel hypergraph knowledge distillation framework named DH2KAN (Distill Hypergraph Neural Network to Kolmogorov-Arnold Network) for fast inference, which breaks through the bottlenecks of traditional hypergraph knowledge distillation.  Methods  This paper designs a three-module knowledge distillation framework DH2KAN(图2). Firstly, we replace the traditional MLP with KAN as the student model, which uses learnable spline-based univariate functions instead of fixed activation functions and linear weights to improve the fitting ability and interpretability of the student model. Secondly, we propose a representation similarity distillation mechanism, which directly aligns the pre-logits representations of HGNN and KAN to avoid information loss caused by the Softmax normalization layer and completely retain the high-order structural knowledge of hypergraphs. Thirdly, we introduce a high-reliable node-aware distillation method(图3), which quantifies the node reliability by information entropy variation, screens out robust nodes with strong anti-noise ability, and takes their soft labels as the core supervision signal to improve the purity of distilled knowledge.  Results and Discussions  The DH2KAN algorithm achieves remarkable performance under both transductive learning (表2) and production learning (表3) settings. Quantitative experimental results reveal that DH2KAN obtains an accuracy improvement of approximately 10.53% over the vanilla KAN student model, 1.73% over the teacher HGNN model, and 1.2% over existing MLP-based distillation methods. Such results verify the effectiveness of knowledge transfer from HGNN to KAN, and demonstrate that the proposed method outperforms conventional MLP-oriented distillation schemes in inference performance. In addition, DH2KAN achieves optimal performance on feature-dominated hypergraph datasets, and possesses strong robustness when handling sparse structures and noisy node samples.  Conclusions  This paper proposes DH2KAN to accelerate HGNN inference for large-scale low-latency applications. Via knowledge distillation, it bridges the performance gap between KAN and HGNN and eliminates structural dependence for efficient reasoning. With representation similarity and reliable node-aware distillation, it transfers effective task knowledge via pre-logit features and soft labels, showing great practical application potential.
Research on Covert Communication Transmission Scheme Combining Relay Selection and Mode Selection over Nakagami-m Fading Channels
HUANG Haiyan, HUANG Yi, ZHANG Ning, LIANG Linlin, ZHANG Xuejun
Available online  , doi: 10.11999/JEIT260287
Abstract:
  Objective  Covert communication enhances the security of wireless communication systems by concealing both transmitted information and communication activities from unauthorized detection. However, practical wireless channels exhibit random and uncertain propagation conditions. The Nakagami-m fading channel, which can characterize a wide range of channel conditions, provides a realistic framework for evaluating the performance of covert communication. Relay-assisted transmission has attracted considerable attention because it improves transmission reliability over fading channels. Moreover, relay selection and transmission mode selection substantially affect system performance. Therefore, investigating their combined effect on covert communication over Nakagami-m fading channels is of both theoretical and practical significance for the design of next-generation secure wireless communication systems.  Methods  This paper proposes a covert communication system incorporating relay selection and transmission mode selection. The source node transmits covert information to the destination node through multiple relays, while a warden monitors transmissions from both the source and relay nodes. A friendly jammer transmits interference signals to degrade the warden’s detection capability. Four transmission schemes are considered: optimal relay selection with fixed Half-Duplex (HD) or Full-Duplex (FD) operation, optimal relay selection with random transmission mode selection, random relay selection with optimal transmission mode selection, and joint optimal relay and transmission mode selection. Closed-form expressions for the warden’s detection error probability under both HD and FD optimal relay selection are derived over Nakagami-m fading channels. Closed-form expressions for the transmission outage probability, asymptotic transmission outage probability, and covert rate are also derived for all transmission schemes. The theoretical analysis is validated through MATLAB simulations.  Results and Discussions  Simulation results demonstrate that an optimal detection threshold exists that minimizes the detection error probability (Fig. 2). As the detection threshold or jamming power increases, the warden’s ability to detect covert communication decreases, causing the detection error probability to approach one (Figs. 2 and 3). Under the same target transmission rate and high Signal-to-Noise Ratio (SNR) conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability, thereby providing the highest transmission reliability (Figs. 4 and 5). At a target transmission rate of \begin{document}$ \text{6.5 bit/(s}\cdot \text{Hz)} $\end{document}, the transmission outage probability of the joint relay and transmission mode selection scheme is 6.9% lower than that of the FD transmission scheme (Fig. 4). As the transmit power and the number of relays increase, the covert rate gradually approaches a constant value. Among all transmission schemes, the joint relay and transmission mode selection scheme consistently achieves the highest covert rate (Figs. 6 and 7).  Conclusions  This paper proposes a covert communication system based on relay selection and transmission mode selection over Nakagami-m fading channels. Closed-form expressions for the warden’s detection error probability and the system’s transmission outage probability are derived under different relay selection and transmission mode selection strategies. The asymptotic transmission outage probability and covert rate are then analyzed. Simulation results show that increasing the detection threshold or jamming power weakens the warden’s ability to detect covert communication, causing the detection error probability to approach one. Under identical target transmission rates and high SNR conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability. These results indicate that appropriate relay selection and transmission mode selection not only reduce the warden’s detection capability and protect covert communication, but also improve both transmission reliability and covertness. Future work will consider practical factors, including imperfect channel state information, residual self-interference, and incomplete knowledge of the warden’s channel.
Research on Channel Multipath Prediction Based on an Environmental Graph
ZHANG Zhaoling, JIN Jing, ZHAO Jingbo, YU Li, CAI Yichen, MA Liang, ZHANG Jianhua
Available online  , doi: 10.11999/JEIT260416
Abstract:
  Objective  Environment-driven channel prediction requires a structured representation that connects physical objects in a propagation environment with the resulting multipath topology. However, conventional data-driven methods generally treat environmental information as unstructured global features and have difficulty representing the interactions among the transmitter (Tx), receiver (Rx), and surrounding scatterers. This study investigates whether an environmental graph can provide an effective intermediate representation for identifying effective scatterers and predicting candidate propagation paths.  Methods  An environmental graph is constructed by representing the Tx, Rx, and scatterers as graph nodes. Spatial distance and visibility between nodes are encoded as edge features to describe their geometric relationships. A ScatterGNN based on the Edge-aware Graph Isomorphism Network (EGIN) is developed to extract structural features and identify effective scatterers involved in signal propagation. Candidate single- and two-bounce paths are subsequently generated from the detected scatterers. A path-ranking network, PathRankingNet, is then designed to estimate the validity scores of candidate paths and rank them using a Listwise Ranking Loss. The proposed framework is evaluated in a controlled indoor Industrial Internet of Things scenario generated using Wireless InSite. The scenario covers an area of 150 m × 63 m × 22 m and contains 3,381 Rx sampling locations.  Results and Discussions  For effective scatterer detection, the proposed method achieves an average precision of 0.957 9, while 77.1% of the test samples achieve complete detection of all effective scatterers. For candidate propagation-path prediction, the model converges stably after approximately 30 training epochs. Precision@3 ranges from 0.60 to 0.64, Recall@3 ranges from 0.42 to 0.45, and Hit@3 ranges from 0.90 to 0.92. These results indicate that the environmental graph preserves useful structural information related to propagation-path topology. In particular, the model retains at least one ground-truth propagation path among the three highest-ranked candidates for more than 90% of the test samples.  Conclusions  The proposed framework provides a graph-based approach for transforming environmental geometry into structured representations of effective scatterers and candidate propagation paths. Rather than replacing ray tracing or channel measurements, the method is intended to reduce the candidate search space before detailed path-parameter calculation or channel reconstruction. The current results demonstrate its feasibility within a single simulated environment and primarily reflect its ability to approximate the propagation-path topology labels generated by Wireless InSite. Further validation using independent environments, measured channel data, and path-level parameters such as power, delay, and complex gain is required before its cross-scenario generalization and practical applicability can be established.
An Incremental Density-Based Clustering Method with TDOA Prior for Mobile Multi-Station Radar Signal Sorting
CHEN Jinli, FAN Yu, WANG Yanjie, ZHANG Jindong
Available online  , doi: 10.11999/JEIT260151
Abstract:
  Objective  Radar signal sorting is a core technology in electronic reconnaissance that aims to deinterleave and classify pulses from multiple radar emitters within dense, overlapping pulse streams. In practical reconnaissance missions, especially those employing mobile platforms such as aircraft, observation stations continuously change position, whereas radar emitters are typically stationary. The resulting variation in the relative geometry causes the Time Difference of Arrival (TDOA) of intercepted signals to evolve over time. Conventional multi-station radar signal sorting methods are generally developed under a static TDOA assumption and cannot effectively characterize this temporal evolution. Therefore, pulses from the same radar emitter are easily split into multiple clusters, resulting in cluster proliferation and degraded sorting performance. To address this problem, an Incremental Density-Based Clustering (ICDC) method with TDOA prior information is proposed for mobile multi-station radar signal sorting. The proposed method exploits the temporal evolution of TDOA to improve sorting stability and accuracy in mobile multi-station scenarios.  Methods  The temporal evolution of TDOA in mobile multi-station scenarios is first analyzed, and the TDOA trajectory is approximated as a linear function of time to establish a linear prior state model. Online micro-clusters are then constructed from multi-station TDOA observations and aggregated into macro-clusters according to their spatiotemporal intersection relationships. For each macro-cluster, an independent Kalman Filter (KF) model is established for each TDOA dimension. The state vector consists of the TDOA value and its rate of change, and recursive state estimation is performed to provide dynamic TDOA priors for newly arriving samples. During incremental clustering, a spatiotemporal joint scoring function is developed by incorporating KF prediction residuals as dynamic constraints. The matching criterion therefore evolves from a conventional density-based rule into a joint spatiotemporal consistency criterion, enabling more accurate assignment of newly arriving TDOA observations. To suppress cluster proliferation caused by TDOA evolution, a concept drift detection strategy based on macro-cluster center evolution is further employed. When concept drift is detected, posterior state estimates and covariance matrices generated by the KF are used to construct a dual Mahalanobis distance criterion that jointly evaluates state-distribution overlap and predicted TDOA overlap. Radar clusters produced by erroneous splitting are then adaptively merged under a 95% confidence threshold, effectively suppressing cluster proliferation caused by concept drift.  Results and Discussions  Simulation data are generated according to the radar parameters listed in Table 1. The multi-station TDOA corresponding to the same radar emitter exhibits an approximately linear evolution with respect to the Time of Arrival (TOA) at the observation station (Fig. 5), consistent with the proposed linear prior model. The ability of the proposed method to suppress radar cluster proliferation is first evaluated (Fig. 6). Radar emitters E4 and E9 are selected from the nine simulated emitters as representative cases because their TDOA trajectories are closely spaced and difficult to separate. Compared with Density-Based Spatial Clustering of Applications with Noise (DBSCAN), ICDC, cloud model-based sorting, and PointNet++ sorting, the proposed method more effectively suppresses radar cluster proliferation and maintains greater cluster stability. The overall sorting performance is further compared with the histogram method, grid-based clustering, DBSCAN, ICDC, cloud model-based sorting, and PointNet++ sorting. When the TOA measurement error ranges from 50 to 300 ns (Fig. 7), the proposed method consistently achieves a sorting accuracy above 96%, demonstrating strong robustness to measurement errors. Under different pulse interference rates (Fig. 8), the sorting accuracy also remains above 96%, indicating excellent interference robustness. The performance under different observation station velocities is further evaluated (Fig. 9). The proposed method maintains high sorting accuracy over the entire velocity range and still achieves approximately 94% accuracy at relatively high observation station velocities, demonstrating strong robustness under dynamic observation conditions. Radar cluster proliferation probability, missed-cluster probability (Fig. 10), and computational complexity (Table 2) are also analyzed. The results demonstrate that the proposed method achieves a favorable balance among sorting accuracy, cluster proliferation suppression, missed-cluster control, and computational complexity.  Conclusions  An ICDC method with TDOA prior information is proposed to address the degradation of radar signal sorting performance caused by TDOA evolution in mobile multi-station scenarios. By incorporating observation station motion into a dynamic TDOA state model and applying KF-based recursive prediction, the proposed method effectively suppresses radar cluster proliferation and erroneous cluster splitting caused by concept drift. The spatiotemporal joint criterion and the adaptive cluster merging strategy further improve robustness in complex dynamic environments. Simulation results demonstrate the effectiveness and stability of the proposed method for mobile multi-station cooperative reconnaissance. Future work will focus on real-time multi-parameter fusion-based sorting in complex electromagnetic environments. Furthermore, adaptive estimation and adjustment of the micro-cluster spatial intersection threshold, process noise covariance matrix, and measurement noise variance will be investigated to further improve performance in complex scenarios.
Difference-aware Adaptive Prompt Learning and Dense Alignment for Weakly Supervised Building Change Detection
CHEN Yanxia, MA Longlong, CHEN Yanhua, HUANG Yuchun
Available online  , doi: 10.11999/JEIT260595
Abstract:
  Objective  Building change detection from bi-temporal high-resolution remote sensing images is important for urban planning, land resource management, illegal construction monitoring, and disaster damage assessment. Existing fully supervised change detection methods generally achieve high detection accuracy but require pixel-level annotations. However, obtaining pixel-level labels for large-scale remote sensing images is labor-intensive and time-consuming, which limits their application to large-scale monitoring scenarios. Image-level weakly supervised change detection reduces annotation costs by using only image-level labels indicating whether an image pair contains changes. However, the lack of spatial supervision makes accurate localization of changed regions difficult. Existing weakly supervised methods generally rely on Class Activation Maps (CAMs) to generate pseudo labels. CAMs tend to highlight only the most discriminative regions, resulting in incomplete coverage of changed areas or background noise. Vision-language models provide semantic priors for weakly supervised learning. However, directly applying Contrastive Language-Image Pre-training (CLIP) to change detection remains challenging. Fixed text prompts are difficult to adapt to the difference semantics of bi-temporal images, and the original CLIP objective mainly focuses on global image-text alignment rather than local pixel-level localization. To address these problems, a Difference-aware Adaptive Prompt Learning and Dense Alignment method for weakly supervised building change detection, termed DAPL-CD, is proposed.  Methods  The proposed framework introduces CLIP-based cross-modal semantic knowledge into image-level weakly supervised building change detection. For a pair of bi-temporal remote sensing images, a shared CLIP visual encoder is first used to extract visual representations from the two temporal images. The local visual features are fused along the channel dimension to obtain bi-temporal difference features containing semantic information related to changed and unchanged regions. Based on the difference characteristics of building change detection, a difference-aware adaptive prompt learning strategy is designed. Instead of using manually designed fixed text templates, learnable context vectors are inserted into the text prompts while preserving category-related semantic words. The resulting foreground and background text embeddings are used as foreground and background text prototypes to provide adaptive semantic guidance for change localization. Furthermore, a pixel-text dense alignment mechanism is introduced to extend CLIP’s global image-text alignment capability to local feature matching. The initial CAM generated by the classification branch is used to obtain preliminary foreground and background regions. Visual-text positive and negative sample pairs are then constructed between local difference features and the foreground and background text prototypes. An InfoNCE-based dense alignment loss is used to pull matched visual and textual features closer and push mismatched features apart. Finally, the classification and segmentation branches are jointly optimized using the classification loss, global alignment loss, dense alignment loss, and segmentation loss, with the segmentation loss introduced only after the quality of the generated CAMs has stabilized.  Results and Discussions  Experiments are conducted on WHU-CD and LEVIR-CD, two public benchmark datasets for building change detection. Only image-level labels are used during training, whereas pixel-level annotations are used only for evaluation. Overall Accuracy (OA), F1-score, and Intersection over Union (IoU) are adopted as the main evaluation metrics. Because changed buildings usually occupy a small proportion of remote sensing images, OA can be strongly affected by the dominant unchanged background pixels. Therefore, F1-score and IoU are emphasized for evaluating the detection quality of changed regions. Quantitative comparisons show that DAPL-CD achieves an OA of 94.7%, an F1-score of 82.8%, and an IoU of 70.6% on WHU-CD, and an OA of 92.3%, an F1-score of 68.0%, and an IoU of 51.5% on LEVIR-CD. The method achieves the best F1-score and IoU among the compared weakly supervised change detection methods (Table 1). Visual comparisons further show that the proposed method produces more complete responses for large-scale building changes and more continuous predictions for small and scattered changed buildings (Figs. 3 and 4). Ablation experiments verify the effectiveness of difference-aware adaptive prompt learning and pixel-text dense alignment. The baseline model using fixed text prompts without foreground or background alignment achieves an F1-score of 63.2% and an IoU of 46.2%. Introducing both foreground and background alignment increases these metrics to 66.3% and 49.6%, respectively, indicating that dense semantic matching between local visual features and text prototypes improves the discrimination of changed regions. After difference-aware adaptive prompt learning is incorporated, the F1-score and IoU increase to 68.0% and 51.5%, respectively, indicating that learnable context vectors reduce the semantic mismatch between fixed text descriptions and bi-temporal difference features. Under the adaptive-prompt setting, foreground alignment alone achieves an F1-score of 64.4% and an IoU of 47.5%, whereas background alignment alone achieves 65.3% and 48.5%, respectively. Combining the two alignment branches yields the best F1-score and IoU, indicating that foreground and background semantic constraints provide complementary guidance for change localization (Table 2). The CAM results further show that pixel-text dense alignment produces stronger and more complete responses over actual changed regions while suppressing irrelevant background activations (Fig. 5).  Conclusions  A weakly supervised building change detection framework based on difference-aware adaptive prompt learning and pixel-text dense alignment is proposed. By introducing CLIP-based cross-modal semantic priors, the proposed method converts text-level semantic knowledge into local change localization capability. The difference-aware adaptive prompt learning strategy improves the representation of change-related semantic descriptions, whereas the pixel-text dense alignment mechanism establishes direct correspondence between local difference features and foreground and background text prototypes. Experimental results on WHU-CD and LEVIR-CD demonstrate that DAPL-CD achieves high performance under image-level supervision and improves the completeness and accuracy of changed building localization. The proposed framework provides an effective approach for reducing annotation requirements in large-scale remote sensing change detection. Future research will focus on improving pseudo-label reliability, reducing dependence on large-scale pre-trained models, and extending the method to multi-temporal and multi-spectral remote sensing data.
A Multi-Dimensional Scenario-Based Evaluation Method for Deep Learning Side-Channel Analysis Using a Multi-Attribute Decision Model
GU Zepeng, CHEN Lin, CAI Juesong, YAN Yingjian
Available online  , doi: 10.11999/JEIT260198
Abstract:
  Objective  Deep Learning Side-Channel Analysis (DL-SCA) has substantially improved the effectiveness of attacks against protected cryptographic implementations. However, the transition of DL-SCA models from research to practical deployment is limited by the lack of systematic, fair, and scenario-specific evaluation methods. Existing evaluations mainly rely on Guessing Entropy (GE) and Success Rate (SR), while overlooking practical factors such as resource overhead and environmental adaptability. Moreover, inconsistent hyperparameter optimization leads to unfair model comparisons and provides limited quantitative guidance for model selection under different deployment constraints, including resource-constrained devices, high-noise environments, and real-time applications. This paper proposes a systems engineering-based evaluation framework that enables comprehensive, quantitative, and scenario-specific assessment of DL-SCA models.  Methods  A multi-dimensional, scenario-based evaluation framework is developed using systems engineering principles. First, a hierarchical evaluation index system is established, comprising three criteria—attack effectiveness, resource overhead, and environmental adaptability—and six evaluation metrics: GE, SR, training time (TC), peak memory consumption (MC), model complexity (MoC), and noise robustness (Rob). Second, a standardized evaluation process based on the V-model is designed to ensure fair comparison. Each candidate model, including a Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), and CNN-LSTM hybrid model, undergoes independent hyperparameter optimization using grid search before multi-dimensional performance evaluation. Third, a hybrid Criteria Importance Through Intercriteria Correlation-Analytic Hierarchy Process (CRITIC-AHP) Multi-Attribute Decision-Making (MADM) framework is developed. The CRITIC method derives objective weights from the statistical characteristics of the evaluation data, whereas the AHP method incorporates scenario-specific preferences through pairwise comparison matrices. The objective and subjective weights are fused to generate scenario-specific weights. Finally, a Multi-dimensional Attack Performance Metric (MAPM) is defined as the weighted sum of normalized evaluation metrics using the fused weights, providing a composite score for each model under a specific deployment scenario.  Results and Discussions  The proposed framework is validated using the ASCAD fixed-key dataset. After independent hyperparameter optimization, the three model architectures are evaluated using all six metrics. The CRITIC method produces the objective weight vector W critic = [0.17, 0.19, 0.15, 0.21, 0.14, 0.14]. Four representative deployment scenarios—Resource-Constrained, High-Performance, High-Noise, and Real-Time—are then defined, and the corresponding AHP preference weights are fused with the objective weights to generate the final scenario-specific weights. For example, MC receives the highest weight (0.52) in the Resource-Constrained scenario, whereas Rob dominates the High-Noise scenario with a weight of 0.57. The resulting MAPM scores (Table 9, Fig. 9, and Fig. 10) clearly differentiate the strengths of the evaluated models and demonstrate the scenario-specific decision capability of the proposed framework. CNN achieves the highest score in the High-Performance scenario (0.894), MLP ranks first in the Real-Time scenario (0.758) because of its shortest training time, and the CNN-LSTM hybrid model performs best in the High-Noise scenario (0.863) because of its superior noise robustness despite higher resource overhead. These results demonstrate that no single model is optimal across all deployment scenarios and that MAPM provides a clear and quantitative basis for model selection under specific deployment constraints.  Conclusions  This paper proposes a systems engineering-based, multi-dimensional evaluation framework to address the major limitations of current DL-SCA model assessment. By integrating a hierarchical evaluation index system, a standardized V-model evaluation process, and a hybrid CRITIC-AHP Multi-Attribute Decision-Making (MADM) framework, the proposed method quantitatively balances the trade-offs among attack effectiveness, resource overhead, and environmental adaptability. Experimental results obtained using the ASCAD benchmark demonstrate that the framework provides clear, quantitative, and scenario-specific guidance for model selection. The proposed Multi-dimensional Attack Performance Metric (MAPM) provides a practical decision basis for selecting DL-SCA models under diverse deployment constraints, narrowing the gap between academic attack development and practical model deployment. Future work will extend the framework to additional model architectures and datasets, improve evaluation automation, and validate its effectiveness in practical deployment environments.
Research on LEO Constellation Interference Prediction and Detection Algorithm Driven by Collaborative Spatial Feature Mapping and Temporal Transformer
YANG Boyu, QIU Kun, CHEN Zhe, ZHAO Jin, GAO Yue
Available online  , doi: 10.11999/JEIT260368
Abstract:
  Objective  The rapid deployment of large-scale Low Earth Orbit (LEO) satellite constellations has intensified competition for orbital and spectrum resources. To improve spectrum utilization, different LEO satellite constellations commonly employ co-frequency reuse, which increases the risk of inter-system co-frequency interference. The high-speed motion of LEO satellites and their highly dynamic, heterogeneous topology further cause rapid variations in interference power, resulting in severe co-frequency interference. Existing interference assessment methods mainly rely on the regulations of the International Telecommunication Union (ITU), with the Interference-to-Noise Ratio (I/Noise) as a key evaluation metric. However, conventional ITU-based physical-iteration methods require repeated calculations of satellite positions, link attenuation and interference contributions for all visible satellites, causing computational cost to increase rapidly with constellation size. A complete interference assessment for a single ground station in a constellation of approximately 10 000 satellites can require more than 30 h. Recent deep-learning-based methods still commonly use all visible satellites as inputs and therefore do not adequately exploit the spatial sparsity of interference features. To address these limitations, an LEO constellation interference prediction and detection method driven by collaborative spatial feature mapping and temporal Transformer is proposed to reduce computational cost while maintaining accurate interference prediction and detection.  Methods  A spatiotemporal framework is developed to model dynamic co-frequency interference in large-scale heterogeneous multi-constellation LEO networks. Based on the ITU physical interference model, three key properties of the interference function are identified: permutation invariance, spatial sparsity and temporal continuity. The conventional physical-iteration process is therefore reformulated as a spatiotemporally decoupled feature-mapping problem. A spatial feature mapping module with adaptive attention and symmetric pooling is constructed to compress unordered and variable-length satellite interference-source features into fixed-dimensional, permutation-invariant spatial features. The attention mechanism adaptively focuses on dominant interference sources, whereas symmetric pooling using maximum and mean pooling captures extreme and global statistical characteristics while suppressing redundant satellite nodes. A temporal Transformer is then employed to model the long-range evolution of interference trajectories. The future I/Noise trajectory is predicted to support rapid interference detection under dynamic constellation configurations.  Results and Discussions  A large-scale heterogeneous LEO constellation scenario consisting of 6 800 Starlink satellites and 650 OneWeb satellites is simulated according to ITU regulations. Real Two-Line Element (TLE) data are used to propagate satellite orbits, and 100,000 time-series samples are generated at 0.5 s intervals. Parameter analysis shows that appropriate feature dimensions and encoder depths provide a favorable balance between feature extraction accuracy and computational cost (Figs. 4 and 5). The sampling interval and sliding-window length are further optimized to balance prediction accuracy and real-time inference performance (Figs. 6 and 7). Compared with the baseline methods, the proposed method produces interference trajectories that closely follow the physical-iteration ground truth (Fig. 8). At a cumulative probability of 90%, the absolute prediction error is maintained within 0.5 dB (Fig. 9). The method also maintains a high interference recall under a low false-alarm-rate constraint (Fig. 10). For a 20.0 s prediction horizon, the Root Mean Square Error (RMSE) remains at 0.45 dB, substantially lower than those of the baseline models. At a 20 s prediction step, the RMSE values of the Multi-Layer Perceptron (MLP) and Long Short-Term Memory (LSTM) network increase to 5.53 dB and 4.82 dB, respectively (Fig. 11). In terms of computational efficiency, the proposed method requires 14.2 ms for a single inference, which is only 4.1% of the computational time required by the ITU-based physical-iteration method (Table 3). The computational cost is therefore effectively decoupled from constellation size.  Conclusions  To address the high computational cost caused by the highly dynamic topology of LEO satellite networks, a collaborative spatial feature mapping and temporal Transformer-based interference prediction and detection method is proposed. The spatial feature mapping module compresses variable-length interference-source features into fixed-dimensional representations, whereas the temporal Transformer captures the long-range evolution of interference trajectories. Simulation results show that the proposed method provides accurate long-term trajectory tracking while remaining consistent with ITU-based interference assessment. Parameter-sensitivity experiments demonstrate that the proposed method can balance feature extraction accuracy and computational cost under different configurations. With its long-range dependency modeling capability, the temporal Transformer maintains an RMSE of 0.45 dB over a 20 s prediction horizon. By filtering redundant satellite nodes and decoupling computational cost from constellation size, the method reduces single-inference time to 14.2 ms. Comprehensive evaluations of prediction error distributions, detection performance and model parameters demonstrate that the proposed method achieves high prediction accuracy and substantially reduced computational complexity, providing a flexible engineering solution for interference monitoring in large-scale LEO constellation deployments.
A Parametric Architecture Description Framework for Embedded FPGAs and Multi-objective QoR-driven Architecture Exploration
ZHOU Jing, ZHANG Shengbing, CHEN Lei, FENG Hanxu, WANG Shuo, TIAN Chunsheng
Available online  , doi: 10.11999/JEIT260609
Abstract:
  Objective  Architecture parameters of commercial off-the-shelf Field-Programmable Gate Arrays (FPGAs) are fixed by vendors and reused across products. Embedded FPGAs (eFPGAs), in contrast, allow architects to select architecture parameters according to specific application requirements. The LUT input count K, the number of LUTs per cluster N, the interconnect topology, and the types of heterogeneous tiles can therefore be configured for the target application. Architecture design space exploration thus becomes an engineering task in which tens to hundreds of architectures may need to be generated and evaluated before a suitable configuration is selected. Existing architecture description practices rely largely on manually written architecture files and batch scripts. After each parameter change, shared fields across multiple backend toolchains must be updated and aligned manually, making large-scale architecture exploration difficult to support. To address this problem, a parametric architecture description framework is proposed for unified description across multiple backend toolchains. Architecture parameters are organized into three layers according to their independence: process invariants, coupled parameters, and independent parameters. Architecture descriptions for different backends are automatically derived from the same source object through independent derivation functions. The framework currently supports VPR, OpenFPGA, and Yosys and has been extended to COFFE. Its operation is validated across the complete toolchain. Based on the framework, a parameter-sweep design space exploration method is developed, and an open Quality of Results (QoR) dataset covering five application domains and 64 benchmark circuits, including homogeneous and heterogeneous architectures, is released as a public benchmark.  Methods  HorizonArch, the proposed parametric architecture description framework, organizes architecture parameters into three layers according to parameter independence (Table 1). L0 contains process invariants that are fixed once the technology node is determined. L1 contains coupled parameters, including K, N, tier, segment length, switch block type, and channel connectivity, for which a single parameter change can trigger updates across multiple fields and backend architecture descriptions. L2 contains independent parameters that can be specified separately. Architecture construction is formalized by an operator B that maps a parameter vector p to a complete architecture object (Fig. 4). Five formal rules are imposed: parameter completeness (R1), fragment independence (R2), type compatibility (R3), explicit coupling (R4), and static checkability (R5). Each backend architecture description is then derived by an independent derivation function from the same source object. Thus, adding a new backend requires only an additional view rather than modifications throughout the existing description structure. Field-level validation and cross-field validation are performed when the architecture object is loaded, before any backend tool is invoked. The class structure (Fig. 3) divides the architecture description into synthesis, circuit, and layout views, with each semantic element declared only once. Three extension levels are defined: G1 adds a black-box model, G2 extends the value set of an existing coupled parameter, and G3 adds a new coupled parameter together with its constrained value set. Based on this framework, a parameter-sweep design space exploration method is developed to scan the (K, N) parameter grid and heterogeneous tile configurations. Each configuration is evaluated using three QoR metrics: area, Critical-Path Delay (CPD), and Area-Delay Product (ADP).  Results and Discussions  End-to-end validation shows that a single source description consistently generates architecture descriptions for VPR, OpenFPGA, and Yosys. COFFE is connected and verified at the interface layer, including SPICE simulation startup (Table 4, Table 5). The G1, G2, and G3 extension experiments pass all cross-field checks. The design space exploration results show different preferences among area, CPD, and ADP across the (K, N) parameter space (Figs. 5 and 6). Area favors smaller K values, with K=4 and N=4 providing favorable area and ADP performance for a large proportion of circuits. CPD, in contrast, favors larger K and N values, with optimal configurations concentrated near (K, N)=(8, 10) and (7, 10). Across the twenty (K, N) configurations, the relative-range distribution shows that parameter selection has a much greater effect on area and ADP than on CPD (Table 6). The mean relative ranges are 50.9% for CPD, 675.1% for area, and 516.3% for ADP. A comparison of default configurations (Table 7) shows that K=4 and N=4 achieves the minimum ADP for 67.9% of the circuits and has an average ADP deviation of 5.8%, although its average CPD deviation reaches 47.9%. In contrast, K=8 and N=10 reduces the average CPD deviation to 5.9%, with 30.8% of the circuits achieving the CPD optimum, but increases the average area and ADP deviations to 673.2% and 487.0%, respectively. A random-forest cross-domain surrogate achieves a top-5 accuracy of approximately 65%. Therefore, parameter sweeping remains necessary when strict design targets are imposed.  Conclusions  HorizonArch, a parametric architecture description framework for eFPGA exploration, is developed and validated. The framework generates VPR, OpenFPGA, and Yosys backend architecture descriptions from a single source object and provides an extensible interface for COFFE. The parameter-sweep exploration shows that area and CPD favor opposite regions of the (K, N) parameter space. Therefore, eFPGA architecture parameters should be selected according to explicit design targets rather than fixed default values. An open QoR dataset covering five application domains and 64 benchmark circuits is also released as a reusable benchmark for eFPGA architecture design space exploration. Future work will complete the COFFE SPICE topology-rewriting component, refit the routing-area coefficient using measured data, and explore more efficient design space exploration strategies.
A Reconfigurable Parallelized Coprocessor Design for the RISC-V-based Grain Cryptographic Algorithm
NAN Longmei, WANG Haoyu, DU Yiran, LI Wei, CHEN Tao
Available online  , doi: 10.11999/JEIT260391
Abstract:
  Objective  To address the performance limitations of Grain cryptographic algorithms on General-Purpose Processors (GPPs), as well as the inflexibility and high hardware overhead of Application-Specific Integrated Circuit (ASIC) implementations, a dedicated cryptographic hardware accelerator is integrated into an RISC-V coprocessor through a custom instruction extension mechanism. A reconfigurable parallelized architecture is proposed for the Grain algorithm family based on the RISC-V coprocessor interface. Corresponding custom instructions are designed to support the flexible and efficient execution of Grain-80, Grain-128, Grain-128a, and Grain-128AEAD on a unified hardware platform. The proposed architecture provides a favorable balance between processing efficiency, design flexibility, and hardware resource overhead, making it suitable for resource-constrained embedded systems.  Methods  A unified shift-register architecture is adopted to support flexible switching among Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, with a configurable parallelization degree of 1 to 8. To implement the custom instructions for the Grain cryptographic algorithms, the software and hardware functions are analyzed, and the cryptographic process is divided between the processor and coprocessor to achieve efficient execution. The proposed coprocessor uses a streamlined architecture that reuses common hardware resources across the four algorithms. Combined with a configurable feedback tap selection network and feedback logic tailored to a predefined set of algorithms, the architecture enables reconfigurable parallel execution with limited additional hardware resource overhead.  Results and Discussions  The proposed custom instructions enable flexible implementation of four Grain cryptographic algorithms on the same hardware platform. Compared with implementations without instruction extensions, the proposed custom instructions reduce the number of clock cycles and instructions required for cryptographic processing while improving throughput (Table 5). On the Hummingbird E203 platform with a parallelization degree of 4, the proposed implementation requires 105, 129, and 183 clock cycles for Grain-80, Grain-128/Grain-128a, and Grain-128AEAD, respectively, with corresponding throughputs of 220.67, 179.74, and 126.70 Mbit/(s·Hz). Compared with purely software-based implementations, the proposed approach substantially reduces both the number of executed instructions and the number of clock cycles (Table 5). Synthesis results further demonstrate that the coprocessor occupies 7 252.55 μm² in a 65 nm process and achieves a throughput of 1.56 Gbit/(s·Hz) at a parallelization degree of 4. Although the reconfigurable architecture requires slightly more area than a dedicated single-algorithm implementation, it supports four Grain algorithms on a unified hardware platform and provides improved hardware resource reuse and design flexibility.  Conclusions  A hardware-software cooperative reconfigurable parallelization scheme is designed to accelerate the Grain cryptographic algorithm family in lightweight embedded systems. The scheme exploits the RISC-V custom instruction extension mechanism and combines a unified shift-register architecture, a configurable feedback tap selection network, and feedback logic tailored to a predefined set of algorithms. Four algorithms, namely Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, can therefore be flexibly selected and processed in parallel on a single hardware platform. This design improves processing efficiency while maintaining design flexibility and limiting hardware resource overhead. The present work focuses on reconfigurable parallelization for the Grain algorithm family. Future research will investigate more general reconfigurable architectures for nonlinear Boolean functions by incorporating configurable units, such as LookUp Tables (LUTs) or programmable logic arrays. Such architectures may further improve compatibility with multiple stream cipher algorithms while maintaining high throughput and low hardware resource overhead.
Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng
Available online  , doi: 10.11999/JEIT260442
Abstract:
  Objective  Accurate grasp detection from language instructions is essential for service robots to achieve natural human-robot interaction. Existing methods primarily rely on large-scale data-driven training or hierarchical feature fusion to align visual perception with textual instructions. However, they generally overlook the strong coupling between target objects and background clutter in low-level visual features, leading to degraded compositional generalization under cross-view and unseen-scene conditions. To address this limitation, a dual-view cross-scene grasp detection and object localization dataset is constructed to systematically evaluate and improve the compositional generalization of existing models. Based on this benchmark, a Simultaneous Grasp detection and object Localization Network (SGL-Net) is proposed to jointly predict object locations and optimal grasp poses. The proposed framework enables service robots to manipulate objects according to natural language instructions in real-world dynamic environments, providing technical support for embodied intelligence.  Methods  The proposed SGL-Net is illustrated in Fig. 1. First, a Cross-modal Global Context Modulation Module (CGCMM) is proposed to exploit semantic priors from language instructions for adaptive viewpoint correction and background suppression during the early stage of visual feature extraction. Second, a Word-Pixel Cross-modal Alignment Module (WPCAM) is designed to achieve fine-grained semantic decoupling through a flattening-based cross-modal attention mechanism, thereby improving semantic understanding in complex dynamic scenes. Finally, a unified decoder jointly predicts object locations and optimal grasp poses from the fused multimodal features.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on the reconstructed dual-view cross-scene dataset containing bottom-view and top-view scenes and on a real-world robotic grasping platform. Comparative results demonstrate that SGL-Net consistently outperforms mainstream CNN-based and CLIP-based methods in both grasp detection and object localization (Tables 2 and 3). Ablation studies further verify the effectiveness of CGCMM and WPCAM in improving fine-grained semantic alignment and semantic decoupling (Tables 4 and 5). Furthermore, qualitative results (Figs. 46) and real-world robotic experiments (Fig. 7) demonstrate that SGL-Net can be reliably deployed in complex physical environments. Overall, the proposed network exhibits strong generalization capability and excellent potential for practical robotic applications.  Conclusions  To improve cross-view and cross-scene generalization in language-conditioned robotic grasp detection, this paper constructs a dedicated validation dataset and proposes SGL-Net, which jointly performs grasp detection and object localization. By integrating CGCMM and WPCAM, the proposed network accurately localizes instruction-specified objects and predicts optimal grasp poses. Experimental results obtained on multiple benchmark scenarios and a real-world robotic platform demonstrate the superior performance and practical applicability of the proposed method. Future work will focus on integrating Large Multimodal Models (LMMs) and adapting the proposed framework through fine-tuning to further improve zero-shot robotic grasp detection.
Physics-Aware Reconstruction for MilliMeter-Wave Radar Gait Recognition Under Complex Wearing Scenarios
HUANG Ling, QIU Liying, WANG Jiacheng, HAN Penglin, ZHOU Qingdi, YAN Huimei
Available online  , doi: 10.11999/JEIT260522
Abstract:
  Objective  MilliMeter-Wave Radar (MMW Radar) gait recognition has demonstrated considerable potential for non-contact biometric identification because of its inherent advantages in privacy preservation and robustness to variable lighting conditions. However, practical deployment remains challenging under complex wearing conditions, such as long coats or backpacks. These external factors introduce non-stationary high-frequency interference, resulting in spectral aliasing and masking of intrinsic micro-Doppler (m-D) features. Conventional deep learning methods generally treat m-D spectrograms as generic images and overlook the physical relationship between Doppler frequency and human motion. Therefore, clothing-induced interference is easily confused with motion-related features, resulting in reduced recognition performance. This study proposes a physics-aware framework that integrates radar signal physics with human biomechanics to achieve frequency-domain decoupling and adaptive interference suppression for robust radar-based gait recognition under complex wearing conditions.  Methods  To reduce clothing-induced interference, this paper proposes PRISM-Net, a physics-aware gait recognition framework for MMW Radar. The framework is built on the biomechanical characteristics of human motion, where the torso, representing the primary body mass, generates relatively stable low-frequency Doppler components, whereas limb motion produces higher-frequency periodic components. (1) Physics-aware Frequency Structural Reconstruction: Instead of uniformly processing the entire m-D spectrogram, the proposed method performs Physics-aware Frequency Structural Reconstruction by exploiting the velocity distribution associated with different body parts. The original aliased m-D spectrogram is reconstructed into torso- and limb-related frequency components. Low-frequency components preserve stable identity-discriminative features, whereas high-frequency components characterize limb motion. This frequency-domain structural reconstruction isolates spectral regions that are most susceptible to clothing-induced interference, thereby reducing interference propagation. (2) Weighted Attention Mechanism (WAM): The WAM adaptively reweights feature responses according to the reliability of different frequency components. Because clothing-induced interference predominantly affects high-frequency regions, the WAM suppresses interference-contaminated high-frequency responses while enhancing stable torso features, thereby improving feature fusion and recognition robustness. (3) Experimental Configuration: The proposed method is evaluated on the MMRGait-1.0 dataset under a subject-independent evaluation protocol. The training set contains data from 74 subjects, while the remaining 47 unseen subjects are used for testing. All m-D spectrograms are resized to 224 × 224 pixels. The network is optimized using the AdamW optimizer with a joint loss comprising Cross-Entropy Loss, Triplet Loss, and Center Loss to improve both classification performance and feature discriminability.  Results and Discussions  Experimental results demonstrate the effectiveness of incorporating biomechanical priors into gait recognition. As shown in Table 2, PRISM-Net achieves an average Rank-1 accuracy of 89.7% under the 90° side-view condition. In the coat (CT) scenario, the proposed method maintains a Rank-1 accuracy of 85.1%, representing an 18.1% improvement over ShuffleNetV2 and a 7.4% improvement over the ResNet-18 baseline. Model stability is verified through ten independent trials. As illustrated in Fig. 1, PRISM-Net achieves a standard deviation of ±0.55%, compared with ±1.75% for the baseline model. An independent-samples t-test yields p<0.001, confirming that the improvement is statistically significant. Ablation results in Table 2 further verify the contribution of each component. Removing Physics-aware Frequency Structural Reconstruction reduces the CT Rank-1 accuracy to 75.5%, demonstrating the importance of physics-aware frequency-domain decoupling for preventing feature distortion. The WAM further improves CT accuracy by 4.2% through adaptive suppression of high-frequency interference. Regarding computational complexity, Table 3 shows that PRISM-Net contains 11.33 M parameters and requires only 0.75 GFLOPs. Compared with computationally intensive 3D convolution-based models requiring more than 10 GFLOPs, the proposed method achieves superior recognition performance with significantly lower computational complexity. Furthermore, the t-SNE visualization in Fig. 5 shows more compact intra-class distributions and clearer inter-class separation, demonstrating improved feature discriminability.  Conclusions  The proposed PRISM-Net demonstrates that incorporating biomechanical priors into deep learning improves the robustness of MMW Radar gait recognition under complex wearing conditions. By combining Physics-aware Frequency Structural Reconstruction with the Weighted Attention Mechanism, the proposed framework effectively performs frequency-domain decoupling and suppresses clothing-induced high-frequency interference. Experimental results on the MMRGait-1.0 dataset demonstrate that the proposed method achieves high Rank-1 accuracy with low computational complexity, indicating its potential for real-time security applications on edge computing devices.
THz Ultra-Massive MIMO Channel Estimation via a Noise-Conditioned Fixed-Point Network
ZHANG Huawei, NIU Yaning, JIANG Zhanjun, LIU Yingting
Available online  , doi: 10.11999/JEIT260420
Abstract:
  Objective  Terahertz (THz) Ultra-Massive Multiple-Input Multiple-Output (UM-MIMO) systems are expected to support future high-capacity wireless communications. However, accurate channel estimation remains challenging under hybrid near-/far-field propagation and Array-of-SubArrays (AoSA) architectures, where limited Radio-Frequency (RF) chains, low Signal-to-Noise Ratio (SNR), noise uncertainty, and structural perturbations degrade compressed observations. Existing compressed sensing, Bayesian inference, and deep unfolding methods generally rely on fixed statistical assumptions, which limit their cross-SNR generalization under varying noise conditions and statistical mismatches. To address these limitations, this paper proposes a Noise-Conditioned Fixed-Point Network (FPN-NCAS) for robust THz UM-MIMO channel estimation. The proposed method aims to improve estimation accuracy, cross-SNR generalization, robustness, and iterative stability by incorporating noise-aware nonlinear recovery.  Methods  FPN-NCAS is developed within an Orthogonal Approximate Message Passing (OAMP)-based fixed-point unfolding framework. A coarse noise power estimate is obtained from repeated pilot differences and injected into the nonlinear recovery module as an explicit conditioning variable. After each linear update, the vector-domain estimate is reshaped into an AoSA-aligned tensor to exploit subarray-level structural priors. The nonlinear recovery chain consists of three components. Token-Gate performs lightweight subarray-level reliability pre-calibration to suppress unreliable responses under low-SNR and structurally inconsistent conditions. Block Shrink serves as the core noise-conditioned block-sparse proximal operator, in which a smooth dual-threshold mechanism balances strong denoising at low SNR with structural preservation at medium-to-high SNR. G-HMTD(Gated Hybrid Multi-scale Transformer Denoiser) further refines residual errors by combining Local Multi-scale Enhancement and Global Context Modeling. Bridge relaxation and nonlinear residual scaling are also incorporated to improve inter-stage stability.  Results and Discussions  Simulation results demonstrate that FPN-NCAS consistently outperforms LS, OAMP, ISTA-Net+, FPN-OAMP, and FPN-OTFN over the 0~20 dB SNR range (Fig. 6). At SNR = 0 dB, FPN-NCAS achieves Normalized Mean Square Error (NMSE) gains of approximately 3.0 dB and 1.9 dB over FPN-OAMP and FPN-OTFN, respectively. At SNR = 20 dB, the gains increase to approximately 4.5 dB and 2.6 dB (Fig. 6(a)). The convergence curves show that FPN-NCAS reaches a stable plateau after approximately three layers at SNR = 5 dB and five layers at SNR = 15 dB, demonstrating stable fixed-point iterative behavior (Fig. 6(b) and Fig. 6(c)). Analysis of repeated pilots shows that four repeated pilot pairs introduce only 3.13% additional pilot overhead while reducing the relative standard deviation of the coarse noise power estimate to 25.00%. Under moderate noise power mismatch, NMSE degradation remains within 0.3 dB. The proposed method also maintains strong robustness under colored Gaussian noise, impulsive noise, near-/far-field distribution shifts, variation in the number of propagation paths, AoSA subarray shuffling, and RF amplitude/phase mismatch (Fig. 7 and Fig. 8). Ablation studies indicate that Block Shrink provides the largest performance gain, whereas Token-Gate and G-HMTD further improve performance through structural calibration and residual refinement (Fig. 9).  Conclusions  This paper proposes FPN-NCAS for noise-conditioned fixed-point channel estimation in THz UM-MIMO systems. By integrating repeated-pilot-based noise conditioning, AoSA-aware feature reshaping, Token-Gate calibration, Block Shrink recovery, and G-HMTD refinement, the proposed method improves NMSE performance, robustness, and iterative stability under different SNR conditions and non-ideal scenarios. The improved performance is achieved at the cost of higher inference complexity. Future work will focus on lightweight implementations and extensions to wideband, multi-user, and hardware-impaired THz communication systems.
A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming
Available online  , doi: 10.11999/JEIT260512
Abstract:
  Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is essential for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often reduce the stability of target region segmentation, leading to centroid deviation, trajectory jitter, missed detections, and trajectory discontinuity. To address these challenges, a Spatial-Temporal Collaborative Optimization Method is proposed for stable and continuous grab trajectory extraction. While maintaining high inference speed, the proposed method improves both trajectory extraction accuracy and trajectory stability, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates Spatial Representation Enhancement (SRE) and Temporal CONSistency constraint (TCONS). First, CBAM, BiFPN-lite, and shallow feature aggregation are incorporated to improve target-background separability, enhance multi-scale feature representation, and preserve boundary details. TCONS is then imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss to suppress the temporal accumulation of local errors. In addition, a stage-wise training strategy with warm-up epochs and a joint optimization objective is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The proposed method achieves the best segmentation performance on the real portal crane grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also improves performance on DAVIS2016 while maintaining comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, and decreases the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further demonstrate that, under different occlusion ratios, the proposed method maintains good region integrity and continuous target extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementary effects of SRE and TCONS, whereas parameter analysis shows that a TCONS weight of 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  A Spatial-Temporal Collaborative Optimization Method is proposed to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves high segmentation accuracy and stable trajectory extraction on DAVIS2016 and the real portal crane grab dataset, while maintaining comparable segmentation performance on SegTrackV2. It also exhibits strong continuous target extraction capability under occlusion without significantly sacrificing inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further improve the robustness and applicability of the proposed method in real-world environments.
Conditional Generative Adversarial Network-Based Channel Estimation for RIS-Assisted ISAC System
LIU Yu, ZHENG Zelin, LIU Gang
Available online  , doi: 10.11999/JEIT251168
Abstract:
  Objective  Accurate channel estimation is essential for the reliable operation of RIS-assisted ISAC systems. Traditional deep learning methods provide partial solutions, but their generalization ability and estimation accuracy remain limited in complex multi-user channel environments. To address this issue, this study proposes a two-stage channel estimation method based on Conditional Generative Adversarial Network (CGAN) for RIS-assisted multi-user ISAC systems to improve estimation accuracy and stability.  Methods  A two-stage CGAN-based method is proposed for channel estimation in RIS-assisted multi-user ISAC systems. By adjusting the RIS switching states, the overall estimation task is divided into subproblems, which enables sequential estimation of the direct and reflected channels. Within the CGAN framework, adversarial training between the generator and discriminator is used to learn the mapping from observed signals to true channels. Feedback from the discriminator is further used to optimize the output, thereby improving training efficiency and estimation accuracy.  Results and Discussions  Extensive simulations are conducted to evaluate the effectiveness of the proposed method. Channel estimation performance is first assessed under different Signal-to-Noise Ratio (SNR) conditions. The CGAN-based approach achieves substantially better Normalized Mean Square Error (NMSE) performance than the Least Squares (LS) benchmark and conventional models such as FNN and ELM (Fig. 4). The effects of antenna number and RIS element count on channel estimation are then examined. Across different channel sizes and SNR conditions, the CGAN-based method consistently outperforms the LS benchmark (Figs. 5 and 6).  Conclusions  This study investigates channel estimation in RIS-assisted multi-user ISAC systems and proposes a two-stage CGAN-based method. By adjusting the RIS switching states and applying adversarial training between the generator and discriminator, accurate estimation of the direct and reflected channels is achieved. Simulation results show that the proposed method has strong generalization ability across different SNR levels and channel dimensions, and achieves substantially higher estimation accuracy than benchmark schemes. This method provides a promising solution for improving the accuracy and stability of channel estimation.
Blind Parameter Estimation Method for PSK Modulated Frequency-Hopping Signals Based on Improved Maximum Likelihood
ZHANG Tianhao, ZHANG Yushu, XU Zhongqiu, TANG Xinyi, DANG Wenhua, LI Guangzuo
Available online  , doi: 10.11999/JEIT260005
Abstract:
  Objective  Blind parameter estimation of non-cooperative Frequency-Hopping (FH) signals is a key task in electronic reconnaissance and countermeasure systems. Estimation methods based on time-frequency analysis typically suffer from limited resolution or high computational cost. Methods based on compressive sensing also rely heavily on consistency between the predefined dictionary and the actual signal characteristics, and their estimation accuracy is often degraded by grid mismatch or modulation-induced energy dispersion. Maximum Likelihood (ML)-based methods provide high theoretical estimation accuracy at relatively low computational cost. However, existing studies usually assume an ideal unmodulated signal model with a single frequency transition. Therfore, severe model mismatch arises when these ML-based methods are applied to digitally modulated FH signals, such as Phase Shift Keying (PSK), or to multi-hop signals. In addition, conventional iterative solutions in ML-based methods are prone to divergence or convergence to local optima. To address these issues, an improved ML-based method is proposed for blind parameter estimation of PSK-modulated FH signals.  Methods  To process received multi-hop signals, a signal-slicing method based on the Short-Time Fourier Transform (STFT) is proposed to extract slices that contain individual frequency transitions. To reduce the model mismatch caused by digital modulation in conventional ML-based methods, a model-matching signal extraction method based on the ML objective function is developed for PSK-modulated FH signals. Furthermore, a weighted iterative algorithm is designed for ML estimation to improve convergence and thus achieve robust and accurate estimation of FG parameters.  Results and Discussions  To verify the effectiveness of the model-matching signal extraction method, ablation experiments are conducted under several modulation schemes, including Binary PSK (BPSK), Quadrature PSK (QPSK), and 8-ary PSK (8PSK). The results show that the proposed method (Group D) significantly reduces the Mean Square Error (MSE) of hopping-frequency estimation compared with the method without the proposed extraction procedure (Group ND). These findings indicate that the proposed method effectively reduces model mismatch (Fig. 5). Simulation results also show that the designed weighted iterative algorithm provides better convergence than linear-weighting and non-weighting schemes (Fig. 6). The experiments further confirm that the algorithm is insensitive to initial frequency offsets, and offsets of up to 2 MHz are tolerated at an Signal-to-Noise Ratio (SNR) of –10 dB with little performance degradation (Fig. 7). Comparative experiments with representative existing methods also show that the proposed method achieves higher estimation accuracy (Fig. 8).  Conclusions  An improved ML-based method is proposed for blind parameter estimation of PSK-modulated FH signals. By using an STFT-based signal-slicing method, the applicability of the ML-based estimator is extended to continuous multi-hop signals. To reduce the model mismatch caused by PSK modulation, a model-matching signal extraction method is developed to isolate valid signal segments that satisfy the ML model. Furthermore, a weighted iterative algorithm with a dynamic weighting function is proposed to address the instability of the conventional iterative ML solver. Simulation results confirm that the proposed method effectively reduces model mismatch, provides superior convergence, and remains insensitive to initial frequency offsets. High estimation accuracy is achieved for both hopping frequency and hopping time.
FPGA Hybrid Programmable Logic Block Architecture for Highly Efficient Resource Utilization
WANG Yanlin, GAO Lijiang, YANG Haigang
Available online  , doi: 10.11999/JEIT260108
Abstract:
Six-input Look-Up Tables (6-LUTs) are widely used in commercial Field-Programmable Gate Arrays (FPGAs) to construct programmable logic blocks. However, related experiments show that their average utilization in circuits is less than 30%, which leads to substantial waste of programmable resources. In this paper, 6-LUTs are fractured according to fracturable factors and then recombined at different granularities to construct several new Hybrid Basic Logic Elements (HBLEs). Based on these HBLEs, several novel Hybrid Programmable Logic Block (HPLB) architectures are proposed. The programmable logic blocks in Xilinx devices are then replaced with these HPLB architectures. Concurrently, a statistical evaluation algorithm for the mapped netlist is proposed. Finally, several HPLB architectures are experimentally verified and evaluated. Experimental results for the three enhanced architectures show that the HPLBs achieve an average area reduction of more than 30% compared with Xilinx PLBs, without increasing the number of input ports. Among them, the hybrid HPLB architecture with a fracturable factor of N = 3 achieves the best overall optimization when both HPLB utilization and area reduction are considered. Based on the MCNC and VTR benchmarks, the proposed architecture results in average HPLB count increases of 8.27% and 27.64%, respectively, while improving programmable resource utilization.  Objective  Currently, modern commercial Field-Programmable Gate Array (FPGA) architectures use Six-input Look-Up Tables (6-LUTs) as the fundamental building blocks of basic logic elements. Experimental results show that when circuits are mapped to 6-LUT-based basic logic elements, only about 30% of the logic elements are ultimately implemented as 6-LUTs. When 6-LUTs are used to implement functions with fewer than six inputs, more than half of the logic resources are wasted. This leads to substantial underutilization of programmable resources. Experimental data show that a circuit design mapped to 100 4-LUTs can be fractured into 78 6-LUTs during 6-LUT mapping, with a {6,5,4,3,2}-LUT function distribution of {23,32,17,9,13}. These results indicate that only about 25% of the 6-LUTs are assigned to 6-input functions, whereas the remaining 6-LUTs are underutilized. This further demonstrates the inefficiency of technology mapping for LUTs with a large input size K.Methods The fracturable factor N, defined as the number of sub-LUTs that can be obtained from a single LUT, characterizes the fracturable and reconfigurable nature of LUT architectures in FPGAs. To address the low resource utilization described above, a 6-LUT is fractured into several granularities according to the fracturable factor. Three novel hybrid-granularity divisible logic structures are then constructed by reconnecting and reconfiguring the resulting sub-LUTs with additional input ports and multiplexer modules. The optimization effects of these three Hybrid Basic Logic Element (HBLE) topologies on FPGA performance are then investigated. The HBLE2 structure consists of one intact 6-LUT and one divisible 6-LUT split into two 5-LUTs with a fracturable factor of N = 2. The HBLE3 structure consists of one intact 6-LUT and one divisible 6-LUT split into one 5-LUT and two 4-LUTs with a fracturable factor of N = 3. The HBLE4 structure consists of one intact 6-LUT and one divisible 6-LUT split into four 4-LUTs with a fracturable factor of N = 4. All three HBLE structures support adder units and allow either latched output or direct combinational logic output. They also support direct latched output without passing through combinational logic. A Hybrid Programmable Logic Block (HPLB) is formed by combining several HBLEs. Two widely used academic benchmark sets, the MCNC circuit set and the VTR circuit set, are selected for experimental evaluation. Each circuit set is mapped onto a Xilinx Virtex-7 FPGA. The mapped netlist is then analyzed to count the types and numbers of LUTs used. After the data are organized with the corresponding greedy algorithms, the minimum number of Configurable Logic Blocks (CLBs) required is determined. Because each Xilinx CLB contains eight 6-LUTs, the greedy algorithm uses the total LUT number fractured by 8 to estimate the minimum number of CLBs required after benchmark mapping. To ensure comparable conditions, each structure is also reorganized with the greedy algorithm after the Xilinx CLB structure is replaced by the HPLB structure proposed in this study. This yields the minimum number of HPLBs required. In practical packing, not every LUT in the mapped CLBs can be used because of routing constraints. Therefore, the optimized result obtained after greedy restructuring represents the theoretical lower bound under ideal optimization conditions.  Results and Discussions  For the MCNC circuit set, replacing CLB structures with HPLBs reduces the average number of required blocks by about 8% for both the HPLB2 and HPLB3 structures. However, the HPLB4 structure increases the required block count by more than 30% on average. For the VTR circuit set, fewer HPLBs are required than CLBs after replacement. On average, the counts for HPLB2 and HPLB4 decrease by less than 10%, whereas the count for HPLB3 decreases by about 30%. This allows more efficient SRAM scheduling and fuller use of input pins. In contrast, the uniform CLB structure requires more CLBs when functions with a small LUT input size K are implemented because of resource waste. According to the post-mapping HPLB counts, the HPLB4 structure performs less effectively than the HPLB3 structure. Analysis of post-mapping area optimization shows that both the MCNC and VTR circuit sets achieve average area reduction ratios of more than 30%. On the MCNC benchmark set, all three HPLB structures achieve area optimization ratios of about 31%. On the VTR benchmark set, the optimization effects differ: HPLB2 achieves an average area reduction of 30.63%, whereas HPLB4 achieves an average reduction of 51.21%. HPLB3 achieves a 45.22% area reduction, which is slightly lower than that of HPLB4. Detailed analysis of the area optimization results shows that a higher fracturable factor N provides greater benefits for integrating small-scale LUTs in circuits, resulting in larger area reduction ratios in the enhanced architectures.  Conclusions  To address the low resource utilization of 6-LUTs, this study proposes three HPLB enhancement architectures based on split granularity. These HPLBs replace the Xilinx CLB structure, and an evaluation procedure and matching algorithms are established to examine the advantages of the proposed structures in resource utilization. Evaluation experiments based on the MCNC and VTR benchmark suites show that although HPLB4 achieves substantial area optimization, it also requires more HPLBs, which increases interconnect area. Both HPLB2 and HPLB3 achieve average area reductions of more than 30%. As the scale of the test circuits increases, HPLB3 provides a greater increase in HPLB count and a stronger area optimization effect than HPLB2. Therefore, after the CLB structure is replaced, HPLB3 provides a better balance between HPLB usage and area optimization, and substantially improves the utilization of programmable resources.
Deep Side-Channel Attack Method Integrating Convolutional Block Attention Mechanism and Triplet Metric Learning
XU Yang, LI Kaibin, HE Xingxing
Available online  , doi: 10.11999/JEIT260140
Abstract:
  Objective  Side-Channel Attack (SCA) is one of the primary threats to the physical security of cryptographic chips, and deep learning methods for secret key recovery have attracted considerable attention in the field of SCA. However, existing deep learning-based side-channel attack methods have limited capability to focus on critical leakage intervals during feature extraction, particularly for long traces with high-dimensional noise. Therefore, irrelevant background noise interferes with feature extraction, leading to reduced attack efficiency and slower convergence of Guessing Entropy (GE). To address these limitations, a deep side-channel attack method integrating the Convolutional Block Attention Module (CBAM) and triplet loss is proposed to improve the extraction of weak leakage features under complex noise conditions and enhance secret key recovery efficiency.  Methods  CBAM is embedded into a Convolutional Neural Network (CNN) to construct an adaptive feature extraction network. CBAM consists of a Channel Attention Module (CAM) and a Spatial Attention Module (SAM). CAM adaptively recalibrates feature-channel weights to emphasize leakage-related features with a high Signal-to-Noise Ratio (SNR), whereas SAM identifies Points Of Interest (POI) in the temporal domain and suppresses background noise outside the leakage intervals. After attention-based feature refinement, triplet loss is adopted as the optimization objective to optimize the distribution of embedding features, encouraging compact intra-class clusters while maximizing inter-class separation in the embedding space. Finally, a multivariate Gaussian template attack is performed using the optimized embedding features to recover the secret key. The overall framework is illustrated in (Fig. 2).  Results and Discussions  The proposed method is evaluated on two public benchmark datasets, ASCAD and AES_HD, using GE and the minimum number of attack traces required for GE to converge to 1 (\begin{document}$ {T}_{{\mathrm{GE0}}} $\end{document}) as evaluation metrics. On the ASCAD dataset, the proposed method requires only 144 attack traces in the ASCAD_f (HW) scenario, representing a 51.0% reduction compared with the conventional CNN model. In the ASCAD_f (ID) scenario, only 61 attack traces are required, corresponding to a 68.0% reduction. In the ASCAD_r dataset, GE converges with 176 attack traces under the HW leakage model and 137 attack traces under the ID leakage model, outperforming representative methods, including RL-SCA and Metric Learning (Table 2 and Fig. 3). On the low-SNR AES_HD dataset, the proposed method requires 1 219 attack traces, outperforming MHA and NLS while maintaining smooth and stable GE convergence (Table 2 and Fig. 4). Furthermore, desynchronization experiments demonstrate that the proposed method maintains accurate localization of effective leakage points under severe desynchronization noise, indicating strong robustness to time-domain jitter (Table 3). Ablation studies further verify the synergistic effect of the proposed architecture and confirm the effectiveness of its core components (Table 4).  Conclusions  A deep side-channel attack method integrating CBAM and triplet-loss-based metric learning is proposed. The CBAM module enables the network to adaptively focus on leakage-related features, improving feature extraction over conventional CNN-based methods. Triplet loss enhances the discriminability of embedding features, thereby improving template matching accuracy. Experimental results on the ASCAD and AES_HD datasets demonstrate that the proposed method substantially reduces the number of attack traces required for successful secret key recovery and accelerates GE convergence. The proposed method consistently outperforms existing mainstream approaches under fixed-key, random-key, and low-SNR conditions. Future work will focus on improving robustness under more severe desynchronization conditions and enhancing generalization in small-sample scenarios.
A Semantic-Enhanced Cybersecurity Named Entity Recognition Approach Oriented to Lightweight Adaptation of Large Language Models
HU Ze, XU Tongwu, YANG Hongyu
Available online  , doi: 10.11999/JEIT251260
Abstract:
  Objective  Named Entity Recognition (NER) in cybersecurity is a core technology for threat intelligence analysis, vulnerability management, and security incident response. However, this field faces several challenges, including dense technical terminology, limited labeled data, dynamic entity categories, and highly complex semantic features. These factors reduce the domain adaptability and semantic fusion capacity of traditional deep learning models and existing Large Language Models (LLMs). To address these issues while meeting the need for lightweight deployment, a cybersecurity NER approach is proposed to strengthen domain semantic representation, improve rare-entity recognition, and support low-resource environments. This approach provides a reliable technical path for intelligent threat analysis in cybersecurity scenarios.  Methods  To address the complex semantic features of cybersecurity texts, a semantically enhanced, lightweight, and LLM-adaptable cybersecurity NER approach is proposed. LLM2Vec is used to achieve bidirectional semantic reconstruction of large-model decoders, and Low-Rank Adaptation (LoRA) is combined for low-rank fine-tuning. This design preserves deep semantic encoding capacity while substantially reducing the number of updated parameters. To address sparse keywords and severe noise interference in cybersecurity texts, a sparse gated attention mechanism is proposed to strengthen keyword-focused feature extraction. High-contribution cybersecurity terms are selected dynamically through global gating and sparse inference. A SecRoBERTa-based semantic enhancement module is also proposed. This module uses a domain-pretrained model to generate similar-word embeddings, improves feature robustness in small-sample settings, and reduces the difficulty of identifying out-of-vocabulary words and low-frequency terms. Finally, a Masked Conditional Random Field (MCRF) is used to constrain label transitions and ensure BIO-compliant output sequences, thereby achieving robust and consistent entity boundary prediction.  Results and Discussions  Extensive experiments are conducted on two public cybersecurity datasets, DNRTI and APTNER. The proposed approach achieves an F1 score of 91.91% on DNRTI, exceeding the previous state-of-the-art model by 2.14%. On APTNER, it achieves an F1 score of 80.37%, exceeding the best baseline by 2.97%. Ablation studies confirm the contribution of each key component. The sparse gated attention mechanism improves F1 by 3.57% over standard multi-head attention on DNRTI. The semantic enhancement module contributes a 2.32% increase in F1. The model also shows efficient training and inference, consistent with the goals of lightweight design.  Conclusions  A lightweight LLM-based adaptation approach is proposed for NER in the cybersecurity domain. The approach effectively addresses the limitations of existing LLM-based NER methods in domain adaptation and rare-entity recognition. By integrating LLM2Vec and LoRA for lightweight fine-tuning, a sparse gated attention mechanism for domain feature fusion, and a SecRoBERTa-based semantic enhancement module for similar-word precomputation, the proposed approach achieves strong performance on the DNRTI and APTNER datasets. This research provides an efficient technical path for NER tasks in low-resource cybersecurity scenarios and supports downstream tasks such as automated threat intelligence analysis.
Advances and Challenges in Intelligent Damage Assessment of Objects in Remote Sensing Imagery
ZHANG Yidan, FENG Yingchao, WANG Tianqi, LIU Yu, WANG Mengyu, HOU Zhongyan
Available online  , doi: 10.11999/JEIT251297
Abstract:
  Significance   Rapid and accurate damage assessment of high-value objects following disasters is essential for effective emergency response and post-disaster recovery. Deep learning-enabled remote sensing provides a scalable, objective, and efficient approach for assessing disaster damage over large and complex environments, including densely populated urban areas, transportation hubs, and critical infrastructure. By exploiting high-resolution satellite and aerial imagery, these methods provide timely situational awareness to support rescue prioritization and recovery planning. Despite substantial advances in algorithms and applications, the field still lacks a comprehensive review, leading to fragmented technical development and inconsistent evaluation practices. This paper systematically reviews the technical foundations of intelligent damage assessment in remote sensing, including damage classification standards, publicly available datasets, evaluation metrics, and representative deep learning methods. The review aims to facilitate the practical deployment of intelligent remote sensing technologies for disaster response under increasing natural and human-induced hazards.  Progress   Deep learning-based damage assessment methods for remote sensing imagery have advanced rapidly, with substantial improvements in assessment accuracy, automation, and scalability. Representative developments include Bi-temporal Change Detection methods, which identify damage by comparing pre-disaster and post-disaster imagery, and Multi-temporal Sequence Modeling methods, which characterize the temporal evolution of damage using image sequences. Multi-modal Data Fusion methods that integrate optical imagery, Synthetic Aperture Radar (SAR), and Light Detection And Ranging (LiDAR) data further improve damage assessment under complex imaging conditions by exploiting complementary information from multiple sensors. In addition, methods designed for data-constrained scenarios, including transfer learning, semi-supervised learning, self-supervised learning, and domain adaptation, improve model robustness and generalization when labeled data are limited. These advances substantially improve the efficiency, reliability, and applicability of intelligent damage assessment systems for emergency response and resource allocation.  Conclusions  This paper systematically summarizes the technical landscape of deep learning-based damage assessment of high-value objects in remote sensing imagery. Existing methods are categorized into four major groups: Bi-temporal Change Detection, Multi-temporal Sequence Modeling, Multi-modal Data Fusion, and methods for Data-Constrained Scenarios. Their technical characteristics, strengths, and limitations are systematically analyzed and compared. Although these methods have demonstrated promising performance in post-disaster damage assessment, several challenges remain, including limited robustness across diverse environments, insufficient exploitation of temporal and multimodal information, and inadequate generalization under limited training data. In addition, unified damage classification standards and comprehensive evaluation frameworks remain unavailable, limiting the consistency, comparability, and practical applicability of current assessment systems.  Prospects   Future research should focus on developing hierarchical collaborative frameworks for damage assessment across multiple object types, spatial scales, and functional levels to characterize both direct physical damage and cascading functional degradation. Complex environments, including airports, industrial facilities, and ports, contain static infrastructure, moving objects, and highly interconnected functional units, requiring hierarchical scene understanding and object-level reasoning. Physics-informed and hybrid learning frameworks that integrate structural mechanics, material degradation mechanisms, and domain knowledge are expected to improve model interpretability and generalization. Furthermore, lightweight model architectures and edge deployment strategies will be essential for real-time damage assessment on unmanned aerial vehicles and satellite platforms. Standardized evaluation systems that jointly consider physical damage and functional degradation will further facilitate practical deployment in emergency response and post-disaster recovery.
Hierarchical Prototype Learning with Shared Subspace Factorization for Generalizable Deepfake Detection
PENG Shufan, LU Tianliang, HE Chunhao, ZHANG Lu, ZHAO Kai
Available online  , doi: 10.11999/JEIT260426
Abstract:
  Objective  Deepfake detectors often exhibit performance degradation when applied to unseen manipulation methods, cross-dataset distribution shifts, diffusion-generated faces, and common image degradations. In forensic applications, the generation process, data source, and post-processing conditions of a questioned sample are usually unknown. Existing methods often represent the fake class with a single feature center, although different generation methods, data sources, and processing conditions produce heterogeneous patterns. Transferable forensic cues may therefore be mixed with mode-specific artifacts, limiting generalization to unknown domains. To address this problem, a hierarchical prototype learning framework with shared subspace factorization (HPL-SF) is proposed. Within-class diversity is first used to observe latent fake modes, followed by estimation of a low-rank structure shared across these modes and sample-level exploitation of the shared structure during inference.  Methods  A pretrained DINOv2 Vision Transformer (ViT-L/14) is used as the backbone. Its original parameters are frozen, and Low-Rank Adaptation (LoRA) modules are inserted into the query and value mappings of the self-attention layers for parameter-efficient training. All features and prototypes are L2-normalized, and cosine similarity is used for prototype assignment, binary classification, and test-time representation adaptation. HPL-SF consists of three successive stages (Fig. 1). First, one real prototype and multiple mode-specific fake prototypes are maintained in the normalized feature space. Each training sample is softly assigned to the fake prototypes according to its cosine similarity to each prototype, and the weighted prototypes are aggregated to obtain a sample-adaptive fake representation. The response distribution thus provides an observable representation of latent within-class modes without using forgery-source labels. A binary classification loss, a sample–prototype contrastive loss, and a prototype diversity loss are jointly optimized to separate real and fake samples, improve sample–prototype alignment, and prevent prototype collapse. Second, the normalized mode-specific fake prototypes are arranged into a prototype matrix. Singular Value Decomposition (SVD) is applied to this matrix, and the largest gap between consecutive singular values is used to determine the dimension of the shared fake subspace. Each fake prototype is then decomposed into a projection onto the shared fake subspace and a mode-specific residual. Only the shared projections are aggregated to update the shared fake prototype, whereas the residuals retain mode-specific information and are excluded from this update. The real prototype and shared fake prototype are updated using Exponential Moving Average (EMA). Gradients are not propagated through subspace construction, dimension selection, or semantic prototype updates. Third, the backbone, LoRA modules, prototypes, and subspace basis are fixed during inference. A test feature is compared with the real and shared fake prototypes, and their relative responses determine a sample-specific adaptation weight. The original feature is blended with its projection onto the shared fake subspace and then L2-normalized. The resulting feature is classified according to its cosine similarities to the two semantic prototypes. Thus, the estimated shared structure is exploited on a sample-by-sample basis without updating model parameters during testing.  Results and Discussions  The method is evaluated under cross-dataset, cross-forgery-type, diffusion-forgery, repeated-run, image-degradation, ablation, and mechanism-analysis protocols. Across seven unseen datasets, HPL-SF achieves the highest area under the receiver operating characteristic curve (AUC) on all seven datasets, with an average AUC of 91.67%, which is 2.10 percentage points higher than the second-highest average (Table 1). When trained on FaceForensics++ and evaluated on the Diffusion Facial Forgery dataset, HPL-SF achieves the highest AUC on the text-to-image, image-to-image, face-swapping, and face-editing subsets, with an average AUC of 84.39% (Table 2). Across four cross-forgery-type settings on FaceForensics++, the average accuracy and AUC reach 85.29% and 92.41%, respectively, both ranking first among the compared methods. When DeepFakes under high-quality compression is held out for testing, HPL-SF trails the best-performing method by only 0.91 percentage points in accuracy and 0.39 percentage points in AUC (Table 3). Repeated experiments with multiple random seeds yield the highest mean values for all four aggregate metrics, with standard deviations ranging from 0.38 to 0.62 percentage points (Table 4). Under five severity levels of compression, blur, and noise on the Deepfake Detection Challenge Preview dataset, HPL-SF achieves the best or joint-best AUC at most severity levels and remains relatively stable under moderate and severe degradations (Fig. 2). All six ablated variants perform worse than the complete HPL-SF model on the four unseen test sets, indicating complementary contributions from mode-specific prototype learning, sample–prototype contrastive loss, prototype diversity loss, shared subspace factorization, and test-time representation adaptation (Fig. 3). Performance generally increases as the number of mode-specific fake prototypes increases and plateaus when the number reaches 10. The largest spectral gap occurs between the fifth and sixth singular values, yielding a shared subspace dimension of 5. With this setting, HPL-SF achieves a diffusion-forgery AUC of 84.39% ± 0.62%, compared with 79.68% for direct mean aggregation and 81.56% ± 0.80% without test-time representation adaptation (Figs. 4(a)–4(c)). Linear-probe analysis further shows that the shared component provides stronger discrimination in unknown domains and weaker domain-identifying capability than the mode-specific residual, supporting the separation of transferable and domain-related cues (Fig. 4(d)). The t-distributed stochastic neighbor embedding (t-SNE) visualization shows clearer real–fake separation and closer distributions of same-class samples from different data sources (Fig. 5). Prototype allocation statistics show differentiated responses across fake types, with prototype usage remaining close to the uniform baseline. The adaptation weights are higher for fake samples, particularly for correctly classified fake samples, whereas misclassified samples show responses closer to the balanced-response line (Fig. 6). False positives are mainly associated with low resolution, compression, filters, or occlusion, whereas false negatives usually contain high-quality forgeries or weak forgery traces (Fig. 7).  Conclusions  HPL-SF organizes generalizable deepfake detection into three successive stages: observation of within-class diversity, estimation of shared structure, and sample-level exploitation of that structure. Experimental results show that separating shared projections from mode-specific residuals provides transferable discriminative cues under dataset shifts, unseen forgery types, diffusion-generated manipulations, and common image degradations. The framework requires no parameter updates during testing and adaptively exploits the shared structure according to each sample’s relative responses to the real and shared fake prototypes. Errors remain for degraded real images whose artifacts resemble forgery traces and for high-quality forgeries with weak traces. Future work will extend the framework to temporal prototypes for video, multimodal forensic evidence, and adaptive discrimination in broader open-world settings.
Dual-Domain Differentiated Feature Extraction Network for MRI Reconstruction
XUE Nan, QIAO Han, WANG Peng
Available online  , doi: 10.11999/JEIT251093
Abstract:
  Objective  In magnetic resonance imaging (MRI), undersampling of k-space data is an effective approach to accelerate image acquisition. Reconstructing high-quality MR images from undersampled data is of great clinical significance for diagnostic efficiency. Currently, dual-domain reconstruction methods that jointly exploit spatial- and frequency-domain features have become mainstream. However, existing dual-domain MRI reconstruction methods fail to design differentiated feature extraction strategies for the two domains, and their physical prior constraints are insufficient, leading to loss of original information during the reconstruction process.  Methods  To address these issues, this study proposes a Dual-Domain Differentiated Feature Extraction Network (DDF-Net) for MRI reconstruction. In the spatial domain, an interlaced row-column self-attention mechanism combined with depthwise convolution is designed to accurately capture anisotropic structures and texture details, thereby addressing the limitations of conventional convolutional feature extraction. In the frequency domain, amplitude and phase characteristics are independently modeled to capture intensity and structural positional information, respectively, and a frequency-domain feature enhancement module is introduced to fully exploit spectral information for improving reconstruction fidelity. Finally, a data consistency layer and a cross-domain adjustment module are incorporated to integrate physical priors and cross-domain information, thereby reinforcing measurement consistency and stabilizing the reconstruction process.  Results and Discussions  The proposed DDF-Net was evaluated on two publicly available MRI datasets, CC359 and IXI, and compared with six representative reconstruction algorithms, including DAGAN, KIKI-Net, MD-Recon-Net, SwinMR, Reconmer, and KTMR. As shown in Fig. 6 and Fig. 7, under a Gaussian 1D 30% undersampling mask, DDF-Net achieves the most faithful anatomical restoration with clearer cortical edges and finer texture details, while suppressing aliasing artifacts effectively. The error maps exhibit more uniform residual distributions, indicating better consistency between the reconstructed and reference images. Quantitative comparisons (Table 1 and Table 2) show that DDF-Net attains the highest PSNR and SSIM values across all sampling patterns. Specifically, it improves PSNR by 0.17 dB and 0.23 dB, and SSIM by 0.0063 and 0.0019 on the CC359 and IXI datasets, respectively, achieving the best overall performance over the second-best competing method. These results demonstrate that the differentiated spatial-frequency feature extraction enables DDF-Net to leverage complementary information between domains for more precise recovery. To further verify robustness, a noise experiment was conducted by adding Gaussian noise of varying intensity levels to the k-space data, following the procedure in SwinMR. As illustrated in Fig. 8 and summarized in Table 3, DDF-Net exhibits stronger robustness to noise perturbations, maintaining higher PSNR and SSIM values than SwinMR under all noise conditions. Even at NL = 80%, DDF-Net effectively preserves most structural details, confirming that the amplitude-phase separation and multi-level data consistency jointly improve noise resilience and stability.  Conclusions  This study proposes an end-to-end Dual-Domain Differentiated Feature Extraction Network to address the limitations of existing dual-domain MRI reconstruction methods, which often lack domain-specific feature extraction strategies and sufficient physical prior constraints, leading to potential information loss during reconstruction. In the spatial domain, an interlaced row-column self-attention mechanism is designed to more effectively model local structures and texture details. In the frequency domain, amplitude-phase separation and a frequency-domain feature enhancement module are introduced to effectively improve the decoupling and utilization of spectral information. Furthermore, by integrating a cross-domain adjustment module with multiple data consistency layers, DDF-Net enhances cross-domain interaction and measurement fidelity. Experimental results on the CC359 and IXI datasets demonstrate that the proposed method achieves superior performance in both quantitative metrics and visual reconstruction quality, successfully recovering fine anatomical details. In addition, noise experiments confirm the robustness of DDF-Net under complex conditions, showing that the network can effectively preserve image details across various noise levels. Future work will focus on extending DDF-Net to unsupervised or semi-supervised learning frameworks to reduce dependence on labeled data and further enhance its potential for clinical applications in fast MRI reconstruction.
Multi-RAT Fusion Architecture and Intelligent Routing for Marine Heterogeneous Wireless Networks
CHEN Jin, ZHOU Xuan, LIN Haitao, YU Huagang, LI Yun
Available online  , doi: 10.11999/JEIT260482
Abstract:
  Objective  Marine Heterogeneous Wireless Networks (MHWNs), which deeply integrate multi-dimensional resources spanning space, air, and sea, must accommodate multiple coexisting communication systems and thus face the dual challenges of interconnection and resource coordination. Meanwhile, existing routing algorithms based on Deep Reinforcement Learning (DRL) exhibit insufficient representation capability for dynamic topologies, making it difficult to make efficient routing decisions when the network topology changes frequently. This is mainly because mainstream frameworks typically adopt standard Graph Neural Networks (GNNs) or fully connected networks for state encoding, which cannot effectively capture the structural features of highly dynamic topologies.  Methods  This paper designs a modular Multi-Radio Access Technology (Multi-RAT) gateway supporting the fusion of 5G and ad hoc (Mesh) networks, and proposes an intelligent routing method combining a Contrastive Message Passing Graph Neural Network with Deep Reinforcement Learning (CMPGNN-DRL), so as to improve Quality of Service (QoS) and forwarding efficiency and achieve optimized resource allocation. In the gateway, service data are IP-encapsulated, protocol-identified, and semantically converted by communication interface modules before being forwarded to the target interface. The gateway periodically collects node features and link states to construct the input graph for routing decisions, and monitors key indicators such as link bandwidth utilization, queue depth, packet loss rate, and end-to-end delay. To compensate for the information delay introduced by periodic reporting, short-term trend terms of key indicators are incorporated into the state vector, and an asynchronous decision-execution architecture is adopted. Under a centralized Software-Defined Networking (SDN) control plane, the network is modeled as a graph with continuously monitored node and link features. For each node, the CMPGNN module synchronously constructs homophily and heterophily views along two message-passing paths and constrains the resulting embeddings with a contrastive loss, yielding discriminative node representations that are robust to edge perturbations. The learned representations are then fed into a Double Deep Q-Network (DDQN) agent that makes hop-by-hop routing decisions with an ε-greedy exploration strategy. Specifically, the topology prior values predicted by CMPGNN for neighboring nodes are fused with the DDQN Q-value estimates through weighted summation to form joint action values, which can correct inaccurate Q-value estimates when training is insufficient or observations are noisy. A normalized multi-objective reward function is designed to be negatively correlated with latency, packet loss rate, and link load, and positively correlated with throughput, while explicitly penalizing routing loops.  Results and Discussions  The proposed solution was validated through hardware prototype measurements and extensive simulations. Prototype tests showed that the average CPU utilization of the fusion gateway was 14%, 25%, and 37% in Mesh-only, 5G-only, and dual-mode operation, respectively; the dual-mode aggregate throughput reached 108 Mbps, compared with 32 Mbps for Mesh-only and 84 Mbps for 5G-only (uplink); and the average ping latencies between the gateway and the application server were 6 ms for Mesh and 16 ms for 5G. CMPGNN-DRL was compared with six baseline methods, namely OSPF, AODV, GNN, DQN, MPNN-DQN, and DAR-DRL, on the GEANT2, GBN, Germany, and Synth50 topologies, covering dynamic traffic, random link failures with failure rates of 3%–24%, and large-scale topology variations. The training reward increased rapidly and then stabilized, and ablation experiments verified the effectiveness of the contrastive learning mechanism. Compared with the optimal baseline, the proposed method reduces the average end-to-end delay by 20.8%–47.7%, reduces the packet loss rate by 0.3%–5.3%, and increases the average throughput by 5.2%–14.2%. In maritime heterogeneous wireless network scenarios constructed according to the environmental constraints of the Maritime Internet of Things (MIoT), i.e., a 1500 m × 1500 m area with 50–100 randomly deployed nodes evaluated through repeated Monte Carlo simulations, the method adapted stably to variations in network scale and node mobility in terms of Packet Delivery Ratio (PDR) and packet loss rate. As the load rate increased from 20% to 50%, it improved PDR by 4.2%–17.3% over MPNN-DQN and DAR-DRL while maintaining lower latency, higher bandwidth utilization, and a lower packet retransmission ratio under medium-to-high loads.  Conclusions  Aiming at sea-air cross-domain heterogeneous networks, this paper designed a Multi-RAT fusion gateway supporting ad hoc networks and 5G, and proposed an intelligent multi-path routing algorithm integrating CMPGNN with DRL. The contrastive learning mechanism strengthens the topology representation capability of the graph neural network and improves the robustness of routing policies. Experimental results show that the proposed method outperforms existing mainstream algorithms in key performance indicators such as PDR, end-to-end delay, and throughput, and exhibits good cross-topology generalization capability. Future work will focus on verification in real maritime environments and optimization of training efficiency, so as to support the practical deployment and application of integrated sea-air communication systems.
Cooperative Search and Tracking of Moving Ships Using Constellation Multi-Functional Payloads Based on Dynamic Information Gain
ZHANG Yumo, ZHAO Fuhai, LI Xiaobin, FAN Shenghua, QU Tao
Available online  , doi: 10.11999/JEIT260500
Abstract:
  Objective  Wide-area maritime surveillance requires satellite constellations to search for and revisit non-cooperative maneuvering ships whose positions become uncertain after missed observations. Meanwhile, multi-functional payloads are subject to coupled constraints on observation timing, attitude maneuvering, payload mode, and energy consumption. To address dynamic target uncertainty and executable constellation scheduling, a cooperative search-and-tracking method based on dynamic information gain is proposed.  Methods  A closed-loop rolling-horizon framework inspired by Model Predictive Control is constructed to perform prediction, optimization, execution, and feedback. At each decision epoch, candidate atomic tasks are generated over a planning horizon, while only tasks within the current execution window are committed. Each task specifies the executing satellite, target, candidate pointing grid, start/end times, and payload mode. Target uncertainty is represented by a parameterized probabilistic grid derived from the latest confirmed state, speed and heading perturbations, and elapsed time since the last successful observation. Hit/Miss feedback updates the uncertainty baseline for the next rolling step, where the probability grid, candidate tasks, and observation plan are regenerated. A state-driven dual-mode benefit model is established according to target information entropy and consecutive successful observations. In the robust tracking state, narrow-field tasks are evaluated by the prior capture probability, namely the probability mass covered by the task footprint. In the lost-search state, wide-field tasks are evaluated by the binary entropy of Hit/Miss events as an approximation of search information value. This approximation is motivated by Kullback-Leibler divergence and avoids explicit posterior reconstruction for every candidate task. A dynamic priority coefficient increases scheduling urgency for long-unobserved targets and moderately down-weights repeatedly confirmed targets. The resulting multi-objective model maximizes weighted task benefit and information gain while minimizing energy consumption, subject to hard constraints on single-satellite temporal exclusivity, attitude-transition stabilization time, and available energy. Based on NSGA-II, the Cooperative Evolutionary Planning-Multi-Objective (CEP-MO) algorithm employs global integer-index encoding, constraint-aware Top-K heuristic initialization, satellite-group crossover, and adaptive repair to improve feasible-solution generation. Feasible Pareto solutions are normalized, and the solution closest to the ideal point (1,1,0) is selected for execution.  Results and Discussions  Simulations with a 48-satellite Walker constellation demonstrate the effectiveness of the proposed method. In the 200-target scenario, Standard NSGA-II obtains an average revisit interval of 53.5 min and a weighted coverage of 13.9%, whereas CEP-MO achieves 19.9 min and 41.7%, respectively, reducing the average revisit interval by 62.8% (Fig. 7). Removing the binary-event-entropy benefit increases system-average uncertainty and revisit interval, while replacing constraint-aware initialization with random initialization degrades early convergence and weighted coverage. As the target number increases from 100 to 200, CEP-MO maintains acceptable scalability (Fig. 8). At 200 targets, its weighted coverage is 16.6 percentage points higher than that of CEP-MO w/o Entropy, the computation time per rolling decision is approximately 22 s, and the Gini coefficient of remaining satellite energy stays below 0.3, indicating that energy consumption is not excessively concentrated on a small subset of satellites.  Conclusions  The proposed framework integrates probabilistic-grid uncertainty representation, state-driven search/tracking benefit evaluation, rolling feedback, and constraint-aware multi-objective evolutionary planning. CEP-MO improves revisit and weighted-coverage performance while maintaining temporal, attitude, and energy feasibility. The method provides an effective approach for large-scale resource-constrained maritime surveillance and a basis for future extensions involving identification errors, communication delays, and constrained inter-satellite links.
Preamble-Referenced Cyclic Cross-Correlation Chirp Spread Spectrum Communication Technology in Complex Multipath Environments
YE Yun, ZHANG Chengyu, PANG Haodong, MA Wenfeng, LI Xuejiao, ZHANG Xiaokai
Available online  , doi: 10.11999/JEIT260702
Abstract:
  Objective  Ground unmanned platforms operating in urban streets, industrial parks, and underground passages require short-burst reliable command-and-control communication. These complex near-ground environments simultaneously impose strong multipath fading, large Carrier Frequency Offset (CFO), residual Timing Offset (TO), Sampling Frequency Offset (SFO), and in-band interference from coexisting wireless systems. Conventional Chirp Spread Spectrum (CSS) receivers based on single-peak decisions in the dechirp–Discrete Fourier Transform (DFT) domain suffer from multipath-induced spectral splitting and interference-induced bin masking, while Direct-Sequence Spread Spectrum (DSSS) baselines exhibit synchronization fragility under combined offsets and degraded energy efficiency under multipath. Simultaneously addressing these impairments is essential for enabling robust low-latency control links and for the coexistence of unmanned platforms with legacy wireless infrastructure in dense deployments.  Methods  This paper proposes a Preamble-Referenced Cyclic Cross-Correlation CSS (PRCC-CSS) scheme that jointly designs frame structure, synchronization estimation, and payload detection. Each frame comprises multiple identical up-chirp preamble symbols, a down-chirp Start Frame Delimiter (SFD), and CSS-modulated payload symbols. The complementary frequency-domain indices at the dechirp-DFT outputs of the up-chirp preamble and the down-chirp SFD are exploited to jointly estimate integer CFO and TO via closed-form linear combinations. Fractional CFO is recovered from inter-symbol phase differences across adjacent preamble symbols; fractional TO is extracted from the centroid of the main-peak neighborhood of the averaged preamble spectrum; and under a common oscillator-reference assumption, SFO-induced bin drift is compensated using the CFO-derived clock-offset relation. After offset compensation, the averaged preamble spectrum serves as a frame-specific reference spectrum that captures the instantaneous multipath fingerprint of the channel. Payload symbols are then detected by computing the cyclic cross-correlation between this reference spectrum and each candidate-shifted payload spectrum, taking the maximum-correlation index as the demodulated symbol. This formulation converts the multipath-induced frequency-domain structure from an adverse perturbation into a matchable intra-frame reference feature, thereby enabling multipath-robust structure-matched payload detection without explicit path-by-path channel estimation.  Results and Discussions  PRCC-CSS is evaluated under identical bandwidth, sampling rate, and processing gain at a target Bit Error Rate (BER) of 10–4. First, it is compared against three DSSS baselines using Binary Phase-Shift Keying (BPSK), Quadrature Phase-Shift Keying (QPSK), and 16-ary Quadrature Amplitude Modulation (16QAM) under three channel conditions. Under Additive White Gaussian Noise (AWGN, Fig. 1), PRCC-CSS reaches the target at approximately 5 dB Eb/N0 versus 8.5–9 dB for the best DSSS baseline, indicating that chirp index modulation combined with preamble-referenced correlation provides inherent frequency-domain energy aggregation independent of any specific multipath profile. Under the Extended Typical Urban (ETU) channel without interference (Fig. 2), PRCC-CSS requires approximately 6.5 dB versus approximately 10 dB, as the preamble reference spectrum captures and reuses the per-frame multipath structure that finite-finger DSSS-RAKE cannot fully exploit due to path-capture and code-synchronization errors. Under ETU with in-band interference at Jamming-to-Signal Ratio (JSR) = 5 dB (Fig. 3), PRCC-CSS requires approximately 8.1 dB versus 10.8–11.0 dB, since the cyclic cross-correlation preserves decision separability—reference and payload spectra share nearly identical channel structure within the same frame—whereas DSSS-RAKE accumulates interference residue across all combining fingers. Second, against Dechirp Non-Coherent (DNC) and coherent peak detection at Spreading Factor (SF) = 7 and SF = 9 under ETU with interference (Fig. 4), DNC fails to reach the target within the tested Signal-to-Interference-plus-Noise Ratio (SINR) range at SF = 7 and requires approximately 4–5 dB at SF = 9, whereas PRCC-CSS reaches the target at approximately –5 dB SINR at SF = 7 and –11 dB at SF = 9, yielding a 10–15 dB SINR threshold improvement and confirming that the gain stems from frame-wide cyclic matching rather than phase compensation alone. Third, a Software-Defined Radio (SDR) prototype on an ETU-emulated channel at JSR = 5 dB (Fig. 5Fig. 6) retains 300 of 332 received frames as reliable (90.4%). In the retained reliable frames, no symbol errors were observed among 6,000 payload symbols and no bit errors were observed among 54,000 payload bits, corresponding to a one-sided 95% upper confidence bound of 5.56 × 10-5 on the retained-frame conditional BER.  Conclusions  This paper proposes the PRCC-CSS scheme that jointly integrates integer/fractional CFO–TO and SFO estimation with cyclic cross-correlation payload detection. Results demonstrate that: (1) at BER = 10–4, PRCC-CSS lowers the required Eb/N0 by approximately 2.7–4.0 dB relative to the best DSSS-BPSK/QPSK/16QAM baseline across AWGN, ETU, and ETU with JSR = 5 dB cases; (2) under ETU with interference, PRCC-CSS lowers the required SINR by approximately 10–15 dB relative to DNC at SF = 7 and SF = 9, with a further consistent margin over coherent peak detection; (3) the SDR prototype retains 90.4% of received frames, and the retained-frame conditional BER has a one-sided 95% upper confidence bound of 5.56 × 10–5. By exploiting the multipath-induced frequency-domain structure as a matchable intra-frame reference feature rather than as a perturbation, PRCC-CSS provides a candidate physical-layer solution for short-burst reliable communication in complex near-ground environments. Future work will extend the scheme to higher-order CSS modulations, multi-antenna diversity reception, and adaptive reference-spectrum updating for time-varying channels.
Recent Advances in Remote Sensing Image-Text Retrieval Driven by Vision-Language Foundation Models
WU Hui, ZHAO Yan, ZHANG Peirong, HOU Yingyan, QI Xiyu, WANG Lei
Available online  , doi: 10.11999/JEIT260189
Abstract:
  Significance  Remote Sensing Image-Text Retrieval (RS-TIR) connects large-scale Earth observation imagery with natural-language queries and has become an important interface for geospatial intelligence systems. Compared with conventional content-based retrieval, RS-TIR allows users to search for scenes, objects, spatial layouts, and functional regions through semantic descriptions rather than handcrafted visual cues. This capability is increasingly needed in natural resource monitoring, urban governance, disaster response, environmental assessment, and on-demand retrieval from rapidly growing satellite archives. However, RS-TIR remains challenging. Remote sensing imagery is captured from nadir or near-nadir perspectives, shows strong rotation invariance, and contains extreme scale variation, ranging from tiny vehicles to large airports. It also requires domain-specific semantic descriptions, such as land-use attributes, spatial distributions, and geoscientific relations. Meanwhile, high-quality image-text annotations remain limited relative to the scale of remote sensing data. These properties widen the cross-modal semantic gap between images and language and limit the generalization ability of traditional cross-modal retrieval methods. Against this background, this review examines how Vision-Language Foundation Models (VLMs) reshape RS-ITR through large-scale contrastive pre-training, stronger transferable representations, and more flexible multimodal interaction mechanisms. It also explains why remote sensing adaptation is needed and why a focused synthesis of architectures, datasets, alignment mechanisms, and future directions is timely for this field.  Progress   The technical development of RS-ITR is reviewed from three complementary perspectives. First, this review summarizes the domain-specific challenges that shape the task, including visually isotropic topology with extreme scale variation, professional and fine-grained textual semantics, and the compounded cross-modal semantic gap between overhead imagery and natural-language descriptions (Fig. 3). The overall survey structure is then presented to show the logical progression from task formulation to future challenges (Fig. 1). From a methodological perspective, RS-ITR has evolved from handcrafted visual descriptors and shallow semantic mapping to deep representation learning, and then to VLM-driven paradigms with stronger generalization and zero-shot transfer capability (Fig. 4, Table 2). Early methods rely on color, texture, shape, and hash-based retrieval. However, they struggle to model high-level geospatial semantics and complex scene composition. Deep learning methods improve retrieval by learning joint embedding spaces, adopting dual-encoder or interaction-based architectures, and using multi-scale feature fusion and region-aware matching. These methods improve semantic consistency, but they still depend heavily on labeled data and often show limited robustness in open or cross-sensor scenarios. Second, this review summarizes the benchmark ecosystem used to evaluate these methods. Representative datasets range from small-scale test sets, such as Sydney-Caption and UCM-Caption, to mainstream benchmarks, such as RSICD and RSITMD, and recent large-scale training resources, such as RS5M and SkyScript (Table 1). These datasets show a clear transition from small manually annotated corpora to web-scale or automatically generated image-text pairs. This transition supports domain pre-training and large model adaptation. Third, this review analyzes the core VLM techniques that now drive progress in RS-ITR. The model spectrum and representative architecture families are systematically summarized, including contrastive dual-encoder models, multimodal interaction models, and remote sensing foundation models integrated with large language models (Fig. 5, Fig. 6, Table 3). Domain adaptation routes are further grouped into continued remote sensing pre-training, parameter-efficient transfer learning, adapter-based tuning, prompt learning, and instruction tuning. At the semantic alignment level, this review focuses on contrastive joint embedding, fine-grained multi-scale alignment, and the use of remote sensing priors, such as spatial topology and geolocation. Performance comparisons on RSICD and RSITMD show that remote sensing VLMs, especially RemoteCLIP, GeoRSCLIP, iEBAKER, and LRSCLIP, yield consistent gains in mean Recall (mR) and overall retrieval robustness (Table 4). In parallel, this review tracks the extension of retrieval capability into unified multi-task remote sensing models, in which retrieval, grounding, segmentation, and reasoning begin to share a common multimodal representation space.  Conclusions  Several conclusions are drawn from the comparative analysis. First, VLMs establish a dominant paradigm for RS-ITR because they narrow the cross-modal semantic gap and improve transferability across datasets and scenes. Second, no single architecture is universally optimal. Dual-encoder models remain attractive for large-scale retrieval because of their efficiency, whereas interaction-based or instruction-enhanced models provide finer semantic alignment at a higher computational cost. Third, domain adaptation is indispensable. Continued pre-training on remote sensing image-text corpora, parameter-efficient tuning, and prompt-based adaptation consistently outperform direct reuse of internet-trained VLMs. This finding indicates that remote sensing imagery differs too strongly from natural-image distributions for generic pre-training alone to be sufficient. Fourth, the most effective recent methods do not improve performance through scale alone. They also exploit remote sensing-specific information, including multi-scale structures, foreground objects, explicit keyword reasoning, and spatial priors. Finally, this review shows that the field is shifting from isolated retrieval models toward more general geospatial multimodal systems. Retrieval is no longer treated only as a matching task. It is also becoming a key capability that supports question answering, instruction following, knowledge augmentation, and coordinated reasoning in remote sensing applications.  Prospects   Future research is expected to advance in four closely related directions. The first direction is the unified representation of multi-source heterogeneous data, especially the integration of optical imagery with Synthetic Aperture Radar (SAR), hyperspectral data, thermal infrared observations, and multi-temporal acquisitions. The second direction is knowledge-enhanced retrieval, in which geospatial priors, land-use rules, remote sensing terminology, and external knowledge bases are incorporated into multimodal alignment and retrieval-augmented reasoning. The third direction is lifelong and open-world learning. Real deployment requires models to remain reliable under seasonal variation, sensor updates, regional domain shifts, cloud contamination, and newly emerging categories, while avoiding catastrophic forgetting. The fourth direction is efficiency and deployability. Practical remote sensing systems often operate under tight computational budgets. Therefore, lightweight tuning, sparse computation, token reduction, model compression, and on-orbit and edge inference will become increasingly important. Interactive and explainable retrieval is also likely to gain importance. It allows analysts to refine queries through dialogue and inspect the image regions or semantic cues that support retrieval decisions. Overall, continued progress in data construction, domain adaptation, semantic alignment, and efficient multimodal modeling is expected to make RS-ITR a more robust infrastructure capability for Earth observation applications.
Energy Efficiency Analysis of Discrete Phase-shifted Active RIS Enhanced Communication Systems
SHU Feng, LIN Zhiyuan, ZHENG Weihai, WANG Yan, JIANG Hao, WANG Jiangzhou
Available online  , doi: 10.11999/JEIT260462
Abstract:
  Objective  Active Reconfigurable Intelligent Surface (RIS) enhances wireless communication performance by integrating radio frequency amplifiers to mitigate the multiplicative fading inherent to passive RIS. However, amplification noise and additional power consumption are introduced. Furthermore, high-precision digital phase control at the base station incurs considerable communication overhead. Employing low-precision phase shifters is therefore an effective approach for practical RIS deployment. Therefore, characterizing the Energy Efficiency (EE) performance of active RIS-assisted communication systems and quantifying the effect of finite-bit phase quantization errors on EE are essential for system design and practical implementation. To this end, a discrete phase-shifted active RIS-assisted communication system over Rayleigh fading channels is investigated. The EE loss caused by phase quantization errors is analyzed, approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE are derived, and the relationship between RIS EE and user EE is established, providing theoretical guidance for the practical deployment of active RIS.  Methods  Based on the law of large numbers and Taylor series expansion, closed-form expressions for the user EE loss and its approximation are derived. The effects of system parameters on EE are investigated by expressing EE as explicit univariate functions. Ferrari’s method and the Lambert W function are then employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE. Finally, the relationship between RIS EE and user EE is established using the law of large numbers and the Lambert W function.  Results and Discussions  User EE is expressed as a function of six parameters: the number of quantization bit (\begin{document}$ k $\end{document}), power allocation factor (\begin{document}$ \beta $\end{document}), the number of RIS elements (\begin{document}$ N $\end{document}), the total power sum of base station and active RIS (\begin{document}$ {P}_{\text{t}} $\end{document}), the noise at active RIS (\begin{document}$ \sigma _{\text{r}}^{2} $\end{document}), and the noise at user (\begin{document}$ \sigma _{\text{u}}^{2} $\end{document}). First, the EE loss decreases as \begin{document}$ k $\end{document} increases. When \begin{document}$ k $\end{document}=3, the difference between the approximate EE loss and the lossless case is less than 0.026 8 Mbit/J, while the difference between the EE loss and the lossless case is less than 0.026 5 Mbit/J (Fig. 3). Therefore, 3-bit to 4-bit discrete phase shifters achieve performance close to that of continuous phase shifters. Second, user EE exhibits a unimodal dependence on both \begin{document}$ \beta $\end{document} and \begin{document}$ N $\end{document}. The approximate optimal solution for \begin{document}$ \beta $\end{document} differs from the exact optimal solution obtained by the Dinkelbach algorithm by less than 0.01 (Fig. 4), whereas the approximate and exact optimal solutions for \begin{document}$ N $\end{document} are identical (Fig. 5), demonstrating the high accuracy of the proposed approximations. Third, user EE exhibits a unimodal trend as \begin{document}$ {P}_{\text{t}} $\end{document} increases. Higher phase quantization precision produces a higher EE peak while requiring a lower optimal \begin{document}$ {P}_{\text{t}} $\end{document} to achieve the maximum EE (Fig. 6). In addition, user EE decreases as both \begin{document}$ \sigma _{\text{r}}^{2} $\end{document} and \begin{document}$ \sigma _{\text{u}}^{2} $\end{document} increase. User EE is more sensitive to the amplification noise introduced at the RIS, indicating that reducing the RIS noise power yields a greater EE improvement (Fig. 7). Finally, user EE first increases and then decreases sharply to zero as RIS EE increases. The signal-to-noise ratio at the RIS is identified as the key factor governing the relationship between RIS EE and user EE (Fig. 8).  Conclusions  The EE performance of active RIS-assisted wireless networks employing discrete phase shifters over Rayleigh fading channels is investigated. First, closed-form expressions are derived for the user EE in the lossless case, the lossy case, and the approximate-loss case. Simulation results demonstrate that 3-bit to 4-bit discrete phase shifters closely approach the performance of continuous phase shifters. Next, explicit functions describing the effects of key system parameters on user EE are established. Ferrari’s method and the Lambert W function are employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE, and both exhibit negligible errors relative to the exact solutions. Finally, the relationship between RIS EE and user EE is established, demonstrating that user EE initially increases and subsequently decreases to zero as RIS EE increases.
Channel Estimation for MIMO-OFDM Based on Adaptive Transformer Network in High-speed Mobile Scenarios
LIAO Xi, HE Xiangni, ZHANG Zhe, WANG Yang
Available online  , doi: 10.11999/JEIT260075
Abstract:
  Objective  Accurate Channel State Information (CSI) is essential for coherent detection, beamforming, and adaptive resource allocation in Multiple-Input Multiple-Output Orthogonal Frequency Division Multiplexing (MIMO-OFDM) systems. In high-mobility scenarios, large Doppler shifts and multipath propagation jointly produce doubly selective fading, destroy subcarrier orthogonality, and intensify inter-carrier interference. Therefore, conventional Least Squares (LS) and Linear Minimum Mean Square Error (LMMSE) estimators exhibit substantial performance degradation. Existing deep learning-based channel estimation methods provide strong nonlinear modeling capability but often lack sufficient adaptability to variations in Signal-to-Noise Ratio (SNR), delay spread, and maximum Doppler shift. Moreover, channel physical parameters are not efficiently exploited as prior information. To address these limitations, this paper proposes AdaFiT, an adaptive Transformer-based channel estimation network incorporating Feature-wise Linear Modulation (FiLM) for high-mobility MIMO-OFDM systems.  Methods  AdaFiT performs channel estimation by jointly exploiting local time-frequency features, global time-frequency dependencies, and channel-adaptive feature modulation. The network takes LS estimates at pilot positions together with SNR, delay spread, and maximum Doppler shift as inputs. A separable two-dimensional linear upsampling module first reconstructs sparse pilot estimates over the complete OFDM time-frequency grid by independently processing the real and imaginary components in the frequency and time dimensions. A convolutional feature enhancement module then extracts robust local feature representations. Specifically, a complex feature-mixing layer jointly models the real and imaginary components of multi-antenna complex channel responses, while multi-scale convolutional blocks with channel attention capture local time-frequency features and suppress noise. Subsequently, an FiLM-based channel-adaptive feature modulation module embeds the three channel physical parameters through independent multilayer perceptrons and combines them into a channel-condition representation. The resulting representation generates feature-wise scaling and shifting coefficients to dynamically recalibrate block-embedded feature sequences according to changing channel statistical characteristics. Finally, the modulated feature sequences are processed by a Transformer encoder with learnable two-dimensional positional encoding to capture long-range dependencies across the time-frequency grid and antenna dimensions. A residual reconstruction module combines global and local feature representations to generate accurate channel estimates while preserving fine local details.  Results and Discussions  Simulation results are obtained under the CDL-C and CDL-A channel models. The LS-based bilinear interpolation method, LMMSE, AdaFortiTran, and AdaFiT without the channel-adaptive feature modulation module are selected as benchmark methods. Their Mean Squared Error (MSE) performance is evaluated under different SNR, maximum Doppler shift, and delay spread conditions. Under the CDL-C channel model, AdaFiT achieves the lowest MSE across the entire SNR range (Fig. 3). At an SNR of 0 dB, the MSE is approximately 2.5 dB lower than that of AdaFortiTran, and the performance gain increases to approximately 7 dB at an SNR of 30 dB. Compared with AdaFiT without the channel-adaptive feature modulation module, the proposed model achieves a maximum MSE improvement of approximately 2.1 dB, confirming the effectiveness of the proposed channel-adaptive feature modulation module. When the maximum Doppler shift increases from 200 Hz to 1 400 Hz, AdaFiT consistently achieves the lowest MSE (Fig. 4). In the high-Doppler region (1 000~1 400 Hz), AdaFiT outperforms AdaFiT without the channel-adaptive feature modulation module and AdaFortiTran by approximately 2 dB and 3.5 dB, respectively, demonstrating superior robustness under rapidly time-varying channel conditions. For delay spreads ranging from 100 ns to 700 ns, AdaFiT also achieves the lowest MSE throughout the entire range (Fig. 5), providing gains of approximately 3 dB and 5.5 dB over AdaFiT without the channel-adaptive feature modulation module and AdaFortiTran, respectively. These results demonstrate that the proposed channel-adaptive feature modulation mechanism effectively improves model adaptability to time-selective and frequency-selective fading. To further evaluate the generalization capability of AdaFiT, additional simulations are conducted under the 3GPP CDL-A channel model. At SNRs of 0~5 dB, AdaFiT outperforms LMMSE by approximately 4~5 dB, and the performance gain increases to approximately 7 dB at SNRs of 20~30 dB (Fig. 6(a)). Compared with AdaFortiTran, AdaFiT achieves an MSE gain of approximately 2 dB under low-SNR conditions, which increases to approximately 5.5 dB under high-SNR conditions. In the maximum Doppler shift experiment, the MSE of AdaFiT remains between approximately –38 dB and –37 dB over the range of 200~800 Hz and is approximately 5 dB lower than that of AdaFortiTran (Fig. 6(b)). Although the MSE increases when the maximum Doppler shift exceeds 1 000 Hz, AdaFiT still provides an approximately 6 dB gain over LMMSE at 1 400 Hz and continues to outperform AdaFortiTran. Across the entire delay spread range, AdaFiT achieves gains of approximately 4~6 dB over LMMSE and approximately 4~5.5 dB over AdaFortiTran (Fig. 6(c)).  Conclusions  AdaFiT, an adaptive Transformer-based channel estimation network, jointly exploits local time-frequency features, global time-frequency dependencies, and channel-adaptive feature modulation to improve channel estimation accuracy. Simulation results under the CDL-C and CDL-A channel models demonstrate that AdaFiT consistently achieves lower MSE than the LS-based bilinear interpolation method, LMMSE, AdaFortiTran, and AdaFiT without the channel-adaptive feature modulation module under different SNR, maximum Doppler shift, and delay spread conditions. These results confirm the effectiveness of the proposed channel-adaptive feature modulation mechanism and demonstrate that AdaFiT maintains high estimation accuracy and stable performance across different channel models, indicating strong adaptability to dynamic channel environments.
Multi-Frequency Feature Interaction and Adaptive Fusion for cooperative Spacecraft 6D Pose Estimation
TONG Wei, LIN Xi, YAN Ying, LIN Jinxing, LI Tao, WU Qi
Available online  , doi: 10.11999/JEIT260452
Abstract:
Spacecraft 6D pose estimation aims to determine the relative pose between the target spacecraft and the service spacecraft in the spatial coordinate system, which is a critical step for a series of close-range operation tasks, such as failed satellite cleaning, space debris capture, on-orbit spacecraft manipulation, and space station rendezvous and docking. In recent years, CNN-based 6D pose estimation methods have received widespread attention. However, their over-reliance on convolutional network architectures makes them sensitive to image textures and limits their capacity to effectively model long-range contextual information. Moreover, current mainstream methods typically adopt a pipeline design consisting of object detection followed by pose estimation, which suffers from limited diversity in feature extraction and is sensitive to low-light conditions and complex background interference. To address these issues, this paper proposes a spacecraft pose estimation network based on multi-spectral feature interaction and dynamic fusion. Specifically, the network first leverages backbone networks with different receptive fields to separately extract spatial high-frequency features (such as semantic and edge details) and spatial low-frequency features (such as global structural information). On this basis, a Transformer-based feature matching mechanism is employed to perform self-attention and cross-attention feature interactions, thereby aggregating long-range contextual information. To further exploit the rich frequency-domain feature representations, a frequency-guided feature module is introduced to dynamically fuse multi-spectral features. Finally, extensive experiments on spacecraft pose estimation datasets demonstrate that the proposed method achieves competitive performance and strong generalization ability, showing advantages over existing methods.  Objective  Due to the limitation of local receptive field of CNN convolution operator, the mining of remote correlation information is often not ideal. In contrast, Transformer with multi-head intra-attention mechanism is more effective in globally modeling long-range context information and is good at processing multi-scale objects and small spacecraft at a long distance. Therefore, enhancing the processing ability of low-frequency and high-frequency features is of great significance for improving the accuracy.  Methods  This work proposes a 6D pose estimation network with Transformer-based multi-frequency feature interaction and fusion. The framework consists of two parts. One is to extract high-frequency features such as image semantics and edges via ResNet18 and leverage them for spacecraft semantic segmentation directly. The other is to extract global low-frequency features such as image texture through DarkNet53. Then Transformer-based feature interaction module is designed to perform attention between high-frequency and low-frequency features, enabling the aggregation of long-range contextual information, which can enhance the diversity of texture-less spacecraft image features and promote network optimization.  Results and Discussions  The proposed method estimates spacecraft 6D pose via multi-frequency feature interaction and adaptive fusion. On the SwissCube dataset, it achieves an overall ADI-0.1d accuracy of 82.31%, outperforming CA-SpaceNet (79.39%) and WDR* (78.78%). On the SPEED dataset, the combined error eq+et is 0.0299, which is 22.3% lower than CA-SpaceNet (0.0385) and 25.2% lower than WDR* (0.0400). Qualitative results in Fig. 5 and Fig. 7 show that predicted key points are significantly closer to the ground truth, especially under challenging low-frequency conditions. Ablation studies confirm that the feature interaction module alone increases overall accuracy from 79.39% to 80.67%, and adding frequency-guided fusion further raises it to 82.31%. These results demonstrate that the proposed framework enhances pose estimation robustness and meets the stringent requirements for on-orbit servicing and space rendezvous.  Conclusions  To enhance the accuracy of spacecraft 6D pose estimation under extreme atmospheric environments, this work innovatively designs a sub-branch based on multi-frequency feature interaction and additional semantic edge segmentation, and overcomes the limitation of low feature extraction efficiency of existing methods by aggregating long-range multi-frequency features. In addition, a frequency-guided dynamic feature fusion module is introduced to fully leverage the rich frequency-domain feature representation. Comprehensive experimental comparison with mainstream methods on SwissCube and SPEED datasets verifies that the proposed work can enhance the representation of feature information and improve the robustness of spacecraft pose estimation.
Frequency-Domain Decoupling and Spatial-Prior-Constrained Detection Method for Infrared Dim and Small Targets
LIU Minglong, JIANG Tingyao, LI Yulan
Available online  , doi: 10.11999/JEIT260847
Abstract:
  Objective  Infrared dim and small target detection exploits thermal-radiation differences between targets and backgrounds for passive sensing in air-ground inspection, maritime surveillance, and wide-area warning. Under long-range imaging, low signal-to-noise ratios, and complex backgrounds, targets occupy few pixels and exhibit weak textures, blurred contours, and low contrast. Bounding-box detection is well suited to edge-deployed rapid detection and subsequent tracking initialization by directly predicting class confidence and location. Although the Real-Time Detection Transformer (RT-DETR) provides an efficient end-to-end framework, its application faces three limitations: repeated downsampling weakens local details and initial localization cues; intra-scale global interaction insufficiently distinguishes high-frequency target responses from low-frequency background context; and cross-scale fusion may propagate homogeneous clutter without explicit spatial constraints. Therefore, a Frequency-Domain Decoupling and Spatial-Prior-Constrained DETR (FSP-DETR) is proposed to coordinate shallow detail preservation, target-background decoupling, and localization constraints.  Methods  FSP-DETR is built on RT-DETR-R18 and follows backbone feature extraction, intra-scale interaction, cross-scale fusion, and query-based decoding (Fig.1). The three modules operate at successive stages to address the three identified limitations. First, a Subspace Progressive Attention Modulated Inverted Residual Block (SPA-MIRB) is embedded in the Residual Network 18 (ResNet18) backbone to preserve shallow details (Fig.2). Its inverted-residual branch combines pointwise and depthwise convolutions with a residual connection to retain edges, spots, and weak textures; its attention branch partitions channels into subspaces and progressively transfers interactions to strengthen weak targets. Second, a Frequency-Domain Asymmetrically Decoupled Attention-based Intra-scale Feature Interaction module (FD-AIFI) replaces the Attention-Based Intra-Scale Feature Interaction module (AIFI) (Fig.3). Unified global interaction is divided into high-pass local-attention and low-pass global-attention paths. The former uses local-window attention to preserve compact target responses and limit distant clutter. The latter retains query resolution but pools key and value features, compressing low-frequency context while preserving spatial queries and target-position indices. Their outputs are concatenated channel-wise and linearly projected. Third, a Spatial-Prior-Constrained CNN-based Cross-scale Feature-fusion Module (SPC-CCFM) retains the upsampling, concatenation, and convolution path while introducing a High-Order Spatial Representation Module (HSRM) and an Overall-Level Mapping Injection Mechanism (OMIM). HSRM aligns multilevel features and constructs spatial passthrough, high-order spatial-relation, and low-order fidelity branches (Fig.4). The branches preserve original spatial responses, model spatially separated yet response-similar clutter through hypergraph aggregation, and supplement local edges and weak textures. OMIM adapts and injects the prior into each fusion node to constrain cross-scale refinement.  Results and Discussions  Experiments are conducted on IRSTD-1k and NUAA-SIRST, with pixel-level masks converted into single-class minimum bounding rectangles. Component ablation shows that SPA-MIRB, FD-AIFI, and SPC-CCFM each improve mean Average Precision at an Intersection over Union threshold of 0.5 (mAP@0.5) on both datasets (Table 1). FSP-DETR achieves 89.03% and 98.75% mAP@0.5 on IRSTD-1k and NUAA-SIRST, exceeding RT-DETR-R18 by 3.39 and 3.08 percentage points. Parameters decrease from 19.87 million to 17.24 million and computational cost from 56.9 to 50.4 billion floating-point operations, by 13.2% and 11.4%. Although some variants yield higher precision or recall at one operating point, the complete model obtains the highest mAP@0.5 with lower complexity. Replacement experiments support the structural choices: SPA-MIRB and FD-AIFI achieve the highest mAP@0.5 among their counterparts, while SPC-CCFM exceeds both fusion alternatives by 0.48 and 0.26 percentage points, respectively, highlighting spatial constraints over complexity reduction (Table 2). SPA-MIRB reaches 96.75% mAP@0.5 with 15.28 million parameters and 46.8 billion floating-point operations, providing a balanced accuracy-complexity trade-off. A balanced path-allocation factor of 0.50 performs best, reaching 97.27% mAP@0.5 and 53.04% mAP@0.5:0.95 on NUAA-SIRST (Table 3). Attention maps show concentrated target-neighborhood responses in the high-pass path and broader, smoother background responses in the low-pass path (Fig.5). Under unified settings, FSP-DETR obtains the highest mAP@0.5 on both datasets, with 172 frames/s and an average per-image forward time of 5.81 ms (Table 4). It also surpasses deeper RT-DETR variants with fewer parameters, less computation, and shorter forward time, showing the advantage of task-oriented adaptation over backbone enlargement. Detection visualization shows better agreement between predicted boxes and target regions in cluttered scenes (Fig.6). Instance-level analysis shows missed-detection rates decrease from 15.70% to 13.22% on IRSTD-1k and from 10.71% to 5.36% on NUAA-SIRST, while average false positives per image decrease from 0.490 to 0.380 and from 0.209 to 0.070 (Table 5). Center-offset errors also decrease on both datasets. However, the scale error on NUAA-SIRST increases slightly from 10.63% to 11.05%, indicating that scale regression for extremely small targets remains challenging.  Conclusions  FSP-DETR coordinates shallow detail preservation, frequency-domain decoupled intra-scale interaction, and spatial-prior-constrained cross-scale fusion for end-to-end bounding-box detection. It improves accuracy, missed-detection control, false-alarm suppression, and center localization while reducing complexity and maintaining efficient inference. An accuracy-efficiency balance is achieved in complex infrared scenes. Future work will investigate multi-frame spatiotemporal modeling for thermal-crossover backgrounds and dense dim-target scenes.
Lightweight image-to-image Steganography Based on Improved Emd and Dual-domain Graph Convolutional Network
DUAN Xintao, CHEN Rusheng, LI Sen, QIN Chuan
Available online  , doi: 10.11999/JEIT260857
Abstract:
  Objective  Image steganography embeds secret information into a cover image to achieve secure transmission, and it serves as an important technique in confidential communication and privacy protection. With the development of deep learning, learning-based steganography has notably improved embedding capacity and reconstruction quality. However, three requirements, namely high steganographic performance, strong resistance to steganalysis, and lightweight design, are difficult to satisfy simultaneously, and this trade-off has become the main bottleneck for practical deployment. Single-domain processing methods cannot balance the three objectives, whereas existing high-performance models are structurally complex and hard to deploy in resource-constrained environments. Moreover, mainstream dual-domain schemes usually assume that the embedding distortion follows a continuous distribution in both the spatial domain and the frequency domain, yet actual steganographic modifications tend to concentrate in the discontinuous and complex texture regions of the cover image. To address these problems, a lightweight image steganography network is designed in this study to jointly optimize the embedding and extraction paths, so that the stego-image quality and the anti-steganalysis ability are improved while a low computational cost and a low inference latency are maintained.  Methods  A lightweight steganographic network named GISNet is proposed, in which an encoder-decoder architecture with skip connections is adopted and the hiding network and the extraction network are made structurally symmetric without weight sharing. First, an Improved Bidimensional Empirical Mode Decomposition (IBEMD) module is applied, by which the secret image is adaptively decomposed into several intrinsic mode components and one residual component, so that the hidden information is dispersed hierarchically among components of different morphology. Gaussian blur is used to replace extremum interpolation for estimating the local mean envelope, the separability of the Gaussian kernel is exploited to reduce the computational complexity, and a parameter-binding mechanism is introduced to guarantee deterministic and reversible reconstruction. Next, a multi-scale spatial-frequency block (IMFB) is designed, in which multi-branch dilated convolutions, an attention mechanism, dynamic gating, and a frequency-domain perception unit are integrated, so that the spatial features and the frequency features are deeply fused, the anti-detection ability is enhanced, and the high-fidelity extraction of the secret information is ensured. Finally, an improved graph fusion neural network (GFNN) is employed as the bottleneck layer, in which a sparse graph is constructed in the feature space through the K-nearest-neighbor algorithm and messages are propagated only among non-local nodes with high similarity, so that the long-range pixel associations are explicitly modeled, the modification patterns in discontinuous texture regions are characterized, and the model complexity is substantially reduced. The hiding network and the extraction network are jointly optimized by a four-term loss function that combines the hiding loss, the restriction loss, the Laplacian pyramid loss, and the perceptual loss.  Results and Discussions  GISNet achieves the best overall image-hiding and recovery performance on DIV2K, COCO, and ImageNet, demonstrating high reconstruction quality and stable cross-dataset generalization (表1). On DIV2K, the PSNR values reach 58.07 dB for cover/stego image pairs and 60.10 dB for secret/recovered-secret image pairs. The cover and stego images are visually indistinguishable, and the residual maps remain nearly black after 30-fold magnification (图5). Ablation experiments show that replacing DWT with IBEMD improves the PSNR values by 8.04 dB and 5.57 dB, respectively (表2). The complete combination of IBEMD, IMFB, and GFNN provides the best results, confirming the complementary effects of hierarchical information dispersion, multiscale spatial-frequency mapping, and nonlocal feature association (表3). Multi-dilation-rate branches and joint spatial-frequency processing further improve hiding quality and recovery accuracy (表4,表5). The detection accuracies under four steganalysis methods range from 49.35% to 49.85%, indicating that the stego images are difficult to distinguish from natural cover images (表6). GISNet requires only 8.00 M parameters and 7.88 GFLOPs, with an inference time of 76 ms (表7). These results demonstrate that GISNet effectively balances visual quality, recovery accuracy, resistance to steganalysis, and computational efficiency.  Conclusions  A lightweight dual-domain graph convolutional network for image steganography, named GISNet, is proposed in this paper. The experimental results demonstrate the following. (1) The improved bidimensional empirical mode decomposition disperses the secret information hierarchically and supports deterministic and reversible reconstruction, by which a reliable basis is provided for high-quality hiding and recovery. (2) The multi-scale spatial-frequency block and the graph fusion neural network jointly improve the stego-image quality and the anti-steganalysis ability, so that the stego images can resist detection by multiple steganalysis tools. (3) The lightweight design, which is based on depth-wise separable convolutions and sparse graph construction, significantly reduces the model complexity and the computational cost, by which the model is made suitable for deployment in resource-constrained environments. Future work will focus on the robustness against channel interference and the extension to multi-image and cross-modal steganography, so that the security and applicability of the scheme are further enhanced.
A Structure-Preserving Semantic Transmission Method for Low-Bandwidth Networks
HU Tianwei, ZHANG Xiangrui, CHEN Jian, DUAN Haodong, JIA Jie
Available online  , doi: 10.11999/JEIT260525
Abstract:
  Objective  Low-bandwidth visual transmission is essential for edge-intelligence applications such as disaster inspection, underwater exploration, and remote assistance. These scenarios require visual communication systems to simultaneously achieve low bitrate, high structural fidelity, stable color reconstruction, and low end-to-end latency. However, conventional image coding methods suffer from severe quality degradation at extremely low bitrates due to the digital-cliff effect and compression artifacts. Although semantic communication provides a promising solution by transmitting task-relevant representations rather than pixel-level information, existing approaches still face challenges in balancing bitrate efficiency, reconstruction fidelity, and decoding complexity. Therefore, a structure-preserving semantic transmission method, termed Edge-Link, is proposed for low-bandwidth networks.  Methods  Edge-Link adopts a multimodal decoupled representation framework that separates image information into three complementary streams: a semantic stream, a structural stream, and a low-frequency appearance stream. The semantic stream is extracted using a frozen CLIP encoder to provide global semantic guidance, while the structural stream is obtained from Canny edge information to explicitly preserve object boundaries. A low-resolution color map is introduced as the appearance stream to maintain global color distribution and illumination characteristics with minimal transmission overhead. Furthermore, a FastSAM-based semantic gateway is developed to distinguish foreground objects from background regions, and an object-aware bitrate allocation strategy is designed to prioritize important semantic regions under bandwidth constraints. At the receiver, a deterministic dual-stream generation network based on SPADE is proposed, where structural information provides spatial constraints and semantic features guide the reconstruction process, avoiding the high latency caused by iterative diffusion sampling. A real LoRa communication prototype based on GNU Radio and USRP is also implemented to validate transmission feasibility and robustness under practical wireless conditions.  Results and Discussions  Experiments are conducted on Set14 and DIV2K to evaluate transmission efficiency, reconstruction quality, perceptual naturalness, and latency. The proposed method achieves an average payload of 10.96 KB for 512×512 images, corresponding to about 0.31 bpp, while keeping the end-to-end inference latency below 60 ms; under a LoRa narrowband link of about 8 kbps, the airtime of a single frame is about 11 s, verifying its feasibility in bandwidth-limited environments (Table 1). The ablation study shows that the object-aware coding strategy improves the bitrate-quality trade-off by reducing the average payload on DIV2K from 126.09 KB to 93.55 KB while preserving better foreground reconstruction quality (Table 2). Compared with JPEG, the proposed method also shows better structural preservation and perceptual quality at low bitrates; for example, on image 0843 with a payload of 15.54 KB, the global SSIM is improved from 0.783 to 0.963, and the subject-region LPIPS is reduced from 0.561 to 0.307 (Table 4). Under the same bitrate constraint, the proposed method further outperforms VQGAN and ControlNet+Canny, achieving PSNR, SSIM, LPIPS, and NIQE values of 23.738, 0.622, 0.223, and 4.570, respectively, indicating a better balance between fidelity and perceptual quality in low-bandwidth semantic reconstruction (Table 5).  Conclusions  A structure-preserving semantic transmission framework for low-bandwidth networks is presented. By combining multimodal decoupled representation, object-aware bitrate allocation, and deterministic dual-stream reconstruction, the framework balances transmission efficiency, structural fidelity, perceptual quality, and decoding latency. The reported results show that the proposed method is well suited to high-reliability edge visual communication scenarios in which accurate contours, stable color appearance, and efficient inference are simultaneously required. The real-link validation on a LoRa prototype further suggests its practical potential for bandwidth-constrained edge networks.
Anomaly Detection on Irregular Signals in Adaptive Decay Reservoir Network Model Space
CHEN Ao, LI Wenpeng, XIE Xiaoyan, CHEN Pengpeng
Available online  , doi: 10.11999/JEIT260423
Abstract:
  Objective  Signals acquired from industrial systems often exhibit irregular sampling due to sensor instability, intermittent operation, communication dropout, and multi-source asynchronous acquisition. This irregularity violates the uniform sampling assumption of most signal analysis methods and challenges anomaly detection. Interpolation and resampling may distort the underlying temporal dynamics, while continuous-time models based on neural ordinary differential equations as well as time-aware Transformers suffer from high training costs and strong dependence on large training sets, making them impractical under limited training resources. Model space learning offers an alternative by fitting each signal with a dynamic model and analyzing the fitted models instead of the raw signals. However, existing model space methods for irregular sampling rely on fixed reservoir configurations and lack adaptive optimization of the model space. This paper aims to develop an anomaly detection framework for irregularly sampled signals that is simultaneously robust to non-uniform intervals and efficient to train.  Methods  An adaptive model space learning framework based on the Adaptive Decay Reservoir Network (ADRN) is proposed (Fig. 1). ADRN extends the echo state network by introducing an exponential decay mechanism derived from a linear ordinary differential equation (Fig. 2). Between consecutive observations, the hidden state decays according to a learnable decay rate over the actual elapsed time, so that the state update naturally adapts to non-uniform sampling intervals without numerical ordinary differential equation solvers. Each new observation then updates the decayed state through a nonlinear activation. Ridge regression with a closed-form solution then fits a readout model mapping hidden states to the original signal, and the fitted readout weights serve as a compact fixed-dimensional representation regardless of signal length. The model space is further optimized by two complementary losses. A time-interval-weighted reconstruction loss assigns higher weights to larger intervals, preventing densely sampled segments from dominating the optimization and improving fitting quality under non-uniform sampling. A separability loss inspired by Fisher discriminant analysis acts through a learnable projection matrix to minimize intra-class scatter and maximize inter-class separation in the projected model space. The two losses are combined into a joint objective that simultaneously updates the reservoir parameters, the decay rate, and the projection matrix, and gradients propagate through the differentiable closed-form ridge regression to enable end-to-end optimization. A downstream classifier, namely a support vector machine on CWRU and SU and a random forest on the higher-dimensional TEP model space, performs the final detection. The echo state property of ADRN is formally established, and the resulting spectral-norm condition is more relaxed than the classical one, allowing richer reservoir dynamics.  Results and Discussions  Experiments cover the CWRU bearing dataset (five subsets, 50% missing rate), the SU gearbox dataset (30%, 50%, and 70% missing rates), and the Tennessee Eastman Process (TEP) chemical dataset (19 classes, three missing rates), with only 200 labeled signals per subset for training on CWRU and SU. The proposed method achieves the highest accuracy in 7 of the 8 CWRU and SU settings (Table 1), with accuracies ranging from 87.3% to 93.8% on CWRU. On SU, it maintains 94.4% accuracy even at a 70% missing rate, and the fluctuation across missing rates is only 2.8%, in contrast to 18.7% for ODE-RNN, while Neural CDE drops from 86.8% to 62.6%. Ablation studies confirm the contribution of each component (Table 1, Table 2). Removing the exponential decay reduces accuracy by up to 28.4 percentage points, and the interval weighting and the separability loss contribute complementary gains of 2.0 and 3.1 percentage points on SU at the 70% missing rate. t-SNE visualization shows that the optimized model space exhibits compact and clearly separated classes (Fig. 3). Training on a CWRU dataset completes in about 50 seconds, over two orders of magnitude faster than neural ordinary differential equation methods, which require 3000 to 7000 seconds (Table 3). Hyperparameter analysis indicates that a loss balance coefficient between 0.1 and 0.3 performs well and that a small reservoir suffices (Table 4). On TEP, a non-rotating-machinery industrial object whose faults manifest as changes in process dynamics, the proposed method attains the highest accuracy among all compared methods at every missing rate, with the largest margin of 7.0 percentage points at the highest missing rate (Table 5).  Conclusions  The ADRN based model space learning framework provides an effective solution for anomaly detection on irregularly sampled signals. The exponential decay mechanism enables interval-aware state updates without numerical ordinary differential equation solvers, ridge regression yields efficient closed-form readout fitting, and the joint optimization of the reconstruction and separability losses produces a model space with both high fitting quality and strong discriminative structure. The framework requires only a small reservoir of 10 to 50 dimensions and completes training within one minute, making it well suited to scenarios with limited training resources and irregular sampling. Future work includes adaptive reservoir sizing, extension to multivariate joint modeling, and validation in further domains such as structural health monitoring and biomedical or meteorological time series.
An ECO Repair Method for Max Transition Violations in Multi-Load Nets
FAN Lingyan, XU Xinchen, HUANG Cankun, SHEN Zhengnuo, DENG Jiangxia, LIU Hailuan
Available online  , doi: 10.11999/JEIT260600
Abstract:
  Objective   Max transition violations in digital integrated circuit physical design may increase gate delay, reduce timing margin, and introduce additional power-consumption and signal-integrity risks. With technology scaling and the increasing complexity of high-performance SoC and CPU designs, long interconnects and heavy effective loads make such violations more prominent in post-routing optimization and ECO stages. Conventional repair methods usually rely on manual analysis or global heuristics, such as total-wire-length-ratio-based buffer insertion, which may lead to low repair efficiency, inaccurate repair targeting, and redundant buffer insertion in multi-load nets. Although these approaches can alleviate some violations in simple cases, they often fail to accurately identify the real violating branches in multi-load nets. As a result, repair targeting becomes insufficient and redundant buffer insertion is likely to occur. To address these limitations, a path-level buffer insertion method is proposed for max transition violation repair in multi-load nets.  Methods   The proposed method first reconstructs the physical topology of the target net from routing information extracted from the Design Exchange Format (DEF) file. Routing endpoints, turning points, and via connection points are abstracted as physical nodes with coordinate and metal-layer attributes, and a weighted undirected graph is established to preserve branch structures and cross-layer connectivity (Fig. 3). To reduce the influence of small coordinate deviations, a spatial-tolerance-based node merging strategy is introduced during graph construction. Since the violation coordinates reported by timing analysis tools may not lie exactly on valid routed segments, a vector-projection-based coordinate snapping strategy is then adopted to align logical violation coordinates with the actual physical topology (Fig. 4). After endpoint binding, the actual physical propagation path from the driver to each violating load is recovered by Dijkstra shortest-path search. Based on the Elmore-model intuition that inserting a buffer near the midpoint of a long interconnect can effectively segment the distributed RC load, the midpoint of each recovered path is selected as the initial candidate insertion point. The exact insertion coordinate is obtained by accumulating segment lengths along the path and interpolating on the segment where half of the total path length is reached. To improve robustness, the initial candidate point is further expanded into an effective candidate interval with a spatial tolerance factor. For multi-load nets, different violating branches may share long common physical segments. Therefore, an interval-intersection-based shared-buffer optimization strategy is introduced to merge overlapping candidate intervals into shared insertion regions, thereby reducing redundant buffer insertion (Fig. 5-Fig. 7).  Results and Discussions   Experiments are conducted on five designs, including CPU, SAS, RAID, PCIe, and HBA, implemented in the UMC 28 nm process. Synopsys IC Compiler II is used for physical implementation, and PrimeTime is used for timing analysis. A violation-margin threshold of -8 ps is adopted, and only paths below this threshold are included in the repair and evaluation. The complete ECO flow retains the existing repair method for single-load nets and applies the proposed path-level method to multi-load nets. To illustrate the repair mechanism, a representative multi-load violating net is selected for detailed analysis. The net contains one driver and fourteen loads, among which thirteen violating load paths share a long common routed trunk and exhibit max transition violation margins ranging from -75.2 ps to -38.3 ps. If these paths are repaired independently, thirteen nearby buffer insertion demands are generated on the shared trunk. After interval-intersection-based merging, one shared buffer is sufficient to repair all thirteen violating paths jointly (Fig. 8). This case study indicates that path recovery improves repair targeting, while shared-buffer optimization improves resource efficiency. Across the five designs, the complete ECO flow reduces the total number of max transition violations from 10,616 to 108, corresponding to an overall repair rate of 98.98% (Table 1). The CPU, SAS, and RAID designs are completely repaired, while only 17 and 91 violations remain in PCIe and HBA, respectively. Among the 8,216 multi-load violating paths, 8,166 are successfully repaired, corresponding to a multi-load violation repair rate of 99.39%. If these successfully repaired paths were handled independently, 8,166 buffer insertions would be required. After shared-buffer optimization, only 1,548 buffers are inserted, giving an overall compression ratio of 81.04%. The setup worst negative slack, total negative slack, and number of violating paths remain generally stable before and after repair, with only minor fluctuations observed in individual designs (Table 2). These results show that the repair flow does not cause obvious degradation in setup timing quality. Further analysis indicates that the residual violations in PCIe and HBA mainly occur when candidate buffer locations fall inside standard-cell placement blockages around hard macros or IP cores. Metal routing is allowed through these regions, but buffers cannot be legally placed, revealing a current limitation in the physical-feasibility handling of candidate insertion locations.  Conclusions   A path-level buffer insertion method is proposed for max transition violation repair in multi-load nets. By combining physical topology reconstruction, violation-coordinate snapping, shortest-path-based path recovery, midpoint-guided candidate generation, and interval-intersection-based shared-buffer optimization, the proposed method improves repair targeting and reduces redundant buffer insertion. Experimental results on five designs show that the complete ECO flow reduces the number of max transition violations from 10,616 to 108, achieving an overall repair rate of 98.98%. Among the 8,216 multi-load violating paths, 8,166 are successfully repaired, corresponding to a repair rate of 99.39%. For these successfully repaired paths, shared-buffer optimization reduces the number of buffer insertion demands from 8,166 to 1,548, corresponding to a compression ratio of 81.04%, without causing obvious degradation in setup timing quality. The remaining violations are mainly associated with standard-cell placement blockages around hard macros or IP cores, indicating that the physical-feasibility handling of candidate insertion locations should be further improved.
Design of a Channel-Adaptive Denoiser for Digital Semantic Communications
WU Yanjun, LIU Zhangyuhang, YANG Wenxin, YAN Mubiao, ZHOU Hao, ZHAO Yajun, XIE Zhuochen, LIANG Xuwen
Available online  , doi: 10.11999/JEIT260523
Abstract:
  Objective  Practical semantic communication should simultaneously satisfy two requirements: compatibility with existing digital communication infrastructures and robustness under varying channel conditions. Semantic-oriented modulation (SOM) provides a feasible way to map continuous semantic features into layered digital constellation symbols, thereby making semantic transmission compatible with conventional digital systems. However, the digitization process also introduces structured quantization distortion, which makes receiver-side recovery more difficult than in continuous semantic transmission. Although diffusion models have shown strong capability in channel-adaptive semantic recovery, directly applying them to SOM-based digital semantic communication is still limited by the structured distortion introduced by SOM. Therefore, this paper focuses on channel-adaptive receiver design for digital semantic communication and investigates how to compensate SOM-induced structured distortion before subsequent recovery.  Methods  An SOM-based digital semantic communication system for image transmission over an additive white Gaussian noise (AWGN) channel is considered. The proposed receiver adopts a two-stage structure composed of a Quantization Noise Predictor (QNP) and a diffusion recovery module. In the first stage, QNP estimates and compensates the structured quantization distortion introduced by SOM from the layer-wise soft received symbols. In the second stage, the compensated semantic representation is further refined by a diffusion denoiser, whose inference step number is adaptively selected according to the estimated signal-to-noise ratio (SNR). The QNP includes a shared feature extraction frontend, a classification branch exploiting discrete SOM constellation priors, and a regression branch performing fine-grained continuous distortion compensation. A Feature-wise Linear Modulation (FiLM) mechanism is used to incorporate SOM parameters and channel-state information, so that the same QNP can adapt to different modulation configurations and channel conditions. In addition, a composite loss with classification loss, regression loss, and distribution regularization is designed to improve the statistical properties of the compensated residual noise.  Results and Discussions  Experiments are conducted on the CLIC dataset using PSNR and MS-SSIM. First, the proposed method is compared with VAE, VAE+Diff, VAE+SOM, VAE+SOM+QNP, VAE+SOM+Diff, and JCM. The results show that direct SOM-based digitization causes noticeable performance degradation, while the proposed method consistently improves reconstruction quality over digital semantic baselines. In particular, QNP alone already provides stable gains over the SOM-only receiver, indicating that its effectiveness does not rely on diffusion recovery itself. Moreover, VAE+SOM+QNP achieves performance close to VAE+SOM+Diff while requiring much lower computational cost, and combining QNP with diffusion yields the best overall performance. Second, two training strategies, namely independent QNP training and diffusion-assisted fine-tuning, are compared. The results show that diffusion-assisted fine-tuning provides only limited additional gains but significantly increases training cost and complexity, so independent training offers a more practical balance. Third, experiments under different SOM configurations and different SNR conditions verify that QNP provides stable gains across different modulation orders and SOM layer settings. Latency analysis further shows that QNP introduces only a small fixed overhead, whereas the diffusion module dominates the total inference time; therefore, the adaptive diffusion-step schedule is selected according to the measured latency-PSNR trade-off. Fourth, Gaussianity analysis based on the Kullback-Leibler divergence and Wasserstein distance shows that QNP compensation significantly improves the Gaussianity of the residual noise, while the version with distribution regularization achieves the best statistical consistency.  Conclusions  This paper proposes a channel-adaptive receiver for digital semantic communication, in which QNP-based front-end compensation is combined with diffusion-based semantic recovery. The main contribution lies in introducing a lightweight and independently effective QNP module to compensate SOM-induced structured quantization distortion before subsequent recovery. Experimental results show that QNP alone can already stably improve digital semantic reconstruction under different SNR conditions and different SOM configurations, while its combination with diffusion recovery yields the best overall performance. Therefore, the proposed method provides an effective way to improve semantic reconstruction quality and channel adaptability while preserving compatibility with existing digital communication infrastructures.
Resource Allocation Optimization in Dual-RIS Cooperative Rate-Splitting Multiple Access Networks
CHEN Yuang, WU Chang, PENG Mingyu, LU Hancheng
Available online  , doi: 10.11999/JEIT260171
Abstract:
  Objective  In Rate-Splitting Multiple Access (RSMA) systems, the achievable common-stream rate is limited by the user with the weakest channel quality. This constraint reduces scalability, robustness, and user fairness in dense 6G networks. Existing cooperative RSMA architectures partly reduce this bottleneck, but they remain constrained by fixed channel conditions and limited interference management. To address these issues, this paper proposes a dual Reconfigurable Intelligent Surface (RIS) cooperative RSMA system. Two cooperatively deployed RISs create additional controllable propagation paths through cascaded double reflection. The objective is to maximize the system sum rate by jointly optimizing Base Station (BS) Beamforming (BF), Rate Splitting (RS) strategies, and dual-RIS phase configurations, thereby improving spectral efficiency, robustness, and user fairness under users’ Quality of Service (QoS) constraints.  Methods  A tractable system model is developed for the dual-RIS cooperative RSMA system. The model captures cascaded multi-link channels, multi-node channel structures, and interference coupling. Based on this model, a joint optimization problem is formulated to maximize the system sum rate by optimizing BS BF, RS strategies, and the discrete phase shifts of both RISs. Because of strong variable coupling and non-convexity, a low-complexity Alternating Optimization (AO) algorithm is designed. The original problem is decomposed into three subproblems: BS-side RIS phase optimization, user-side RIS phase optimization, and BS BF optimization. Semidefinite Relaxation (SDR) and Successive Convex Approximation (SCA) are used to transform these subproblems into tractable convex forms, which are then solved iteratively with fast convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed dual-RIS cooperative RSMA system. The proposed AO algorithm converges within six iterations under different numbers of RIS reflecting elements, and it reaches 97.8% of the steady-state sum rate within three iterations when M = 170 (Fig. 3). Compared with SOPS and RPS, the proposed phase-configuration scheme obtains 10.6% and 31.8% sum-rate gains when M = 190, respectively (Fig. 4). The proposed RSMA scheme also outperforms NOMA and SDMA by 10.0% and 14.6%, respectively (Fig. 5). Under M = 160 and b = 4, dual-RIS cooperation provides an 11.9% sum-rate gain over the single-RIS scheme, and its performance is close to the CPS upper bound (Fig. 6). Balanced allocation of reflecting elements between the two RISs further improves the sum rate (Fig. 7). The proposed BF strategy also outperforms ZF and RBF, achieving 33.2% and 336.5% gains at a transmit power of 30 dBm, respectively (Fig. 8). Under different RIS cooperation modes, the proposed joint optimization scheme achieves the best overall performance (Fig. 9). These results show that dual-RIS cooperative RSMA improves common-stream decoding, interference suppression, robustness, and user fairness.  Conclusions  This paper investigates a dual-RIS cooperative RSMA communication system. The proposed architecture improves common-stream decoding while mitigating complex interference. To maximize the system sum rate, BS BF vectors, RS vectors, and the discrete phase matrices of two RISs are jointly optimized. A low-complexity AO algorithm based on SDR and SCA is developed to solve the strongly coupled non-convex problem. Simulation results show that the proposed dual-RIS cooperative RSMA scheme achieves clear sum-rate gains over advanced benchmark schemes. Compared with the single-RIS mode and SDMA, it obtains 11.9% and 14.6% rate gains, respectively, while improving system robustness and user fairness.
Intelligent Resource Allocation Algorithm Based on Outdated CSI for Multi-Node URLLC
ZHAO Yizhen, GAO Wei, HU Yulin, ZHU Yao
Available online  , doi: 10.11999/JEIT260216
Abstract:
  Objective  Ultra-Reliable and Low-Latency Communications (URLLC) is widely used in Industrial Internet of Things (IIoT) systems. However, in mobile industrial scenarios such as transportation and inspection, instantaneous Channel State Information (CSI) is difficult to obtain because of feedback overhead. Resource allocation decisions therefore need to be made using outdated CSI. This mismatch restricts system energy efficiency. Traditional convex optimization methods have difficulty addressing this problem. Classical Deep Reinforcement Learning (DRL) algorithms also have limited convergence stability and policy performance under the stringent latency and reliability constraints of URLLC. To address these challenges, this paper considers a multi-node URLLC system under outdated CSI in dynamic scenarios. An energy-efficiency maximization problem is formulated under the Finite BlockLength (FBL) regime, with communication latency and reliability constraints. An efficient and stable algorithm is then designed for joint power and blocklength allocation.  Methods  A Successive Convex Approximation (SCA)-assisted DRL framework is proposed to maximize energy efficiency under outdated CSI. First, an SCA-based algorithm is developed to obtain a pre-allocation solution for transmit power and blocklength. This solution is feasible and physically interpretable, but relatively conservative. Based on this baseline, a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is used for incremental refinement through interaction with the dynamic environment. This process reduces the conservatism of SCA. The SCA solution is used as prior knowledge in the state representation. Node location information is also incorporated into the state space. These designs narrow the policy search space and enable the DRL agent to better capture large-scale channel characteristics and system dynamics under outdated CSI. Learning efficiency and stability are therefore improved.  Results and Discussions  The proposed algorithm is evaluated through simulations and compared with three benchmark algorithms: an SCA-based optimization algorithm, a TD3 algorithm without SCA guidance, and a TD3 algorithm without node location information. The results show that the proposed method outperforms all benchmarks in convergence stability and system energy efficiency. In the training phase (Fig. 3), the average reward of the proposed algorithm increases steadily and converges stably. By contrast, removing node location information leads to lower rewards and stronger fluctuations. Removing SCA guidance causes the algorithm to converge to a much lower reward level. These results confirm the roles of SCA-based prior guidance and location-aware state representation in improving training stability. In the actual operation stage (Fig. 4), the proposed algorithm achieves high and stable energy efficiency and outperforms all comparison algorithms. Under outdated CSI, DRL-based methods can obtain higher energy efficiency than conservative optimization methods when transmission succeeds. However, removing node location information reduces energy efficiency, and removing SCA guidance increases transmission failures. These results verify the effectiveness of both designs in improving energy efficiency and maintaining policy feasibility. The effects of key system parameters are also examined. For basic resource parameters, a moderate increase in the blocklength budget (Fig. 5) or power budget (Fig. 6) improves system energy efficiency. For reliability constraints (Fig. 7), the reliability requirement should be set according to service requirements to avoid resource waste. Finally, the average energy efficiency under different numbers of nodes and different numbers of neurons in the TD3 network is analyzed (Fig. 8). The results provide guidance for algorithm configuration and network-scale design.  Conclusions  This paper addresses energy-efficient resource allocation for multi-node URLLC systems with outdated CSI by integrating SCA and DRL. In the proposed framework, a TD3-based DRL algorithm is guided by an SCA reference solution, and node location information is incorporated into the state representation. This optimization-learning dual-driven framework combines the interpretability and feasibility of model-based optimization with the adaptivity of data-driven learning. Simulation results show that the proposed method achieves higher energy efficiency than SCA-based optimization and conventional TD3 while satisfying URLLC latency and reliability constraints. The SCA reference solution improves policy stability and effectiveness under outdated CSI. Node location information further supports efficient decision-making. This work focuses on a single-cell multi-node scenario under Time Division Multiple Access (TDMA). Practical issues such as multi-cell interference, cooperative scheduling among multiple base stations, and more complex mobility patterns are not considered. Future work will extend the proposed framework to multi-cell and multi-agent scenarios and test its applicability under more severe CSI imperfections.
Load Optimization of Inverter Air-Conditioning Clusters Driven by Constraint Surface Projection and Spatial-Fitness Synergy
ZHENG Bowen, PAN Mingming, WANG Lei, LIU Chang, ZHENG Qingrong, TANG Zhuofan, ZHAO Jianli
Available online  , doi: 10.11999/JEIT260149
Abstract:
  Objective  Supply-demand imbalances in modern distribution networks are intensified by the increasing penetration of distributed renewable energy and frequent extreme high-temperature events. Large-scale Inverter Air-Conditioning (IAC) clusters can be aggregated as virtual energy storage resources for Demand Response (DR), providing an effective way to improve grid flexibility. However, existing dispatch strategies are often limited by the curse of dimensionality. Conventional penalty-function-based soft constraints also fail to strictly satisfy aggregate power equality constraints and may introduce steady-state errors. This paper develops an optimization framework in which grid-side power commands are accurately tracked while user thermal discomfort is reduced and fairness among heterogeneous users is maintained.  Methods  A multi-objective optimization framework based on an Equivalent Thermal Parameter (ETP) model is established to describe the thermodynamic states of heterogeneous buildings. To balance collective comfort and individual fairness, a composite fitness function is designed by integrating a weighted mean-squared error term, a fairness variance term, and a maximum violation suppression term. To eliminate the steady-state errors of traditional penalty-based methods, a Spatial-Fitness Adaptive Particle Swarm Optimization (SFA-PSO) algorithm is proposed. A geometric constraint surface projection mechanism maps particles strictly onto the power-conservation hyperplane, thereby satisfying the aggregate power equality constraint. In addition, the learning factors are dynamically adjusted through a Spatial-Fitness Adaptive (SFA) strategy. This strategy measures the mismatch between a particle’s fitness rank and spatial distance rank, which helps prevent premature convergence in high-dimensional search spaces.  Results and Discussions  Extensive continuous scheduling simulations are conducted in a complex dynamic environment. The environment includes multi-source thermal disturbances, a bidirectional communication packet loss rate of 1%, and Part Load Ratio (PLR) values of 20%, 50%, and 80%. First, ablation experiments confirm that constraint surface projection guarantees power tracking accuracy. Traditional penalty-based methods, such as Penalty Particle Swarm Optimization (Penalty-PSO), produce steady-state power deviations of approximately 10^-1 kW. By contrast, SFA-PSO limits aggregate power tracking errors to within 10^-9 kW (Fig. 3). The SFA strategy also prevents the premature convergence observed in Physical Particle Swarm Optimization (Phy-PSO). It enables continuous fitness reduction, especially in low-load scenarios with narrow feasible regions (Fig. 4). This improvement is attributed to the dynamic evolution of the learning factors. The cognitive factor remains high at the early stage to promote global exploration. It then decreases as the social factor increases, which strengthens local exploitation and improves convergence precision (Fig. 5). Second, continuous dynamic scheduling performance is evaluated through a 6-hour simulation during the peak load period from 12:00 to 18:00. The dispatch interval is 5 min, yielding 72 decision steps. Under tight peak-load constraints, Genetic Algorithm (GA) and Whale Optimization Algorithm (WOA) show severe power-limit violations because their population update rules do not cooperate well with the projection mechanism. By contrast, SFA-PSO maintains strict constraint satisfaction (Fig. 7). SFA-PSO remains at the lowest fitness level throughout the real-time evolution curves, indicating strong robustness against environmental thermal noise and uplink and downlink communication packet loss (Fig. 8). Quantitatively, compared with eight baseline algorithms, including Social Learning Particle Swarm Optimization (SLPSO), Competitive Swarm Optimizer (CSO), and Dynamic State Cluster-Based Particle Swarm Optimization (DSCPSO), SFA-PSO achieves the best overall performance. It obtains an average fitness of 904, a minimum fitness of 243, and the lowest standard deviation of 551 (Table 2). Finally, scalability analyses across cluster sizes from 100 to 1,000 nodes further validate the high-dimensional optimization capability of SFA-PSO. In all scale scenarios, SFA-PSO shows the strongest optimization capacity. It achieves rapid initial descent within the first 20 iterations and maintains continuous exploration in later stages (Fig. 9). Although the projection and SFA mechanisms increase computational time by 30% to 50% compared with basic Particle Swarm Optimization (PSO) (Fig. 6), the absolute optimization time remains stable at approximately 1.5 seconds even for a 1,000-node cluster (Fig. 9). This computational overhead is acceptable for minute-level control cycles and meets the real-time dispatch requirements of modern smart grids.  Conclusions  The proposed SFA-PSO algorithm effectively addresses the steady-state error of traditional soft-constraint methods in aggregate power control. By ensuring accurate tracking of dispatch commands and mitigating high-dimensional search traps, it provides a robust and scalable solution for flexible scheduling of large-scale IAC loads in smart grids. It also maintains a practical balance between grid-side regulation and user-side comfort. The method still has limitations. The constraint projection mechanism depends on the host algorithm, which restricts cross-algorithm generalization. High-precision tracking also increases computational cost. Future work will focus on adaptive constraint handling and lightweight algorithm design. Coordinated scheduling for heterogeneous loads, such as electric vehicles and energy storage, will also be investigated.
A Cryptographic Side-Channel Security Modeling and Formal Verification Method
WANG Xingxin, HU Wei, HUANG Xuan, LIN Chenyu, ZHOU Yi
Available online  , doi: 10.11999/JEIT260631
Abstract:
  Objective  Compared with post-silicon side-channel security analysis, pre-silicon side-channel security verification during the design phase enables the earlier identification of potential side-channel security vulnerabilities in cryptographic core designs, thereby effectively reducing the cost and time of post-silicon remediation. However, most existing pre-silicon side-channel security assessment approaches rely on data-driven statistical analysis or artificial intelligence techniques and require complex calculations on large amounts of data to mitigate the impact of insufficient coverage on the assessment results. In addition, existing methods typically adopt independent modeling strategies for different types of side channels, lacking a unified side-channel security modeling approach. A cryptographic side-channel security modeling and formal verification method is proposed, supporting unified and automated modeling of different types of side channels by constructing a side-channel security model. The method can identify potential timing side-channel, power side-channel and fault injection vulnerabilities in cryptographic core designs, and analyze the effectiveness of side-channel countermeasures based on side-channel security property checking.  Methods  The proposed cryptographic side-channel security modeling and formal verification method includes side-channel security model construction, side-channel security property extraction, and side-channel security verification. The side-channel security model uses information flow analysis to characterize timing side-channel leakage, power side-channel leakage, and fault propagation behavior in cryptographic core designs, providing an effective mathematical model for side-channel security verification. Specifically, the side-channel security model utilizes changes in signal labels to analyze information flows during the encryption process by assigning a label to a signal bit and defining label propagation rules. Side-channel security properties formally describe the behavioral characteristics of side-channel leakage, including timing properties, power properties, and fault properties, providing theoretical support for side-channel security verification. Side-channel security verification uses the extracted security properties as verification constraints and employs formal verification tools to identify potential timing side-channel, power side-channel, and fault injection vulnerabilities in cryptographic core designs. Furthermore, the method can analyze the effectiveness of masking and fault injection countermeasures against side-channel vulnerabilities.  Results and Discussions  The proposed side-channel security verification method utilizes formal verification techniques to accurately identify potential side-channel security vulnerabilities in various block cipher core designs, and evaluate the effectiveness of side-channel countermeasures based on side-channel security property constraints. The timing side-channel verification results demonstrate that the proposed method can accurately identify timing side-channel security vulnerabilities in AES, SM4, LED, PRESENT, and IDEA core designs within 20s (Table 2, Fig. 7). No timing side-channel vulnerabilities are identified in the other cryptographic core designs, except for IDEA, which exhibits timing side-channel vulnerabilities caused by modular multiplication operations. The power side-channel security verification results show that formal checks based on controllability property, key–power distinguishability coupling property and key–power nonlinear coupling property can accurately identify target modules with potential power side-channel security vulnerabilities in AES, SM4, LED and PRESENT within 1 minute (Table 3). The key expansion module in cryptographic core designs does not cause key leakage through key-dependent power consumption, as it fails to satisfy the controllability property. In addition, the experimental results indicate that the masking protection in the RSM core design can prevent the correct key from being distinguished through random masking (Fig. 8). The fault injection security verification results for three AES core designs with infective countermeasures demonstrate that the proposed method can analyze the effectiveness of fault infection countermeasures. The results show that the infection countermeasure requires not only altering the fault propagation path but also disrupting the algebraic relationships among faults (Table 4, Fig. 10).  Conclusions  This paper proposes a cryptographic side-channel security modeling and formal verification method to address the lack of formal mathematical models and the limited completeness of existing data-driven side-channel security assessment methods. The proposed method first achieves unified modeling of timing leakage, power leakage, and fault propagation behaviors from the perspective of information flow analysis. Based on the constructed side-channel security model, side-channel security properties are extracted to formally characterize the behavioral features of side-channel information leakage and propagation during the encryption process. Potential side-channel security vulnerabilities in cryptographic core designs are then identified through formal checking using the extracted security properties as constraints. The proposed method provides an effective solution for the unified modeling and formal security verification of different types of side channels. Experimental results obtained from the side-channel security verification of various block cryptographic core designs demonstrate that: (1) the proposed method can uniformly model timing side-channel leakage, power side-channel leakage, and fault propagation behaviors in cryptographic core designs; (2) the proposed method can accurately identify timing side-channel vulnerabilities, power side-channel vulnerabilities, and fault injection vulnerabilities in cryptographic core designs, including AES, SM4, IDEA, LED and PRESENT; (3) the proposed method can analyze the effectiveness of masking and fault infection countermeasures. However, this study only qualitatively identifies side-channel security vulnerabilities in cryptographic core designs; pre-silicon quantitative assessment of side-channel leakage should be investigated in future work.
A Lightweight Dual-Stream Convolutional Network Feature Fusion Method for UAV RF Recognition
DONG Pengyu, XIANG Xin, LV Siting, LIANG Yuan, WANG Rui, MAO Hu
Available online  , doi: 10.11999/JEIT260464
Abstract:
  Objective  With the rapid proliferation of Unmanned Aerial Vehicles (UAVs) and the escalating demand for airspace security, radio frequency (RF) fingerprint recognition has emerged as a pivotal technology for identifying non-cooperative UAVs. However, existing methods grapple with significant challenges, including poor robustness in low signal-to-noise ratio (SNR) environments and prohibitive computational complexity, which severely hinder their deployment on resource-constrained tactical edge devices. To address these critical limitations, this paper proposes a novel lightweight dual-stream convolutional network tailored for UAV RF recognition. This network is designed to extract static spectral texture features and dynamic temporal gradient features in parallel, complemented by a meticulously crafted lightweight feature fusion strategy.  Methods  The proposed network architecture is ingeniously designed to process RF signals. The input signal undergoes a Short-Time Fourier Transform (STFT) to generate a two-dimensional spectrogram, which serves as the primary input. The network is bifurcated into two parallel streams: a static stream and a dynamic stream. The static stream is engineered to capture the inherent static spectral patterns and energy distributions within the STFT spectrogram. It comprises a series of stacked convolutional blocks, each integrating convolutional layers, batch normalization, and ReLU activation functions, followed by max-pooling layers to progressively downsample the feature maps and increase the channel depth. Conversely, the dynamic stream is dedicated to enhancing feature discriminability, particularly in low-SNR scenarios. It begins by computing the temporal gradient of the input spectrogram, effectively suppressing static background noise and accentuating dynamic signal variations. This gradient map is then processed by a symmetric set of convolutional blocks, mirroring the structure of the static stream. To maintain model efficiency, an element-wise addition fusion strategy is employed to integrate the features from both streams, ensuring a balance between feature complementarity and computational overhead. The fused features are subsequently fed into a classification head, consisting of an adaptive average pooling layer, dropout layers for regularization, and fully connected layers to produce the final classification output. Extensive experiments are conducted on the publicly available DroneRF dataset, encompassing ablation studies to dissect the contribution of each component, comparative analyses of various fusion strategies, and rigorous evaluations of the model’s lightweight characteristics.  Results and Discussions  The experimental results unequivocally demonstrate the efficacy of the proposed method. The dual-stream network achieves a remarkable 95.65% accuracy on the test set, representing a substantial 5 percentage point improvement over the best-performing single-stream network. A critical analysis reveals that the temporal gradient operation contributes significantly to this enhancement by improving the average SNR by 1 dB, thereby bolstering feature discriminability in challenging low-SNR environments. Furthermore, the model’s lightweight design is a standout feature, with a mere 0.58 million parameters, making it eminently suitable for deployment on tactical edge devices. Ablation studies and feature visualization analyses provide compelling evidence for the complementary nature of static and dynamic features. The static stream adeptly captures broad spectral contours, while the dynamic stream focuses on fine-grained temporal variations. The element-wise addition fusion strategy proves superior, outperforming other approaches like feature concatenation and attention-based fusion in terms of both performance and computational efficiency, thereby validating the rationale behind the lightweight design.  Conclusions  This paper presents a comprehensive solution to the challenges of UAV RF recognition in complex environments by proposing a lightweight dual-stream convolutional network. The method effectively enhances recognition accuracy and robustness through the synergistic combination of dual-stream feature extraction and the SNR-enhancing properties of temporal gradient features, all while maintaining a lightweight architecture suitable for edge deployment. The proposed approach offers a significant advancement in the field, providing a robust and efficient solution for UAV identification. Future research endeavors will focus on further enhancing the model’s adaptability to complex electromagnetic environments, incorporating the effects of sensor noise, and extending the framework to multi-UAV cooperative scenarios.
Complex-domain Joint Spectrum Sensing Method for UAV Swarms in Complex Electromagnetic Environments
QIAN Hui, CHEN Li, YIN Huarui, WANG Weidong
Available online  , doi: 10.11999/JEIT260499
Abstract:
  Objective  The rapid development of the low-altitude economy is increasing the use of unmanned aerial vehicle (UAV) swarms in emergency communication, urban logistics, reconnaissance, and low-altitude network coverage. These applications require reliable spectrum awareness to support cooperative communication, dynamic spectrum access, and interference avoidance. However, low-altitude electromagnetic environments often contain low-SNR signals, multipath propagation, non-cooperative interference, and multiple coexisting transmissions. Conventional energy, cyclostationary-feature, and matched-filter detectors are sensitive to noise uncertainty, computational cost, or prior waveform knowledge. Learning-based methods can improve robustness, but many rely on power spectral density (PSD) or short-time Fourier transform (STFT) representations. These representations may weaken phase information or require costly two-dimensional time-frequency preprocessing. Existing methods also focus mainly on spectrum occupancy and provide limited information about overlapping transmissions. This study therefore develops a low-latency joint sensing method that preserves magnitude and phase information while estimating spectrum occupancy and interference overlap for each frequency bin. The method is intended for local spectrum sensing at UAV nodes under resource and latency constraints.  Methods  The proposed RadioSEUnet pipeline contains two stages: magnitude-phase feature construction and joint spectrum-state estimation. First, each complex baseband in-phase/quadrature (I/Q) sequence is multiplied by a Hann window and transformed using a one-dimensional fast Fourier transform (FFT). The resulting complex spectrum is decomposed into logarithmic magnitude and phase components. The two components are normalized separately and stacked as a two-channel feature tensor. Binary labels indicate spectrum occupancy and multi-signal overlap at each frequency bin. RadioSEUnet adopts a U-shaped encoder-bottleneck-decoder architecture with four encoder stages containing 64, 128, 256, and 512 channels. Each RadioSEBlock combines a complex-parameterized convolution, squeeze-and-excitation channel attention, and a residual connection. The convolution couples the magnitude and phase feature streams through constrained cross-channel operations. Two prediction heads convert the shared representation into a spectrum-occupancy probability mask and an interference-overlap probability mask. The model is optimized using an equally weighted sum of two binary cross-entropy losses. Training uses AdamW, cosine-annealing learning-rate scheduling, early stopping, a batch size of 64, and at most 200 epochs. The complete data collection contains 72,000 complex I/Q records, including 48,000 simulated records and 24,000 measured records. The simulated subset covers Wi-Fi, BLE, ZigBee, LoRa, QPSK/16QAM, FM, and AM signals. Signal-to-noise ratios range from –15 dB to 10 dB under additive white Gaussian noise and Rayleigh fading. The measured subset was collected using a USRP N310 in the 2.4–2.5 GHz ISM band at 100 MS/s over a 1 ms observation interval. The controlled quantitative evaluation uses an 8:1:1 split of the simulated subset. A separate simulated-to-measured protocol is defined in the main text to examine cross-domain generalization. RadioSEUnet is compared with six PSD- or STFT-based baselines under matched data splits and hardware conditions. Performance is measured using intersection over union (IoU), precision, recall, preprocessing time, inference time, and total sensing latency.  Results and Discussions  The SNR-dependent quantitative results reported here are obtained using the controlled simulated-data protocol. At -15 dB, RadioSEUnet achieves an IoU of 0.768 and a recall of 0.846 for spectrum occupancy detection. Compared with the second-best STFT-RADN baseline, these values correspond to absolute improvements of 0.186 and 0.166, respectively. For interference-overlap detection, RadioSEUnet achieves an IoU of 0.456 and a precision of 0.768 at –15 dB. The corresponding improvements over STFT-RADN are 0.246 and 0.275. The lower IoU for interference-overlap detection indicates that weak overlap boundaries remain difficult to separate from strong-signal sidelobes and background noise. The latency evaluation is conducted on the workstation specified in the main text. Magnitude-phase preprocessing requires 21.04 ms, and network inference requires 4.77 ms, producing a total sensing latency of approximately 25.81 ms. STFT-YOLOv3 requires 98.9 ms under the same hardware setting, so the proposed pipeline is approximately 3.8 times faster in this comparison. Ablation experiments show that magnitude-phase preprocessing, complex-parameterized feature coupling, and channel attention each improve low-SNR sensing performance. Removing the magnitude-phase preprocessing produces the largest degradation. These results indicate that preserving complementary magnitude and phase information is useful for weak-signal and interference-overlap detection. They do not, however, establish performance on airborne hardware or across unreported radio environments.  Conclusions  RadioSEUnet combines a magnitude-phase representation, constrained cross-channel feature coupling, channel attention, multiscale feature fusion, and dual-head prediction. It jointly estimates spectrum occupancy and interference-overlap states while avoiding two-dimensional STFT preprocessing. Under the controlled simulated-data protocol, the method provides higher point estimates than the six evaluated baselines at low SNR and reduces total sensing latency on the evaluated workstation. The present evidence is limited to the reported signal types, channel models, hardware configuration, and the 2.4–2.5 GHz measurement band. Quantitative simulated-to-measured results, tests on wider bands, additional interference types, repeated trials, and deployment on airborne edge hardware are still required. Future work will therefore focus on cross-domain validation, lightweight deployment, boundary-aware interference modeling, and integration with spectrum resource management for UAV networks.
Decision Learning Correction Network: Fusion Classification of Hyperspectral Images and LiDAR Data
WANG Haoyu, LIU Nuofei, CHENG Yuhu, LIU Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260362
Abstract:
  Objective  HyperSpectral Images (HSI) and Light Detection And Ranging (LiDAR) provide complementary information for land-cover classification. HSI captures rich spectral information for material discrimination, while LiDAR provides elevation and structural information for spatial characterization. However, most existing fusion methods treat multimodal fusion as a one-shot static aggregation process, implicitly assuming that a fixed fusion strategy is applicable to all pixels and regions. This assumption is difficult to satisfy in complex remote sensing scenes, where class-boundary and cross-modal heterogeneous regions exhibit high information density but account for only a small proportion of samples (Fig. 1). To address this limitation, this paper proposes a Decision Learning Correction Network (DLCN) that reformulates static HSI-LiDAR fusion as a context-dependent sequential decision-making process.  Methods  The proposed DLCN consists of feature extraction, fusion decision learning, and classification. First, HSI and LiDAR are processed through two parallel branches to extract spectral and spatial features and elevation and structural features, respectively. The extracted features are then concatenated to form the current state and are fed into an Actor-Critic framework. The Actor network generates fusion actions to adaptively adjust modality contributions, while the Critic network evaluates the long-term value of each action for classification. To improve learning from difficult samples, a key-sample-oriented sampling module assigns higher sampling probabilities to samples with larger modal fidelity loss. Meanwhile, a modal fidelity constraint mechanism evaluates spectral fidelity, feature consistency, structural preservation, and resolution matching, and corrects destructive actions during fusion. Through this closed-loop framework, DLCN performs dynamic generation, evaluation, and correction of fusion actions, thereby producing high-quality fusion features for classification (Fig. 2).  Results and Discussions  Experiments are conducted on the Houston2013, Trento, and MUUFL datasets. DLCN achieves the highest Overall Accuracy (OA) of 97.85%, 99.58%, and 94.38% on the three datasets, respectively, outperforming CHNet, DSymFuser, mPMCL, MEDFN, S3F2Net, and MSAF. The classification maps demonstrate that DLCN effectively reduces misclassification in class-boundary, mixed land-cover, and structurally complex regions, producing results that more closely match the ground-truth maps across all three datasets (Figs. 35). Ablation studies further demonstrate that the value-guided policy optimization mechanism, key-sample-oriented sampling module, and modal fidelity constraint mechanism each improve classification performance. Compared with the baseline models, the complete DLCN consistently increases OA on Houston2013, Trento, and MUUFL, validating the effectiveness of the proposed decision-learning-correction framework. Time-step analysis shows that DLCN progressively improves classification accuracy while maintaining stable spectral-angle variation during sequential decision making (Fig. 6). Furthermore, DLCN achieves inference times of 1.32 s, 0.86 s, and 2.23 s on the three datasets, respectively, ranking first among the compared methods. These results indicate that the additional computation introduced by the Actor-Critic decision framework and modal fidelity constraint mechanism is effectively translated into improved classification performance without imposing excessive computational cost.  Conclusions  This paper proposes a DLCN for HSI and LiDAR fusion classification. Unlike conventional static fusion methods, DLCN formulates multimodal fusion as a sequential decision-making process and adaptively adjusts fusion strategies according to the local context. Its closed-loop framework enables fusion actions to be generated, evaluated, and corrected throughout the decision process, thereby producing high-quality fusion features for classification. Experimental results demonstrate that DLCN produces more accurate classification maps in heterogeneous remote sensing scenes, and the time-step analysis further confirms the stability of the sequential decision-making process. Future work will focus on more fine-grained feature representation and more robust policy optimization to improve model generalization in complex remote sensing scenes.
Radiation-Hardened Ga2O3 MOSFET Design Featuring NiO Heterojunction and Comb-Shaped Gate Modulation
GAO Sheng, ZHANG Lin, WU Yanjun, WANG Qi, JING Liang
Available online  , doi: 10.11999/JEIT260396
Abstract:
  Objective  Gallium Oxide Metal-Oxide-Semiconductor Field-Effect Transistor (Ga2O3 MOSFET) is regarded as a promising power device for high-voltage applications, particularly in aerospace and satellite power systems, because of its ultra-wide bandgap and high critical breakdown field. However, the Conventional MOSFET (C-MOSFET) exhibits limited reliability in space radiation environments. Under off-state conditions, the electric field is highly concentrated near the gate edge. Heavy-ion irradiation generates dense electron-hole pairs along the ion track. Driven by the intense electric field, these carriers undergo avalanche multiplication through impact ionization, causing the drain current to increase sharply without recovery and ultimately leading to irreversible Single-Event Burnout (SEB) at relatively low drain bias. This failure mechanism severely limits the application of Ga2O3 MOSFETs in harsh radiation environments. Furthermore, the lack of reliable and efficient p-type doping restricts the implementation of conventional radiation-hardening techniques, including junction termination extension and junction isolation. Therefore, ionization-induced carriers readily accumulate in sensitive regions, increasing susceptibility to Single-Event Effect (SEE). The extremely low thermal conductivity of Ga2O3 further promotes local heat accumulation following heavy-ion irradiation, producing localized hot spots that increase the likelihood of thermal burnout. Existing hardening approaches, including field-plate optimization and dielectric engineering, provide only limited improvement. Moreover, the application of heterojunction structures for radiation hardening has rarely been investigated, and systematic hardening strategies have not yet been established. To address these limitations, this paper proposes a Comb-Shaped Gate Metal-Oxide-Semiconductor Field-Effect Transistor (CSG-MOSFET) incorporating a NiO heterojunction. The proposed structure redistributes the channel electric field, suppresses electric-field crowding at the conventional gate edge, and significantly improves SEB tolerance, providing an effective solution for Ga2O3 power devices operating in harsh radiation environments.  Methods  Technology Computer-Aided Design (TCAD) simulations are performed to evaluate the electrical characteristics and SEB performance of the proposed CSG-MOSFET in comparison with the C-MOSFET. The simulations incorporate high-field mobility, Shockley-Read-Hall recombination, Auger recombination, impact ionization, and heavy-ion models. Based on the charge-compensation effect of the p-NiO/n-Ga2O3 heterojunction, the proposed structure utilizes the extended depletion region formed at the heterointerface to redistribute the channel electric field. This heterojunction-induced depletion region improves electric-field uniformity and enhances SEB tolerance. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field away from the conventional gate edge, suppress local electric-field crowding, and improve device reliability under high-voltage and radiation conditions.  Results and Discussions  Simulation results demonstrate that the optimized Double Comb-Shaped Gate MOSFET (DCSG-MOSFET) significantly improves radiation hardness compared with the C-MOSFET. The SEB Threshold Voltage (VSEB) increases from 240 V to 2 280 V, while the Breakdown Voltage (BV) increases from 2 000 V to 3 500 V. Meanwhile, the specific on-resistance decreases. Therefore, the Baliga Figure of Merit (BFOM) and the SEB-based figure of merit are substantially improved. The NiO heterojunction and comb-shaped gate columns effectively redistribute the electric field, shifting the peak electric field from the conventional gate edge to the outer gate-column edge and suppressing local electric-field crowding. These improvements substantially enhance the radiation hardness of the device.  Conclusions  A radiation-hardened DCSG-MOSFET incorporating a NiO heterojunction is proposed and evaluated using TCAD simulations. The optimized structure significantly improves SEB tolerance while maintaining excellent electrical performance. Compared with the C-MOSFET, both VSEB and BV are substantially increased, demonstrating enhanced blocking capability. Charge compensation at the p-NiO/n-Ga2O3 heterojunction forms an extended depletion region that effectively redistributes the channel electric field and suppresses electric-field crowding near the conventional gate edge. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field to the outermost gate-column edge, thereby suppressing impact ionization induced by heavy-ion irradiation and effectively mitigating SEB. The reduced specific on-resistance further improves the BFOM and the SEB-based figure of merit. These results demonstrate that the proposed DCSG-MOSFET is a promising candidate for power electronic applications in harsh radiation environments, including aerospace and satellite systems.
A Multi-Station Emitter TDOA Deinterleaving Method for Severe Pulse-Loss Environments
LIU Yuchen, ZHAO Yaqin, WU Longwen
Available online  , doi: 10.11999/JEIT260401
Abstract:
  Objective  Modern electronic reconnaissance systems must deinterleave dense and overlapping radar pulse streams in non-cooperative environments. As radar emitters increasingly employ agile waveforms, similar pulse descriptor words, and low-intercept-probability strategies, conventional single-station methods based on carrier frequency, pulse width, and Pulse Repetition Interval (PRI) become less reliable. Multi-station deinterleaving based on Time Difference of Arrival (TDOA) provides a more stable geometric observable, but severe pulse loss still causes sparse cross-station pairing, weak true TDOA peaks, ambiguity-induced spurious peaks, isolated pulses, and fragmented trajectories across time slices. These effects increase false alarms and weaken track continuity. To address these issues, a closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments, with Time of Arrival (TOA) sequences used as the core observables.  Methods  A slice-based framework is developed for continuous reconnaissance. Residual unmatched pulses are carried forward by a sliding window to alleviate cross-slice misalignment. First, candidate pulse pairs satisfying geometric TDOA constraints are generated, and pulse descriptor word constraints on carrier frequency and pulse width are used to remove inconsistent pairs. To reduce the sparsity and binning sensitivity of conventional histograms, multiscale Kernel Density Estimation (KDE) is introduced to reconstruct the TDOA density from sparse candidate differences. Gaussian kernels with different bandwidths are fused, and candidate peaks are adaptively extracted using local statistics and peak widths. Second, a dynamic memory matrix is designed to suppress ambiguity-induced spurious peaks in high pulse repetition frequency scenarios. Since dependent spurious peaks collapse after the dominant peak is extracted and removed, a collapse-rate criterion is defined, and the spurious regions are recorded in a memory mask for subsequent iterations. Third, Dynamic Time Warping (DTW) is used to compare incomplete TOA sequences of isolated pulses with extracted pulse sequences, enabling reassignment of unequal-length and incomplete sequences. Finally, a Kalman-filter-based state-space model tracks multi-baseline TDOA trajectories across successive slices. Predicted and observed TDOA residuals are jointly used for association, so intermittent observations can still be linked to the correct track. In this way, the proposed method forms a closed-loop processing chain that links weak-peak reconstruction, spurious-peak suppression, isolated-pulse reassignment, and trajectory association (Fig. 3).  Results and Discussions  Four simulation scenarios are designed: a high pulse repetition frequency scenario dominated by ambiguity-induced spurious peaks, a parameter-overlapping scenario dominated by isolated pulse reassignment, an ablation scenario for evaluating the memory matrix and DTW modules, and a 10-emitter mixed-regime scenario including fixed PRI, staggered, jittered, frequency-agile, pulse-group frequency-agile, frequency-agile jittered-PRI, linear-sliding, and sinusoidal-sliding PRI signals. In the mixed-regime scenario, the total reconnaissance duration is 1 s and the slice duration is 0.1 s. The environmental pulse loss rate is fixed at 10%, and the receiver-specific loss rate increases from 0% to 40%. Both loss rates are calculated with respect to the initial theoretical number of transmitted pulses; therefore, the total loss rate is their sum, ranging from 10% to 50%. The proposed method is compared with an extended TDOA histogram method under constrained criteria, a cloud-model-based multi-station sorting method, and a Dirichlet Process Mixture Model (DPMM)-based method (Table 5). In the high pulse repetition frequency scenario, the proposed method maintains near-zero false alarms by identifying the collapse of dependent spurious peaks and suppressing them through the memory matrix, whereas the comparison methods show severe false alarms (Figs. 4 and 5). In the parameter-overlapping scenario, DTW-based reassignment improves isolated-pulse recovery, while the memory matrix suppresses spurious TDOA peaks. Their combination improves extraction reliability and reduces false alarms (Figs. 6 and 7). The ablation results verify their complementary roles: at a 50% loss rate, the memory matrix reduces the TDOA false alarm rate from 29.92% to 3.32%, DTW increases pulse extraction accuracy from 66.71% to 91.99%, and the complete method achieves a TDOA detection rate of 99.25% with a false alarm rate of 0.88% (Fig. 8). In the 10-emitter mixed-regime scenario, the proposed method achieves a favorable overall trade-off. At an overall pulse loss rate of 50%, its pulse extraction accuracy remains 92.72%, and the TDOA false alarm rate is limited to 7.75%, lower than 35.39%, 35.05%, and 37.58% for the DPMM, cloud-model, and constrained recursive histogram methods, respectively. After cross-slice trajectory association, the mean number of identity switches decreases to 3.83, compared with 10.64, 9.85, and 12.64 for the three comparison methods (Fig. 9 and Table 5).  Conclusions  A closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments. By integrating multiscale KDE-based weak peak reconstruction, dynamic memory-matrix-based spurious peak suppression, DTW-based isolated pulse reassignment, and Kalman-filter-based trajectory association, the method addresses the coupled failure mechanisms caused by severe pulse loss. Simulation results demonstrate high extraction accuracy, low TDOA false alarm rates, and strong trajectory continuity in high-loss and mixed-regime scenarios. These results demonstrate the effectiveness of the method under the simulated conditions and indicate its application potential for persistent multi-station passive reconnaissance.
LLM-Aided Secure Routing Method in Industrial IoT Against Flooding Attacks
LI Jieling, XIAO Liang, WANG Chengyao, FANG Mingyang, CHEN Chen, LEI Yan
Available online  , doi: 10.11999/JEIT260400
Abstract:
  Objective  Industrial Internet of Things (IIoT) routing forwards and schedules control commands, equipment status information and sensing data to support critical tasks such as collaborative equipment control, safe system operation and environmental monitoring, but the routing process is prone to congestion and resource exhaustion under flooding attacks. Existing intelligent secure routing methods apply reinforcement learning (RL) to optimize next-hop selection based on network topology, but the heterogeneity in queue capacity and link bandwidth of IIoT terminals is often overlooked, leading to load imbalance and local congestion, and limiting performance under high load or malicious traffic attacks. Therefore, we propose a large language model (LLM)-based global situation-aware assisted secure routing method in IIoT against flooding attacks, which applies RL to optimize multi-path selection and achieve load balancing across the network.  Methods  Based on global security awareness, queue congestion of neighboring nodes, queue capacity, link bandwidth, and service types, the proposed secure routing method applies RL to optimize multi-path selection against flooding attacks. The cloud–edge large model infers global security situational awareness including global load distribution and anomalous traffic distribution based on network topology, node resource occupancy and link state information, and feeds the inference result back to IIoT terminals to construct RL states and evaluate routing policies risks. In addition, a risk-aware function is formulated to quantify the routing disruption potential by integrating end-to-end latency, packet delivery ratio and node vulnerability to attacks. An experience replay buffer that incorporates both reward and risk is constructed, where both factors are considered during routing parameter updates to guide routing policy selection, thereby balancing safe path exploration and optimization efficiency.  Results and Discussions  Simulations are conducted using 30 industrial nodes under varying configurations, including bandwidths of 5 MHz, 10 MHz, 20 MHz, and queue capacities ranging from 100 to 500 packets. The global security situational awareness is inferred by the Qwen3.5-27B-AWQ-4bit, which is deployed on a cloud–edge server equipped with dual 24 GB RTX 4090 GPUs. In each time slot, each terminal sends 5 packets of 2 KB each to the industrial gateway. A flooding attacker injects \begin{document}$ y\in \{10,20,30\} $\end{document} packets into neighboring queues per time slot to excessively consume network resources. Compared with the baseline method EEMR, the proposed secure routing method improves 39.4% packet delivery ratio, reduces 48.2% end-to-end latency and 41.1% routing energy consumption. Compared with the baseline method RLMR, the proposed secure routing method improves packet delivery ratio by a factor of 1.48, reduces end-to-end latency by 53.8% and routing energy consumption by 54.5%. This is because the proposed method leverages an LLM to infer global security situational awareness, integrating load distribution and anomalous traffic patterns to assist in selecting low-load nodes while avoiding high-load nodes, potential attack nodes, and abnormal or faulty nodes.  Conclusions  This paper proposes an LLM-based global situation-aware assisted secure routing method for IIoT against flooding attacks, which applies RL to optimize multi-path selection based on global security situational awareness including load distribution and anomalous traffic distribution. A risk assessment network is constructed based on attack behavior characteristics and service requirements to evaluate the risk level of routing performance degradation, thereby enabling risk-aware rerouting. Simulation results show that the proposed method increases the packet delivery ratio by 39.4%, reduces the end-to-end latency by 48.2% and the routing energy consumption by 41.1%.
Accelerated Broadband Electromagnetic Scattering Analysis via ACA-Driven Measurement Matrix Interpolation
WANG Zhonggen, WU Chenggang, NIE Wenyan, SUN Yufa
Available online  , doi: 10.11999/JEIT260392
Abstract:
  Objective  Broadband electromagnetic scattering analysis is widely used in radar target recognition, stealth technology, and microwave imaging. Although the Method of Moments (MoM) provides high computational accuracy, it incurs substantial computational and memory costs for electrically large or geometrically complex targets because full impedance matrices must be constructed and solved. Existing acceleration techniques, including the MultiLevel Fast Multipole Method (MLFMM) and Adaptive Cross Approximation (ACA), reduce the computational burden but still require repeated matrix construction and equation solving at every frequency during wideband analysis. Methods such as Asymptotic Waveform Evaluation (AWE), Model-Based Parameter Estimation (MBPE), and impedance matrix interpolation have been proposed to reduce this redundancy. However, AWE is prone to error accumulation over wide frequency bands, MBPE requires expensive initial sampling, and conventional impedance matrix interpolation still requires the computation of full high-dimensional impedance matrices at the sampling frequencies. More recently, Compressive Sensing Method of Moments (CS-MoM) and its extension, CS-HBFM, have improved wideband analysis by employing Hyper-Basis Functions (HBFs). By constructing Characteristic Mode Basis Functions (CMBFs) only once at the highest frequency, CS-HBFM eliminates repeated basis-function generation. Nevertheless, existing CS-HBFM methods rely on nondeterministic random or uniform sampling, require expensive large-scale matrix-vector products, and repeatedly reconstruct and solve impedance equations throughout the frequency sweep.  Methods  A CS-ACA-MMI framework is proposed for broadband electromagnetic scattering analysis by combining dual ACA decomposition with Measurement Matrix Interpolation (MMI). First, CMBFs are constructed at the highest frequency, and dominant HBFs are selected according to the Modal Significance (MS) criterion. ACA is then applied to the full impedance matrix to extract deterministic row indices corresponding to the dominant Rao-Wilton-Glisson (RWG) basis functions. These indices are reused throughout the frequency band, eliminating nondeterministic sampling and repeated index extraction. Second, four sampling frequencies are selected using Chebyshev-Lobatto nodes. Low-dimensional measurement matrices are constructed directly from the extracted row indices, avoiding the generation of full high-dimensional impedance matrices. The measurement impedance elements at the sampling frequencies are corrected according to the geometric distance, interpolated to the target frequency, and then restored to the actual measurement impedance elements, thereby eliminating repeated construction of measurement matrices during frequency sweeping. Third, ACA is applied to the far-field component of the interpolated measurement matrix, converting large-scale matrix-vector products into low-dimensional matrix multiplications. The near-field sensing matrix is obtained directly by multiplying the measurement matrix by the basis functions, enabling rapid construction of the complete sensing matrix. Finally, the dense linear system is transformed into an overdetermined system under the compressive sensing framework, and the least-squares method is used to reconstruct the current coefficients, from which the broadband Radar Cross Section (RCS) is calculated. The Root Mean Square Error (RMSE) is used to evaluate numerical accuracy. Three representative targets, namely a cylinder, a slotted cone, and an almond, are analyzed. Broadband RCS, numerical accuracy, total computation time, and single-frequency measurement-matrix memory consumption are compared with those obtained using MoM and CS-HBFM to validate the proposed framework.  Results and Discussions  Three numerical examples, including a perfect electric conductor cylinder, a slotted cone, and an almond, are used to validate the proposed CS-ACA-MMI framework. The ACA-extracted row indices are concentrated near geometric boundaries and structural junctions, demonstrating the physical validity of the deterministic sampling strategy (Fig. 2). Parametric studies show that appropriate ACA thresholds and four sampling frequencies provide the best balance between computational efficiency and numerical accuracy (Figs. 35). The broadband RCS predicted by the proposed framework agrees closely with the MoM results over the entire frequency band (Figs. 68), and the RMSE remains low, demonstrating high numerical accuracy. Compared with CS-HBFM, the proposed framework reduces the total computation time by 93.4% for the cylinder, 96.7% for the slotted cone, and 80.9% for the almond (Table 2). These improvements result from deterministic index reuse, MMI, and dual ACA acceleration, which substantially reduce the computational cost of broadband frequency-sweeping analysis.  Conclusions  A CS-ACA-MMI framework is proposed by integrating ACA with MMI for efficient broadband electromagnetic scattering analysis. The proposed framework eliminates repeated matrix construction and equation solving during frequency sweeping while overcoming the nondeterministic sampling strategy and the high computational and memory costs of conventional CS-HBFM. Dominant row indices extracted by ACA at the highest frequency provide a deterministic measurement-matrix construction strategy and a stable physical basis for broadband interpolation. By shifting the interpolation target from full impedance matrices to low-dimensional measurement matrices, the computational complexity and redundant matrix construction are substantially reduced. A second ACA decomposition further accelerates sensing-matrix construction by converting large-scale matrix-vector products into low-dimensional matrix multiplications. Numerical results demonstrate that the proposed framework achieves numerical accuracy comparable to that of MoM while reducing total computation time by more than 80% and decreasing single-frequency measurement-matrix memory consumption by up to 65%. Because only the measurement matrices at four sampling frequencies need to be stored, the overall memory requirement is further reduced.
Construction and Performance Analysis of Optimal Low-Hit-Zone Frequency Hopping Sequence Sets
TIAN Xinyu, CHEN Xiaoyu, ZHANG Jitao
Available online  , doi: 10.11999/JEIT260343
Abstract:
  Objective  ElectroMagnetic Interference (EMI) is a critical factor limiting the reliability of synchronization systems. Existing Fifth-Generation (5G) synchronization schemes extensively employ Zadoff-Chu (ZC) sequences to distinguish users through cyclic shifts. However, finite sequence lengths and limited orthogonal resources create substantial capacity bottlenecks in high-density access scenarios. To address these challenges, this paper investigates the problem from two perspectives. At the system level, a synchronization framework is developed by integrating Frequency Hopping (FH) with ZC sequences. By jointly exploiting code, time, and frequency-domain resources, the proposed framework improves concurrent access capability for local clusters while enhancing robustness against complex EMI through frequency diversity. At the sequence-design level, a class of multi-subset Low-Hit-Zone (LHZ) Frequency Hopping Sequence (FHS) sets is constructed to provide an efficient sequence allocation scheme for local-cluster synchronization.  Methods  Based on the theoretical framework proposed by Cai et al., the sequence mapping mechanism is reconstructed, and a disjoint Cyclic Perfect Mendelsohn Difference Family (CPMDF) is introduced to construct FHS sets that are optimal with respect to the Peng-Fan bound. The generating units are further expanded through Cartesian products, and a column-incoherent partitioning strategy is proposed to construct multi-subset LHZ FHS sets. It is proved that every nonempty subset satisfies the Peng-Fan-Lee bound with equality. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure. At the system level, an FH-ZC synchronization architecture is developed by combining predefined FH patterns with the frequency-domain correlation properties of ZC sequences for subband signal detection. A Peak-to-SideLobe Ratio (PSLR) decision metric and an early-termination strategy are adopted to evaluate synchronization preamble detection under interference. Furthermore, a multi-user simulation model is established to evaluate synchronization detection performance under accumulated co-channel collisions and EMI.  Results and Discussions  The proposed construction generates an FHS set that is optimal with respect to the Peng-Fan bound and a class of multi-subset LHZ FHS sets in which every nonempty subset is optimal with respect to the Peng-Fan-Lee bound. Example 2 demonstrates the construction procedure and the intra-subset and inter-subset Hamming correlation properties of the proposed multi-subset LHZ FHS sets. Table 1 shows that, under the same frequency-resource constraints, the proposed construction generates more sequences than existing methods under the compared parameter settings, indicating higher sequence-resource utilization. Table 2 compares the parameters of the proposed sequence sets with representative constructions reported previously and demonstrates that the proposed multi-subset optimal sequence family provides a new parameter combination. To the best of our knowledge, an optimal sequence family with a multi-subset structure has not been reported previously. Figures 2 and 3 demonstrate that the proposed FH-ZC synchronization architecture achieves a higher synchronization detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI. Figure 4 shows that the synchronization detection probability decreases as the number of active users increases because accumulated co-channel collisions degrade synchronization performance. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure characterized by strong intra-cluster competition and weak inter-cluster coupling.  Conclusions  To satisfy the sequence-capacity requirements of massive-access scenarios, this paper proposes a class of multi-subset LHZ FHS sets. By expanding the generating sequence sets through Cartesian products and partitioning subsets using a column-incoherent strategy, the proposed construction achieves both a large family size and optimal LHZ performance. The proposed multi-subset structure is well suited to local-cluster synchronization and substantially improves sequence family size and sequence-resource utilization, thereby providing a richer sequence resource pool for high-density multi-user systems. Simulation results under the considered physical-layer model demonstrate that the proposed LHZ FHS subsets reduce the effect of frequency collisions during multi-user synchronization detection. Furthermore, the FH-ZC synchronization scheme achieves a higher synchronization preamble detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI.
A Phase Transition Obstacle Avoidance Method for UAV Swarms Driven by Multistable Potential Fields
HE Ming, CHEN QiYang, HAN Wei, PAN Fan, MA YiSong
Available online  , doi: 10.11999/JEIT260357
Abstract:
  Objective  Unmanned Aerial Vehicle (UAV) swarms have demonstrated considerable potential for complex missions, such as search, surveillance, and disaster response, because of their distributed coordination and robustness. However, in dynamic environments with dense obstacles and rapidly changing risks, conventional swarm control methods often exhibit discontinuous behavior switching and control chattering, which reduce system stability and coordination efficiency. Existing approaches, including threshold-based switching and Artificial Potential Field (APF) methods with fixed potential weights, rely on abrupt transitions between behavioral modes, leading to oscillatory responses. To address these limitations, a phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. Swarm behavior evolution is modeled as a continuous phase transition process within a unified potential field framework, enabling smooth and adaptive transitions between formation flight and obstacle avoidance.  Methods  An environmental risk assessment model is first established by integrating static obstacle risk, dynamic obstacle risk, and inter-agent proximity risk. A distributed consensus protocol is then employed to establish global risk consensus. Subsequently, a morphology factor is generated through nonlinear mapping of the global risk consensus and is used as an order parameter to characterize the macroscopic swarm state. A unified time-varying potential field, comprising formation, obstacle avoidance, and navigation potentials, is constructed, and the relative weights of these potentials are continuously adjusted by the morphology factor. When the risk level is low, the system exhibits a monostable structure dominated by the formation and navigation potentials. As the risk increases, the potential field continuously evolves into a multistable structure dominated by the obstacle avoidance potential, thereby enabling distributed obstacle avoidance. A distributed consensus control law based on the negative gradient of the unified potential field is further developed. A damping term is incorporated to dissipate system energy and improve stability, while a dynamic compensation term addresses nonlinear dynamics. The control law depends only on local information, ensuring good scalability. The global uniform ultimate boundedness of the closed-loop system is established using Lyapunov theory.  Results and Discussions  Simulation results demonstrate that the proposed method enables the swarm to maintain a compact solid-phase swarm formation in low-risk regions and to transition smoothly to a dispersed liquid-phase swarm configuration when obstacles are encountered, followed by rapid formation recovery after obstacle avoidance. The pitch and roll angles of each UAV vary smoothly without abrupt changes, and both the UAV-to-obstacle distance and the inter-UAV separation remain above the prescribed safety threshold throughout the flight, ensuring collision-free operation. Statistical results obtained from 20 independent simulation runs show that, compared with the threshold-switching method, the proposed method reduces the rate of control input variation by approximately 26% and decreases the peak control input by approximately 18%. Compared with the bio-inspired diversion method, the average formation recovery time after obstacle avoidance is reduced by approximately 16%. Ablation experiments further demonstrate that removing the morphology-driven phase transition mechanism significantly increases trajectory oscillation and control oscillation, confirming the critical role of the multistable continuous phase transition mechanism in maintaining smooth swarm motion. In complex narrow-channel environments, the proposed method effectively avoids the local minimum problem encountered by conventional APF methods and generates smoother flight trajectories with substantially reduced oscillation.  Conclusions  A phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. By introducing a morphology factor and constructing a unified potential field framework, swarm behavior evolution is represented as a continuous phase transition process. The distributed control law enables smooth behavioral transitions while maintaining system stability and scalability. Simulation results demonstrate that the proposed method achieves better safety, smoother control, and higher coordination efficiency than conventional methods.
An SO(3)-Manifold-Constrained Registration Method for Twin-Fisheye Panoramic Images
WANG Zhuopeng, LIN Shanling, LIN Jianpu, LÜ Shanhong, LIN Zhixian
Available online  , doi: 10.11999/JEIT260798
Abstract:
  Objective  Twin-fisheye cameras provide near-360° coverage with low hardware complexity and are widely used in immersive imaging, surveillance, and mobile robotics. Their panoramic output depends on registration over a narrow overlapping band, so geometric accuracy and temporal consistency directly affect seam quality and video smoothness. After the two fisheye views are unfolded into the Equirectangular Projection (ERP), three coupled problems arise. First, the near-co-centric lens pair is ideally related by a pure rotation R ∈SO(3), whereas a conventional 8-Degree-of-Freedom (DoF) homography introduces five redundant parameters that may couple with matching noise. Second, ERP sampling is nonuniform with latitude, so identical pixel residuals do not represent identical spherical angular errors. Third, the cyclic ±π longitude boundary splits structures that are continuous on the sphere and weakens correspondences around the seam. Existing planar pipelines and generic learned matchers rarely combine these constraints under a unified rotation-referenced evaluation. This study therefore develops a lightweight registration framework that explicitly exploits spherical rotation geometry while addressing ERP boundary discontinuity, temporal fluctuation, and long-tail residuals.  Methods  The proposed framework contains three modules (Fig. 1). First, an overlapping-band Region-of-Interest (ROI) is cropped around the ERP seam and rearranged with modulo-W wrap-around (Fig. 2). A default longitude half-width of ±15° and an approximately 3° margin on each side preserve cross-boundary feature continuity while restricting the search region. XFeat detects, describes, and matches features under a fixed Top-K budget. Second, the two-dimensional matches are restored to global ERP coordinates, mapped to unit-sphere direction vectors, and processed by rotation-only SO(3)-RANSAC. Spherical angular residuals are used as the inlier criterion with a 0.8° threshold, a maximum of 2000 iterations, confidence 0.999, and at least 12 inliers; the iteration bound is updated adaptively. All inliers are then used for Kabsch/SVD closed-form rotation re-estimation, which reduces the randomness of a minimal sample while preserving the SO(3) constraint. Third, a local increment on the Lie algebra so(3) is optimized under a Huber loss by the Levenberg–Marquardt algorithm. The refinement is triggered only when the inlier-residual P95 exceeds 0.90° and the inlier ratio is below 0.58, thereby concentrating nonlinear optimization on difficult image pairs.  Results and Discussions  Experiments are conducted on PanoraMIS Sequences 3 and 4 under a unified relative inter-frame rotation protocol (Table 1). The proposed method achieves a 97.10% success rate, a 0.549° P95 angular residual, and a 0.313° temporal-stability error. Compared with SuperPoint+LightGlue, the P95 and temporal-stability errors are reduced by 22.8% and 77.3%, respectively. Compared with Efficient LoFTR, peak GPU memory and runtime are reduced by 52.7% and 60.9%, although Efficient LoFTR retains the lowest overall P95. Under a unified SO(3)-RANSAC back-end (Table 2), XFeat provides the largest average inlier count of 625.4 and the lowest temporal-stability error of 0.313° at 38.04 ms. The ablation and sensitivity results (Tables 34) show that the ±15° ROI reduces the P95 from 0.720° for the full ERP to 0.600°. Replacing H-RANSAC with SO(3)-RANSAC reduces temporal instability from 0.931° to 0.313°, a 66.3% reduction, while increasing runtime from 24.91 ms to 35.24 ms. Adaptive refinement operates on approximately one third of the image pairs and improves both P95 and temporal stability with lower overhead than always-on refinement; its five-seed mean and median trigger rate are both 36.23%. A Top-K budget of 2048 reaches the saturated accuracy level, because increasing the budget to 4096 yields no further P95 or stability improvement. Five fixed-seed repetitions produce standard deviations no greater than 0.011° for P95 and stability, indicating that the main conclusions are insensitive to RANSAC randomness. In a 5×5 threshold sweep, the maximum changes in P95 and inter-frame rotation jitter within the central neighborhood are 4.40% and 0.037%, respectively, and changing the robust-error truncation from 3° to 2° or 5° does not alter the relative ranking. On the more difficult Sequence 4, characterized by weak texture and unstable overlap, the proposed method obtains 919.7 average inliers and a P95 of 0.383°, the lowest among the evaluated learning-based matchers, although robust-estimation time increases. With an identical standardized stitching back-end, it produces a lower seam-band gradient than H-RANSAC in the small-rotation example (26.90 versus 28.71) and a lower truth-referenced temporal-stability error (0.313° versus 0.931°; Fig. 6). On outdoor Sequence 7-L2, 324 of 346 correspondences are retained, yielding a 93.6% inlier ratio and a 0.524° P95 residual (Fig. 5). Because sequence-specific calibration is unavailable and the fixed inter-lens baseline may cause depth-dependent parallax, this result serves only as a diagnostic consistency check, not as evidence of absolute pose accuracy.  Conclusions  By restoring feature continuity across the ERP boundary, replacing the redundant planar homography with an explicit SO(3) rotation model, and selectively refining difficult image pairs on so(3), the proposed method balances registration accuracy, temporal consistency, and resource cost. It provides a lightweight front-end for twin-fisheye panorama stitching. The current evaluation is limited to pairwise registration on a small number of sequences; future work will address translation compensation, multi-frame global optimization, end-to-end integration with seam finding, exposure compensation, and blending, as well as generalization across additional platforms, dynamic scenes, and illumination conditions.
Design and Verification of Robust Modulation Recognition Framework Under Blind Adversarial Attacks
ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260019
Abstract:
  Objective  Deep learning-based Automatic Modulation Recognition (AMR) models demonstrate strong performance in non-cooperative communication systems such as cognitive radio and spectrum monitoring. However, deep learning models remain vulnerable to adversarial attacks. In these attacks, imperceptible perturbations lead to severe misclassification and create security risks. Existing defense methods, including adversarial training, often rely on prior knowledge of specific attacks. They also introduce considerable computational overhead and reduce accuracy on clean samples. This study designs and verifies a robust modulation recognition framework that operates effectively under blind adversarial attack scenarios without prior knowledge of attack type or strategy. The goal is to support reliable deployment of intelligent communication systems in adversarial environments.  Methods  The proposed framework integrates a feature-purifying autoencoder module with standard modulation classifiers, including Convolutional Neural Network (CNN) and Transformer architectures. The core component is the autoencoder bottleneck layer, which implements a dynamic purification mechanism. First, an adaptive threshold is calculated from the statistical properties of encoded latent features to detect anomalies. Then a Top-K sparsification operation retains the most significant feature activations. This step suppresses noise and adversarial perturbations and preserves essential signal characteristics. The autoencoder is trained using a three-stage curriculum learning strategy. The stages sequentially optimize reconstruction fidelity, feature sparsity, and semantic consistency between purified signals and original clean signals. This process guides the reconstructed signals toward the true modulation manifold. The module is model-agnostic and can be placed before a trained classifier without retraining.  Results and Discussions  Experiments are conducted on a simulated dataset containing 12 digital modulation types under multipath fading channels. The framework produces clear performance gains. Under targeted white-box attacks, recognition accuracy increases to 82.1% for CNN and 83.2% for Transformer. Under non-targeted black-box attacks, accuracy reaches 87.7% and 89.4%, respectively (Table 1). The Attack Success Rate (ASR) and Attack Effectiveness Index (AEI) remain low, indicating strong defense capability. Figure 4 shows that defense performance improves as the Signal-to-Noise Ratio (SNR) increases. The ablation study in Figure 5 confirms the critical role of the autoencoder. Removing this module reduces accuracy by 4.02% for CNN and 2.36% for Transformer under strong attacks. Further analysis in Figure 6 shows that the framework maintains stable robustness across a wide perturbation range (\begin{document}$ \epsilon \leq 0.1 $\end{document}). Parameter sensitivity experiments in Figures 7 and 8 indicate stable performance when the threshold coefficient \begin{document}$ \xi $\end{document} is within [1.5, 1.9] and the sparsity rate k is around 0.7. These results support practical deployment.  Conclusions  A robust blind defense framework for AMR is presented based on a feature-purifying autoencoder. The framework provides three main advantages. First, it defends against different white-box and black-box attacks without requiring prior knowledge of attack methods. Second, as a preprocessing module, it avoids computationally expensive retraining of the primary classifier and remains compatible with different backbone networks. Third, the multi-stage training strategy balances adversarial robustness with high accuracy on clean samples. Experiments on the simulated dataset confirm the effectiveness of the proposed framework. Future work will explore lightweight architectural designs to reduce inference latency and will further investigate performance limits under extremely low SNR conditions combined with nonlinear channel impairments.
Co-Frequency Interference Analysis and Dynamic Simulation Validation of Satellite-Direct-to-Device Systems Against Terrestrial IMT Networks in Cross-Border Scenarios
LIU Quan, ZHAO Weisong, XIAO Na, SONG Yanjun, ZHOU Meng, ZHANG Zhili, WANG Jinhai, WANG Lichong
Available online  , doi: 10.11999/JEIT260263
Abstract:
  Objective   Satellite-Direct-to-Device (SD2D) systems that reuse terrestrial IMT spectrum may generate harmful downlink interference to incumbent IMT networks in neighboring administrations, particularly in cross-border deployments where SD2D downlinks overlap the receive bands of both IMT user equipment (UE) and IMT Base Stations (BSs). A practical coexistence methodology is therefore required to (i) translate IMT receiver protection criteria into explicit Power Flux Density (PFD) and Equivalent Power Flux Density (EPFD) constraints and (ii) validate these constraints using a dynamic simulation framework so that they can be converted into enforceable geographic coordination measures, such as minimum isolation distances. This study focuses on the dominant interference path, namely SD2D downlink interference to IMT receivers, and establishes a traceable workflow from deterministic protection limits to dynamic simulation validation and the corresponding minimum isolation distances.  Methods  A cross-border scenario is modeled in which Country A deploys an SD2D system and Country B operates a terrestrial IMT network. Two representative downlink frequencies, 1 995 MHz and 2 190 MHz, are evaluated for two representative Starlink configurations, Starlink-1 and Starlink-2. The IMT network is modeled using ITU-R- and 3GPP-compliant parameters, with an I/N protection threshold of –6 dB and a target percentile κ (baseline κ=99.5%) for both IMT UEs and BSs. Satellite transmit antennas follow the ITU-R S.1528 reference pattern, IMT BS receive antennas follow the ITU-R F.1336 sector pattern, and IMT UEs are modeled with omnidirectional antennas. A back-lobe blockage model is incorporated into both satellite and BS antenna patterns to account for rear-side shielding. Signal propagation follows the ITU-R P.619 model, using free-space path loss as the conservative baseline, while an optional clutter-loss term is incorporated through a clutter-occurrence probability. Deterministic protection limits are derived by calculating the maximum permissible aggregate PFD for IMT UE protection and the maximum permissible aggregate EPFD for IMT BS protection. A dynamic simulation framework then validates these limits and searches for the required minimum isolation distances (Fig. 4). Co-channel beam isolation angles are optimized using the C/I Complementary Cumulative Distribution Function (CCDF), and a segmented search algorithm determines the minimum UE- and BS-side isolation distances together with the corresponding κ-percentile PFD/EPFD statistics.  Results and Discussions  The deterministic analysis yields a maximum permissible aggregate PFD of –102.72 dBW/m2/MHz at 2 190 MHz for IMT UEs and a maximum permissible aggregate EPFD of –129.53 dBW/m2/MHz at 1 995 MHz for IMT BSs (Fig. 3). For Starlink-1, the C/I design criterion yields a minimum co-channel beam isolation angle pair of (12°, 12°) (Fig. 5). Dynamic simulation shows that, under the representative baseline configuration with an I/N threshold of –6 dB and κ=99.5%, the minimum isolation distances are 195 km for UE protection and 290 km for BS protection (Fig. 6, Fig. 7, and Table 4). The resulting coordination isolation distance is therefore 290 km, and the simulated κ-percentile PFD and EPFD agree with the deterministic protection limits, with a residual margin below 0.5 dB. For Starlink-2, the optimized co-channel beam isolation angles increase to (15°, 15°), and the corresponding minimum isolation distances increase to 272 km for UEs and 420 km for BSs under the same baseline configuration (Table 5). These baseline distances should be interpreted as representative values for the specified simulation configuration rather than unique, strictly converged results. Stability verification shows that, under different sampling intervals, simulation durations, and random seeds, the UE- and BS-side minimum isolation distances remain within 195~210 km and 290~300 km, respectively, for Starlink-1, and within 266~290 km and 370~420 km, respectively, for Starlink-2 (Table 7). Sensitivity analysis for Starlink-1 further indicates that the required minimum isolation distance is governed by the upper tail of the aggregate I/N distribution (Table 6). Increasing κ from 99.5% to 100% increases the UE- and BS-side minimum isolation distances from 195/290 km to 304/560 km. Clutter attenuation substantially reduces the UE-side minimum isolation distance, decreasing it to 173 km when the clutter-occurrence probability is 0.5, while producing little change in BS protection. Polarization reuse increases the UE- and BS-side minimum isolation distances to 222 km and 360 km, respectively, whereas increasing the number of co-channel beams to 16 increases the BS-side minimum isolation distance to 330 km. The minimum service elevation angle and the link establishment strategy are identified as the dominant operational factors. Changing the minimum service elevation angle from 10° to 35° changes the required UE- and BS-side minimum isolation distances from 340/460 km to 101/150 km, whereas replacing the Sat-MaxElevation strategy with the UE-MaxElevation strategy reduces them to 80/180 km.  Conclusions   The proposed workflow converts IMT receiver protection criteria into deterministic protection limits expressed as PFD and EPFD constraints and validates them using a dynamic simulation framework. Under an I/N threshold of –6 dB and κ=99.5%, the baseline and stability analyses jointly indicate representative UE- and BS-side minimum isolation-distance ranges of 195~210 km and 290~300 km for Starlink-1 and 266~290 km and 370~420 km for Starlink-2, rather than unique, strictly converged values. Sensitivity analysis further shows that κ only changes the statistical criterion used to extract tail events from the sample set, whereas clutter attenuation primarily benefits IMT UEs. In contrast, the minimum service elevation angle, polarization reuse, the number of co-channel beams, and the link establishment strategy reshape the worst-case interference geometry and can produce substantial, and sometimes non-monotonic, changes in the required minimum isolation distances. The proposed framework establishes a traceable link between IMT receiver protection criteria and enforceable border coordination measures.
MG-MoE: Routed Multi-Granularity Expert Ensemble
XIAN Fengyu, JIAN Haifang, XIE Zihui, DU Jun, ZHANG Yuanyuan, NING Xin, DONG Miaomiao, WANG Hongchang
Available online  , doi: 10.11999/JEIT260219
Abstract:
  Objective  Fine-Grained Image Recognition (FGIR) aims to distinguish visually similar subcategories that differ only in subtle local patterns. It must also remain robust to large intra-class variations caused by pose changes, occlusion, illumination shifts, and complex backgrounds. In real-world scenarios, these challenges are further intensified by long-tailed category distributions. Rare or difficult classes are more likely to overfit spurious contextual cues and suffer from unstable decision boundaries. Therefore, a conditional computation paradigm is needed, in which complementary inductive biases are separated into specialized expert branches and adaptively combined for each sample. This work aims to develop a routed multi-granularity mixture-of-experts framework that improves discriminative performance under controllable inference cost. It also enhances robustness for difficult samples and long-tailed categories through adaptive sparse expert activation.  Methods  A Multi-Granularity Mixture-of-Experts (MG-MoE) model is proposed. It is a routed ensemble architecture composed of a shared backbone, four heterogeneous experts, and a learnable router that predicts input-conditioned expert weights (Fig. 2). The experts are designed with complementary inductive biases to address key factors in FGIR. MPSA emphasizes global structure and contour-level semantics. PMG captures fine local details through multi-granularity part modeling. TransFG focuses on pose and deformation modeling. PIM improves robustness in cluttered backgrounds through background suppression. To limit interference and reduce unnecessary computation, MG-MoE adopts sparse fusion. Only the Top-K experts, with K=2 by default, contribute to the final prediction during inference. To improve routing stability and generalization, a two-stage optimization strategy is designed. In the first stage, dynamic cluster-level training is performed. A cluster-level soft teacher distribution is constructed from validation-set statistics and imposed through Kullback-Leibler (KL) divergence regularization. This process stabilizes routing behavior and promotes effective expert specialization. In the second stage, residual fine-tuning is conducted. The feature-driven routing mechanism is kept unchanged, while the classification heads of the Top-2 experts associated with each cluster are selectively unfrozen. The router and expert heads are then jointly optimized with grouped learning rates. This design reduces fusion bias and strengthens discrimination for difficult samples and long-tailed categories.  Results and Discussions  MG-MoE achieves strong performance on standard FGIR benchmarks. On CUB-200-2011, it obtains 92.89% Top-1 accuracy. This result is higher than those of representative expert backbones used individually, including MPSA (91.23%), PIM (91.17%), and TransFG (90.49%). It also outperforms the multi-granularity baseline PMG (88.32%) (Table 1). On the Bird-1445 sampled set, MG-MoE achieves 96.80% Top-1 accuracy and consistently improves over strong baselines (Table 2). These results indicate that routed multi-expert specialization remains effective in data-limited and highly similar fine-grained scenarios. The efficiency-accuracy trade-off is summarized in Table 3. With Top-2 sparse routing, MG-MoE reaches 92.89% accuracy with a compute budget of 143.9 GFLOPs. It avoids dense expert activation during inference by selecting only the Top-2 experts for each sample, thereby achieving a favorable balance between accuracy and efficiency. Ablation experiments show that increasing K beyond 2 does not yield consistent gains, which suggests that indiscriminate fusion can dilute discriminative evidence. Top-2 fusion produces the best performance, whereas Top-1 fusion is more sensitive to routing errors and larger K values may introduce noise and reduce accuracy (Table 4). The role of expert diversity and composition is also analyzed. Two- and three-expert variants generally underperform the full four-expert configuration, indicating that each inductive bias contributes to different fine-grained difficulty factors. In contrast, adding homogeneous experts without new functional diversity brings diminishing or negative gains, which is consistent with increased routing ambiguity and limited expert complementarity (Table 5). These results support the use of a compact set of heterogeneous experts combined with sparse routing. To interpret the learned specialization, category-wise routing statistics are visualized. The expert-category heatmap shows that MPSA receives dominant routing weights across many categories, reflecting the central role of global structure in fine-grained discrimination. PIM and TransFG show higher activation for specific difficult categories, which is consistent with their roles in background suppression and pose and deformation modeling (Fig. 3). Finally, t-SNE visualizations illustrate the qualitative effect of expert fusion on class separability. Shared backbone features show stronger inter-class entanglement among visually similar subcategories. In contrast, fused outputs form clearer clusters with better between-class separation and within-class compactness, indicating a more reliable decision space shaped by routed expert aggregation (Fig. 4).  Conclusions  MG-MoE is a multi-granularity routed mixture-of-experts framework for fine-grained recognition. By combining four complementary experts, Top-2 sparse fusion, and a two-stage optimization strategy for stable routing and calibrated fusion, MG-MoE improves recognition accuracy on CUB-200-2011 and the Bird-1445 sampled set. It also provides interpretable evidence of expert specialization (Table 1, Table 2, Fig. 3, Fig. 4). Ablation results confirm that controlled Top-2 fusion and heterogeneous expert design are key to the observed performance gains. Overly dense fusion or homogeneous expert expansion provides limited benefit (Table 4, Table 5).
Energy-Aware and Attention-Driven Edge–End Collaborative Inference and Resource Allocation
LIU Yiming, TIAN Jie, LI Tiantian, ZHOU Xiaotian, ZHANG Haixia
Available online  , doi: 10.11999/JEIT260086
Abstract:
  Objective   The perception and decision-making capabilities of mobile intelligent applications have been significantly improved by the development of Deep Neural Networks (DNN). Nevertheless, these applications often have high computational loads, which greatly strain mobile terminals with constrained latency and energy consumption. Due to limited wireless bandwidth and edge computing power, severe resource contention is inevitable in multi-user concurrent scenarios, even though mobile edge computing (MEC) lowers latency pressure by moving tasks from the edge to the terminal. More importantly, because edge servers have limited resources, strict budget control of their long-term energy consumption is necessary to ensure the system's dependability and long-term operational efficiency.An energy-aware and attention-driven edge-to-edge collaborative inference and resource allocation technique is proposed in this paper. The majority of research ignores the system's long-term energy consumption limitations in favor of maximizing instantaneous performance. This approach greatly lowers the average end-to-end latency of multi-user inference tasks under various load scenarios and greatly increases the utilization efficiency of general computing resources while closely adhering to the long-term energy budget.  Methods   Based on the DNN collaborative inference model in an edge environment, this paper models the problem as minimizing the long-term average end-to-end processing latency of all user inference tasks, under the premise of satisfying the long-term energy consumption budget constraint of the ES system. This optimization problem and energy consumption constraint both involve long-term averages and stochasticity, and the DNN model partitioning and general computing resource allocation decisions are highly coupled. Therefore, this paper adopts Lyapunov optimization theory to transform this complex problem into a more manageable single-slot deterministic optimization problem. In order to jointly finish the DNN partitioning and general computing resource allocation decisions within each time slot, this paper develops a Joint Collaborative Inference and Resource Allocation Algorithm (JCIRA). There are three steps in the algorithm (Algorithm 1). To determine the ideal DNN model partitioning point for the task, the first step balances the estimated latency against energy penalties based on the comprehensive cost function. Task urgency, remaining computational load, data volume, and global energy deficit state are all mapped into high-dimensional feature vectors in the second stage, which presents a joint attention mechanism based on the key-query-value paradigm. To accomplish cooperative distribution of general computing resources, the matching degree is computed using a scaled dot product attention mechanism. In order to improve the system's utilization rate of general computing resources and lower the average end-to-end latency of multi-user inference tasks while closely adhering to long-term energy consumption constraints, the third stage creates a closed-loop decision-making process by monitoring execution progress and updating queue status.  Results and Discussions   In this simulation experiment, the task arrival process is random and has a Poisson distribution. In this experiment, three distinct deep neural network models were employed: ResNet18, MobileNetV2, and EfficientNet-B0. When compared to other algorithms, JCIRA's actual energy consumption under various load scenarios consistently falls below the budget threshold, demonstrating the proposed algorithm's ability to manage long-term energy consumption constraints (Fig. 2). Furthermore, JCIRA's virtual queue length is always brief (Fig. 3). Among all schemes that meet the energy consumption constraints, JCIRA's latency is consistently below the 300 ms QoS threshold (Fig. 4). It performs better and successfully resolves the problem of other algorithms' insufficient processing power under heavy loads. Finally, the algorithm maintains the highest task completion rate even with 1300 concurrent tasks in a high load scenario (Fig. 5).  Conclusions   To address the challenges of limited computing resources and long-term energy consumption constraints in DNN collaborative inference within a MEC environment, this paper proposes an energy-aware and attention-driven edge-end collaborative inference and resource allocation method. The long-term energy consumption hard constraint is separated into a low-complexity single-slot deterministic optimization subproblem using Lyapunov optimization theory. In order to accomplish systematic scheduling of DNN model partitioning and computing resources, a joint optimization algorithm called JCIRA is created. Simulation results show that the suggested approach significantly lowers the average end-to-end latency of multi-user inference tasks under different load conditions and efficiently optimizes the use of computing resources while guaranteeing strict satisfaction of long-term energy consumption constraints.
Efficient Non-Orthogonal Multiple Access Scheme Based on Modified Alamouti Code Design
WAN Dehuan, HUANG Ronglan, JI Fei, LIU Jingxian, LIANG Yaokun, LÜ Lu, LI Xingwang, YUE Xinwei
Available online  , doi: 10.11999/JEIT260567
Abstract:
  Objective  Existing Alamouti-coding-based Non-Orthogonal Multiple Access (NOMA) schemes adopt an equal-number symbol transmission mode for both cell-center and cell-edge users, which overlooks the significant channel disparity between the two types of users. This leads to two drawbacks: on one hand, the superior channel condition of cell-center users is not fully exploited, resulting in wasted transmission resources; on the other hand, the equal-number transmission inevitably weakens the performance of Space-Time Block Coding (STBC) in suppressing intra-group multi-user interference. Therefore, a novel design that adapts to user channel differences is urgently needed to improve spectral efficiency and interference mitigation capability.  Methods  In light of the significant channel disparity between cell-center and cell-edge users, this paper proposes an unequal-number symbol transmission scheme for Alamouti coding in NOMA. Specifically, when constructing Alamouti group transmission codes, the number of symbols required for the cell-edge user is made smaller than that for the cell-center user. By reducing the number of transmitted symbols for the cell-edge user, two benefits are achieved: first, the multi-user interference imposed on the cell-center user is directly reduced; second, under a total power constraint, reducing the number of symbols for the cell-edge user equivalently increases the per-symbol transmission power, thereby significantly improving the received signal-to-interference-plus-noise ratio (SINR). Furthermore, the paper derives perfect closed-form solutions for the achievable sum rate and outage probability of the proposed scheme, and validates its effectiveness via numerical simulations.  Results and Discussions  The theoretical derivations yield closed-form analytical expressions for the achievable sum rate and outage probability, providing an accurate basis for system performance evaluation. Numerical simulation results demonstrate that, compared with existing equal-symbol Alamouti coding schemes, the proposed unequal-symbol Alamouti coding scheme effectively reduces the intra-group interference from the cell-edge user to the cell-center user, while significantly enhancing the SINR of the cell-edge user through power reallocation. Under typical channel parameters, the system achieves a notable sum-rate gain and a substantial reduction in outage probability, confirming the high efficiency of the proposed scheme.  Conclusions  The proposed unequal-number symbol transmission scheme based on Alamouti coding for cell-edge and cell-center users fully exploits the potential benefits arising from channel disparity. By reducing the number of symbols transmitted by the cell-edge user, the scheme achieves both suppression of intra-group interference and enhancement of the edge user’s transmission power. The theoretical closed-form solutions and simulation results consistently show that the proposed scheme outperforms conventional equal-number transmission schemes, providing an effective new approach for mitigating multi-user interference and optimizing resource allocation in NOMA systems.
Function-Aware Partitioning Driven Hierarchical Circuit Representation Learning
YE Juyang, CHEN Qilin, WANG Yaohua
Available online  , doi: 10.11999/JEIT260645
Abstract:
  Objective  One of the core challenges in applying machine learning techniques to electronic design automation (EDA) lies in learning high-quality circuit representations from large-scale gate-level netlists. As modern digital integrated circuits scale to tens of millions of gates, existing methods based on graph neural networks (GNNs) and graph transformers (GTs) suffer from excessive computational and memory overhead, rendering them impractical for industrial-scale designs. The fundamental issue is the lack of a proper tokenization mechanism for netlists—unlike natural language, where subword tokenization effectively compresses long sequences, the circuit domain lacks an analogous decomposition strategy that preserves functional semantics while reducing the effective graph size. This work aims to bridge this gap by introducing a hypergraph-partitioning-based circuit tokenizer that decomposes massive netlists into functionally cohesive sub-circuits, termed circuit elements, thereby enabling scalable and fine-grained representation learning.  Methods  This paper proposes a function-aware hypergraph partitioning driven framework for large-scale circuit representation learning. Inspired by the tokenization paradigm of large language models, the framework first models a gate-level netlist as a directed hypergraph, where gates are nodes and signals are hyperedges that can connect multiple gates. A novel optimization objective, the Functional Independence Ratio (FIR), is introduced to guide the partitioning process. FIR incorporates circuit structural priors and a bus recognition correction mechanism that identifies bus-structured signals based on structural similarity (gate type purity across predecessor/successor levels) and spatial similarity (topological distance variance within candidate groups). The bus recognition module corrects the effective interface count, ensuring that functionally cohesive modules are not penalized for using wide buses. An iterative greedy refinement procedure accepts only moves that strictly decrease FIR, converging to a locally optimal partition. On top of the partitioned circuit elements, a two-stage self-supervised pretraining framework is designed. In the first stage, a masked autoencoder with edge prediction tasks is applied to the coarse-grained circuit-element graph, learning global inter-element dependencies. The graph transformer encoder is then frozen and circuit element embeddings are saved. In the second stage, within each circuit element, a contrastive learning scheme is employed at the gate level. Positive pairs are constructed via Boolean equivalence transformations (e.g., associativity, De Morgan's laws), which preserve the Boolean function while altering the gate-level structure. Negative pairs are drawn from functionally different circuit elements within the same batch. The training jointly optimizes node-level and local-global alignment losses. A feature-wise modulation mechanism injects circuit-element-level context into gate-level representations, enabling the same gate type to acquire different embeddings depending on its functional context.  Results and Discussions  Extensive experiments are conducted on circuits collected from multiple sources, including ITC99, EPFL, OpenCores, and three RISC-V SoC designs (Rocket, BOOM, and OpenC910), with the largest design containing 22.3 million gates. For the tokenizer evaluation, FIR-based partitioning is compared against Mt-kahypar using the Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) metrics, with the original module hierarchy serving as ground truth. Across all three RISC-V designs, FIR achieves AMI improvements of 24.6% to 27.1% and ARI improvements of 25.5% to 35.8% over Mt-kahypar, demonstrating consistent and substantial gains in functional coherence. On downstream tasks, the proposed method consistently outperforms state-of-the-art baselines including DeepGate4 and NetTAG across all comparable datasets. For full-circuit function recognition, the proposed method achieves F1 scores of 0.896, 0.861, 0.817, and 0.773 on ITC99, EPFL, OpenCores, and Rocket, respectively, representing improvements of 10.0% to 15.4% over the strongest baseline. Crucially, on BOOM (8.12M gates) and OpenC910 (22.3M gates), the proposed method is the only approach capable of completing training without running out of memory, achieving F1 scores of 0.771 and 0.748 for full-circuit tasks and 0.723 and 0.679 for gate-level tasks, respectively. Ablation studies with six model variants reveal a clear division of labor: the module-level pretraining stage dominates full-circuit performance (20.9% F1 drop when removed), while the gate-level pretraining stage dominates gate-level performance (26.1% F1 drop when removed). The FIR partitioning objective contributes 8.6% and 11.7% F1 improvements at the full-circuit and gate levels, respectively. The bus recognition module and the two-stage architecture also show consistent positive contributions across all metrics.  Conclusions  This paper presents a novel framework that addresses the scalability challenge in circuit representation learning by introducing a function-aware hypergraph partitioning tokenizer and a two-stage self-supervised pretraining architecture. The key insight is that by abstracting the intermediate circuit element level between individual gates and the full circuit, one can achieve both scalability to multi-million-gate designs and fine-grained gate-level discriminative capability. The proposed FIR objective effectively captures functional cohesion during partitioning, and the two-stage pretraining framework decouples global context learning from local representation refinement. Experimental results demonstrate state-of-the-art performance on function recognition tasks and, more importantly, the unique ability to scale to industrial-sized designs where existing methods fail. Future work includes extending the framework to other EDA tasks such as logic synthesis and physical design, exploring more aggressive hierarchical strategies for billion-gate designs, and adapting the bus recognition mechanism to non-standard cell libraries.
A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar
HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang
Available online  , doi: 10.11999/JEIT260695
Abstract:
  Objective  Traditional optimization-based power allocation algorithms for collocated MIMO radar have two fundamental limitations. First, they optimize tracking performance only for the next time step and therefore lack a full-time-horizon view of the power allocation process. This myopic strategy cannot achieve optimal multi-target tracking accuracy over extended periods, particularly when target trajectories vary substantially. Second, these algorithms rely on iterative nonlinear constrained optimization, resulting in high computational complexity. Therefore, they cannot satisfy the real-time requirements of dynamic battlefield environments where target states change rapidly. To address these limitations, this paper proposes a Reinforcement Learning (RL)-driven power allocation algorithm. Unlike conventional methods, the proposed approach formulates the power allocation problem as a Markov Decision Process (MDP) that maximizes long-term cumulative tracking accuracy. The algorithm adaptively allocates limited transmit power among multiple beams according to the current system state, balancing immediate tracking performance with long-term cumulative tracking accuracy.  Methods  The Posterior Cramér-Rao Lower Bound (PCRLB) is employed to quantify the theoretical lower bound of the tracking error for each target. The state space is constructed by combining the motion states (position and velocity) of all targets with the normalized PCRLB from the previous allocation step. The action space consists of discrete transmit power levels for each beam, subject to the total power budget and individual beam power constraints. All feasible power allocation vectors are enumerated and encoded to reduce the action-space dimensionality. The reward function is defined as the negative weighted sum of the normalized PCRLB, encouraging the agent to minimize tracking errors. The power allocation process is formulated as an MDP and solved using the Dueling Double Deep Q-Network (D3QN) algorithm. The D3QN framework incorporates three major enhancements: (1) a double-network architecture comprising an online Q-network and a target Q-network to improve training stability; (2) a dueling architecture that decomposes the Q-value into a state-value function and an action-advantage function to improve action discrimination; and (3) off-policy learning with experience replay to improve the use of historical trajectories. An ε-greedy strategy is adopted for exploration, with ε gradually decreasing during training. After offline training, the learned network directly generates real-time transmit power allocation decisions from the current system state without iterative optimization.  Results and Discussions  Simulations are conducted using three targets following the Constant Velocity (CV) model. Fixed power allocation yields the lowest tracking accuracy because of inefficient resource utilization. The traditional optimization method, which minimizes the instantaneous tracking error, achieves moderate tracking performance but remains myopic. When the discount factor \begin{document}$ \gamma =0 $\end{document}, the D3QN algorithm achieves performance comparable to that of the traditional optimization method because both optimize only immediate rewards. In contrast, when \begin{document}$ \gamma =0.99 $\end{document}, the D3QN algorithm significantly improves full-time-horizon tracking accuracy. The resulting power allocation strategy allocates more transmit power to distant, low-Signal-to-Noise Ratio (SNR) targets at earlier stages while reducing redundant power assigned to nearby high-SNR targets. The training curves show that \begin{document}$ \gamma =0.99 $\end{document} achieves a higher steady-state cumulative reward, although convergence exhibits greater oscillation because of the increased difficulty of estimating long-term returns. Furthermore, the trained D3QN network generates transmit power allocation decisions almost instantaneously, whereas the traditional optimization method must solve a constrained optimization problem at every time step, providing a substantial real-time computational advantage.  Conclusions  This paper proposes an RL-driven power allocation algorithm for collocated MIMO radar multi-target tracking that overcomes the myopic behavior and high computational complexity of conventional optimization methods. The proposed algorithm constructs the state space and reward function using the PCRLB, models the power allocation process as an MDP, and solves it using the D3QN algorithm. Simulation results demonstrate that, with an appropriate discount factor (\begin{document}$ \gamma =0.99 $\end{document}), the proposed approach significantly improves full-time-horizon tracking accuracy. This improvement results from the agent’s ability to learn a long-term optimal policy that proactively allocates transmit power to future distant, low-SNR targets. Furthermore, the trained network enables real-time decision-making through direct forward propagation, substantially reducing computational latency compared with iterative optimization. This work provides a new approach for intelligent radar resource management in complex battlefield environments.
Hierarchical Attention Mechanism-based Path Planning for Multi-UAV Inspection
FEI Bowen, XING Wenjie, LIU Daqian
Available online  , doi: 10.11999/JEIT260192
Abstract:
  Objective  In modern power inspection, the use of multiple Unmanned Aerial Vehicles (UAVs) for cooperative inspection is an efficient but challenging task. Existing multi-UAV path planning methods often have limited cooperative scheduling capability. They also fail to accurately capture the topological relationships among heterogeneous nodes, especially those between device nodes and charging stations under strict energy constraints. To address these limitations, this paper proposes Hierarchical Attention mechanism-based Path Planning for multi-UAV Inspection (HAPPI). The objective is to minimize the total flight distance of the UAV fleet while ensuring that all device nodes are inspected and all UAVs safely return to the base station under energy and visit-count constraints.  Methods  The multi-UAV power inspection problem is first formulated as a combinatorial optimization problem with energy constraints. It is then modeled within a Markov Decision Process (MDP) framework. To solve this problem, HAPPI adopts an encoder-decoder architecture with a customized hierarchical attention mechanism. The encoder uses a multi-level attention design to model three types of node relationships. Self-attention among device nodes is used to learn spatial proximity and visit-order preferences. Cross-attention between device nodes and charging stations is used to model energy supply-demand relationships. Self-attention among charging stations is used to explicitly capture the topological structure of the charging-station network. This hierarchical design enables the model to distinguish functional differences and dependencies among heterogeneous nodes. The decoder integrates the global graph embedding, the embedding of the last visited node, and the current remaining energy of the UAV to generate a context vector. A single-head attention mechanism is then used to compute compatibility scores for all candidate nodes. A masking strategy excludes infeasible nodes, including visited nodes, unreachable nodes, nodes that would prevent the UAV from reaching a charging station, and premature returns to the base station. The final node is selected from a probability distribution generated by softmax, which supports both greedy and sampling decoding strategies. The policy network is trained using reinforcement learning, and a baseline network is used to stabilize training. Policy-gradient optimization is used to minimize the expected total path length (Fig. 2).  Results and Discussions  Extensive simulations are conducted on three problem scales: T20C2, with 20 device nodes and 2 charging stations; T60C6, with 60 device nodes and 6 charging stations; and T100C10, with 100 device nodes and 10 charging stations. The training results show that HAPPI achieves faster convergence and a lower final cost than the baseline Attention Model (AM) and Heterogeneous Attention-based Deep Reinforcement Learning (HADRL) methods (Fig. 4). In the comprehensive performance comparison, HAPPI with sampling obtains the shortest total path lengths on T60C6 and T100C10, with values of 6.21 and 8.41, respectively. It outperforms five classical metaheuristic algorithms and two deep reinforcement learning baselines on these two larger-scale scenarios (Table 1). On T20C2, HADRL with sampling achieves the shortest path length, whereas HAPPI remains highly competitive. Overall, HAPPI reduces the total path length by approximately 12% on average compared with the baseline methods. The visualization results show that HAPPI generates routes with fewer route crossovers and a more balanced workload among UAVs, improving safety and efficiency (Fig. 6 and Fig. 7). The single-UAV path length distribution further confirms the superior load-balancing capability of HAPPI across all problem scales (Fig. 8).  Conclusions  This paper presents HAPPI, a hierarchical attention mechanism-based deep reinforcement learning method for cooperative path planning in multi-UAV power inspection scenarios with multiple charging stations. By explicitly modeling spatial relationships among device nodes, energy dependencies between device nodes and charging stations, and the internal topology of the charging-station network, HAPPI improves information aggregation and constraint satisfaction. Experimental results across different problem scales show that HAPPI achieves higher planning quality, greater computational efficiency, and stronger generalization than heuristic and learning-based comparison methods. Future work will extend this framework to multi-objective optimization that considers time, risk, and energy trade-offs, and will further validate the method using real-world inspection data.
Decoupled Learning for Long-tailed Oracle Bone Character Recognition Based on Adaptive Difficulty Sampling
SUN Junwei, GUAN Suyan, CHEN Xinyu, WANG Kun, CAI Yuanqiang
Available online  , doi: 10.11999/JEIT260327
Abstract:
  Objective  Oracle Bone Character (OBC) recognition is challenged by an extreme long-tailed distribution and substantial intra-class variation. Conventional deep learning methods are often dominated by head classes, whereas existing approaches tend to overfit tail classes or fail to account for differences in learning difficulty across classes. To address these limitations, a two-stage decoupled learning framework is proposed to improve the recognition of tail and difficult classes while preserving the discriminative capability of head classes.  Methods  The proposed framework decouples feature representation learning from classifier optimization. In the first stage, the backbone network is trained using a mixed data augmentation strategy that combines CutMix and RandAugment with Label-Distribution-Aware Margin (LDAM) loss to learn robust feature representations and alleviate the effect of intra-class variation. In the second stage, the backbone network is frozen, and only the classifier is optimized. An adaptive difficulty sampling strategy is proposed to dynamically assign sampling weights according to historical and current class-level training difficulty. The classifier is further optimized using a Class-Balanced LDAM (CBL) loss, which combines class-balanced weighting with LDAM to refine decision boundaries for long-tailed classification.  Results and Discussions  Experiments on the highly imbalanced OBC306 dataset demonstrate that the proposed method achieves an overall accuracy of 94.34% and an average class accuracy of 89.89%. Compared with the Inception-v4 baseline, the proposed method improves the average class accuracy by 19.61%. Comparisons with representative long-tailed OBC recognition methods further demonstrate superior overall performance. Comprehensive ablation studies verify the effectiveness of the mixed data augmentation strategy and the adaptive difficulty sampling strategy in improving the recognition of rare and difficult characters. Parameter sensitivity analysis and qualitative error analysis further confirm the robustness and effectiveness of the proposed framework.  Conclusions  The proposed two-stage decoupled learning framework effectively addresses long-tailed OBC recognition by balancing the learning priorities of head, tail, and difficult classes. The mixed data augmentation strategy improves feature robustness, whereas the adaptive difficulty sampling strategy and the Class-Balanced LDAM loss jointly optimize classifier learning and refine decision boundaries without degrading head-class recognition performance. The proposed framework provides an effective solution for the digital recognition of Oracle Bone Characters and offers technical support for low-resource ancient character recognition.
Construction of a DNA Strand Displacement Memristor and Its Filter Circuit Characteristics
WANG Yanfeng, CHEN Guanzhou, SUN Ce, SUN Junwei
Available online  , doi: 10.11999/JEIT260283
Abstract:
  Objective  Filter circuits are widely used in modern control and signal-processing systems for noise suppression and signal integrity enhancement. Conventional Resistor-Capacitor (RC) filters are widely applied, but their fixed parameters limit adaptability and miniaturization in emerging molecular and nanoscale computing platforms. To address these limitations, DNA Strand Displacement (DSD) technology is integrated with memristor theory to develop tunable multistable molecular filter circuits. This study aims to design and validate first- and second-order low-pass filter circuits based on the dynamic response and state-dependent behavior of a DSD-based memristor. The proposed filters are designed to improve frequency selectivity, parameter adaptability, and system stability compared with traditional filter architectures. This approach is intended for molecular signal processing, integrated biocircuits, and adaptive filtering systems that require compact size and reconfigurability.  Methods  The method consists of four stages. First, core DSD reaction modules, including sine, cosine, integration, addition, and multiplication modules, are designed to construct a programmable multistable memristor model. Second, square-wave and sinusoidal input signals are generated through DSD reactions to evaluate the memristor response under different frequencies and amplitudes. Third, the memristor is embedded into low-pass filter structures to construct first- and second-order DSD-based memristor filter circuits. Fourth, simulations are performed using Visual DSD for molecular dynamics analysis and MATLAB for circuit-level analysis. Circuit performance is evaluated using transfer functions, Nyquist plots, Bode diagrams, and time-domain comparisons with classical RC filters. This combined simulation strategy verifies both molecular feasibility and circuit functionality.  Results and Discussions  The DSD-based memristor exhibits multistable behavior and converges to six stable equilibrium points under different initial conditions (Fig. 8). Its hysteresis characteristics further confirm the state-dependent memory behavior of the designed molecular memristor (Fig. 7). The first-order DSD-based memristor filter circuit provides stable attenuation for square-wave and sinusoidal input signals. Its output amplitudes are consistently higher than those of the traditional RC filter across the tested frequencies (Table 3). The second-order DSD-based memristor filter circuit further reduces signal delay and improves stability, especially under high-frequency inputs (Table 4). Frequency-response analyses show that the cutoff frequency can be dynamically tuned by adjusting DSD reaction rates and initial concentrations (Figs. 9 and 11). Time-domain simulations further confirm the filtering performance of the first- and second-order circuits (Figs. 10 and 12). Reliability analysis indicates that lower initial copy numbers increase stochastic molecular noise, whereas higher initial copy numbers make the output distribution closer to the deterministic response and improve the probability of successful filtering. These results verify the feasibility of DSD-memristor integration for adaptive molecular filtering.  Conclusions  A DSD-based memristor with multistable characteristics and its corresponding first- and second-order low-pass filter circuits are designed and validated. Compared with traditional RC architectures, the proposed filters show improved output stability, parameter tunability, and frequency adaptability. By combining DSD technology with memristor theory, this study provides a reconfigurable molecular-scale filtering framework for signal-processing applications. The results provide a basis for future work on adaptive molecular circuits, intelligent filtering, and nanoelectronic system design. Further studies should focus on experimental validation, real-time tuning strategies, sequence optimization, anti-interference design, signal amplification, and circuit integration.
Spatial-domain Anti-jamming for Unmanned Systems Under Limited Prior Information
PAN Zihao, ZHANG Bangning, ZHEN Pan, ZHU Bowen, WANG Ning, GUO Daoxing
Available online  , doi: 10.11999/JEIT260296
Abstract:
  Objective  Unmanned systems play an increasingly important role in emergency response, public safety, intelligent transportation, and other mission-critical applications. Reliable communications in complex electromagnetic environments are essential for autonomous operation. However, communication links are directly exposed to open, non-cooperative electromagnetic environments and are therefore vulnerable to intentional jamming and unintentional interference. In practical scenarios, prior information regarding the desired signal, jamming sources, and multipath propagation is often unavailable, substantially degrading the performance of conventional spatial-domain anti-jamming methods. To address this challenge, this paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information.  Methods  The proposed method first applies a spatial smoothing algorithm to the received signals to decorrelate coherent multipath components. Capon spatial spectrum estimation is then performed to detect the Direction Of Arrival (DOA) of potential incident signals. Spectrum peaks corresponding to individual incident signals are subsequently identified. A Covariance Matrix Reconstruction (CMR)-based beamforming algorithm is then applied by traversing all detected spectrum peaks to sequentially extract the signal associated with each peak, thereby separating the mixed signals. After signal separation, a signal classification method based on spectral similarity and time delay is employed. Kullback-Leibler (KL) divergence between the spectrum of each separated signal and the reference spectrum is calculated to identify jamming signals. The remaining communication signals are further classified into direct-path and multipath signals according to their relative time delays. Finally, different processing strategies are applied according to the identified signal type. Specifically, multipath signals are either suppressed as interference or coherently combined with the direct-path signal after time-delay and phase alignment.  Results and Discussions  Two simulation scenarios, including jamming only and combined jamming and multipath, are designed to evaluate the proposed method in terms of the output Signal-to-Interference-plus-Noise Ratio (SINR), beam pattern, Bit Error Rate (BER), and Error Vector Magnitude (EVM). Simulation results demonstrate that, under the jamming-only scenario, the proposed method achieves performance close to the theoretical optimum. The output SINR increases with the input Signal-to-Noise Ratio (SNR) at a fixed Jamming-to-Signal Ratio (JSR) (Fig. 3(a)) and remains nearly unchanged as JSR increases at a fixed SNR (Fig. 3(b)), indicating stable jamming suppression capability. The recovered time-domain waveform and spectrum remain highly consistent with the transmitted signal (Fig. 4). The BER curve nearly overlaps that of the optimal beamformer (Fig. 5). At \begin{document}$ {E}_{\rm b}/{N}_{0}=10\;{\mathrm{dB}} $\end{document}, the recovered Quadrature Phase-Shift Keying (QPSK) constellation closely matches the ideal constellation, achieving an EVM of –11.52 dB (Fig. 6). Under simultaneous jamming and multipath conditions, the proposed framework flexibly suppresses or exploits multipath signals. Compared with multipath suppression, multipath utilization further improves both the output SINR and BER (Fig. 7(a) and Fig. 7(b)). The corresponding beam pattern forms a beam toward the multipath direction rather than a null, demonstrating effective multipath exploitation (Fig. 7(c)).  Conclusions  This paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information. Using only the received mixed signals, the proposed framework estimates the directions of arrival, separates incident signals, and classifies them as direct-path, multipath, or jamming signals. Appropriate suppression or preservation strategies are then applied according to the identified signal type. Therefore, the framework flexibly suppresses or exploits multipath signals while preserving the direct-path signal and mitigating jamming. Simulation results demonstrate the effectiveness of the proposed method in terms of output SINR and demodulation accuracy, confirming reliable jamming suppression and communication performance even when prior information regarding the desired signal, jamming sources, and multipath propagation is unavailable. Future work will investigate the effects of array perturbations, intelligent jamming, and heterogeneous communication modes on the proposed framework and extend it to more complex unmanned-system communication environments.
A Nested Multi-scroll Memristive Hopfield Neural Network and Its Hardware Implementation
WANG Zhe, WAN Qiuzhen, ZHOU Pan, RAO Huhui
Available online  , doi: 10.11999/JEIT260516
Abstract:
  Objective  In recent years, memristors have been employed to emulate neuronal synapses with dynamically adjustable synaptic weights, enabling the construction of Memristive Hopfield Neural Networks (HNNs). Compared with conventional HNNs, Memristive HNNs more accurately reproduce the nonlinear dynamical behavior of biological neural systems. Multi-scroll attractors have attracted considerable attention in secure communication because of their complex topological structures and strong state-space ergodicity. However, previous studies have primarily focused on conventional multi-scroll attractors with single structural patterns, whereas multi-scroll attractors with special structures remain largely unexplored. Therefore, this paper proposes a nested multi-scroll Memristive HNN system that generates nested multi-scroll attractors, thereby overcoming the limitations of conventional single-structure multi-scroll attractors.  Methods  A Four-Dimensional (4D) Memristive HNN system is constructed from a three-neuron HNN by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of neuron 2. Equilibrium-point and stability analyses are performed to investigate the regulatory effects of the Memristive self-connected synapse coupling strength and system initial conditions on the system dynamics. The number of multi-scroll attractors is regulated by adjusting the memristor control parameters. Building on this framework, a Multi-level Logic Pulse current (IMLP) is introduced to construct a nested multi-scroll Memristive HNN system. The proposed system generates nested multi-scroll attractors with enhanced dynamical complexity. Finally, the MATLAB numerical simulation results are validated through Multisim circuit simulations and Field-Programmable Gate Array (FPGA)-based hardware experiments.  Results and Discussions  The results demonstrate that regulating the Memristive self-connected synapse coupling strength enables the proposed 4D Memristive HNN system to exhibit period-doubling bifurcations and chaotic behavior, as illustrated by the bifurcation diagrams and Lyapunov exponent spectra (Fig. 3). Various types of coexisting attractors are generated under different coupling strengths (Fig. 4). By adjusting the memristor control parameters, multi-scroll attractors with different numbers of scrolls are generated through one-directional extension (Figs. 58). After the introduction of the IMLP, the proposed nested multi-scroll Memristive HNN system generates nested multi-scroll attractors while preserving the controllable scroll-number extension property (Figs. 810). Spectral Entropy (SE) analysis demonstrates that the IMLP increases the dynamical complexity of the proposed system compared with the original 4D Memristive HNN system (Figs. 9 and 10). The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments confirms the physical realizability of the proposed nested multi-scroll Memristive HNN system (Figs. 1214).  Conclusions  A 4D Memristive HNN system is constructed by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of a three-neuron HNN. Equilibrium-point and stability analyses reveal the regulatory effects of the Memristive self-connected synapse coupling strength and the evolution of coexisting attractors associated with different initial conditions. The results show that the system enters chaos through the period-doubling route to chaos and generates single-scroll and double-scroll chaotic attractors. The number of multi-scroll attractors is continuously increased by adjusting the memristor control parameters. Furthermore, introducing the IMLP produces a nested multi-scroll Memristive HNN system capable of generating nested multi-scroll attractors with increased dynamical complexity. The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments validates the physical realizability of the proposed nested multi-scroll Memristive HNN system.
Research on Ka-band Enhanced Active Load Modulation Ultra-wideband High-efficiency Doherty Power Amplifier
YANG Lin, YAN Chengyu, WANG Yanping, ZHANG Ming, WANG Baozhu, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260514
Abstract:
  Objective  The Ka-band has become a key frequency band for satellite communications, placing stringent requirements on the millimeter-wave power amplifier, a core component of the transmitter, to provide high efficiency, compact size, and broadband operation. To maximize spectral efficiency, millimeter-wave satellite communication signals typically exhibit a high Peak-to-Average Power Ratio (PAPR), making high back-off efficiency particularly important. Although the Doherty Power Amplifier (DPA) is widely adopted because of its high efficiency under power back-off conditions, its operating bandwidth is inherently limited. In addition, both saturated efficiency and back-off efficiency degrade substantially at millimeter-wave frequencies. Therefore, extending the operating bandwidth while maintaining high efficiency remains a major challenge for millimeter-wave DPAs used in satellite communication transmitters.  Methods  An ultra-wideband enhanced active load modulation technique is proposed to overcome the trade-off between bandwidth and back-off efficiency in DPAs. The proposed method achieves optimal load impedance modulation for both the carrier and peaking amplifiers over an ultra-wide frequency range by introducing an Impedance Tunable Bias Network (ITBN) and a dual-drive impedance control mechanism. These techniques improve load modulation while extending the load modulation bandwidth, thereby enhancing both efficiency and bandwidth in the millimeter-wave DPA architecture. Furthermore, a broadband phase-compensation technique is integrated into an unequal power division network to achieve sufficient load modulation across the entire operating band while accurately compensating for the phase difference between the carrier and peaking paths. The proposed ultra-wideband phase-compensated power division network further extends the high-efficiency operating bandwidth while reducing the chip area.  Results and Discussions  To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was designed and fabricated using a 0.15-μm GaN process. Across the 24~33 GHz frequency band, corresponding to a relative bandwidth of 31.6%, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB with a gain flatness of ±1 dB. The measured saturated output power is 28.9~31.0 dBm, with a Power-Added Efficiency (PAE) of 25.2%~33.5% at saturation and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating over similar frequency bands, the proposed design achieves the highest reported small-signal gain and saturated PAE while maintaining high saturated output power and high 6 dB back-off PAE. Furthermore, it occupies the smallest chip area among reported three-stage DPA MMICs.  Conclusions  An ultra-wideband enhanced active load modulation method is proposed to achieve sufficient active load modulation through multi-frequency impedance tuning across the entire operating band. A novel ultra-wideband phase-compensated unequal power division network is also proposed to reduce the chip area while maintaining accurate phase compensation. To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was fabricated using a 0.15-μm GaN process. Measurement results demonstrate that, over the 24~33 GHz frequency band, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB, a saturated output power of 28.9~31.0 dBm, a saturated PAE of 25.2%~33.5%, and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating in similar frequency bands, the proposed design achieves a maximum relative bandwidth of 31.6%, the highest reported small-signal gain and saturated PAE, high saturated output power, and high 6 dB back-off PAE, while occupying the smallest chip area among reported two- and three-stage DPA MMICs. These measurement results validate the proposed method and demonstrate its strong potential for millimeter-wave satellite communication transmitters.
Iterative Parameter Estimation Method for Energy Detection Threshold in Ambient Backscatter
QU Wenfeng, YE Yinghui, SHI Liqin, LU Guangyue
Available online  , doi: 10.11999/JEIT260418
Abstract:
  Objective  In Ambient Backscatter Communication (AmBC) systems, a low-complexity Energy Detector (ED) is commonly employed at the reader to recover symbols transmitted by the tag. The detection performance of ED depends strongly on the accurate setting of the detection threshold, which is determined by the average received signal power corresponding to tag symbols “1” and “0”. Existing parameter estimation methods assume that the two symbols are transmitted with equal probability and therefore divide the sorted received signal power samples into two equal groups. However, because the number of transmitted symbols is finite, the actual numbers of symbols “1” and “0” are generally unequal. Therefore, equal partitioning introduces sample misclassification, causing the estimated threshold to deviate from its optimal value and reducing detection performance. To address this limitation, an iterative threshold parameter estimation method is proposed to reduce the parameter estimation bias caused by sample misclassification and improve the accuracy of detection threshold estimation.  Methods  An iterative threshold parameter estimation method is proposed to overcome the sample misclassification introduced by conventional sorting-based grouping. Because the initial detection threshold obtained by the sorting-based grouping method provides reliable decisions for most received samples, these initial decisions are used as the basis for sample reclassification. The received signal samples are then reclassified to iteratively update the threshold parameters, progressively refining the detection threshold. The proposed method is evaluated through simulations under three representative ambient radio-frequency source conditions: complex Gaussian, Phase-Shift Keying (PSK), and Quadrature Amplitude Modulation (QAM) sources.  Results and Discussions  Simulation results show that, over a wide range of Signal-to-Noise Ratio (SNR) values, the proposed iterative method substantially reduces the Bit Error Rate (BER) compared with the conventional sorting-based grouping method and approaches the theoretical lower bound obtained with perfect parameter estimation. At a given SNR, the proposed method improves BER by approximately 0.1, 1.5, and 1.3 orders of magnitude under complex Gaussian, PSK, and QAM sources, respectively (Fig. 2). These results demonstrate that the proposed iterative method effectively corrects sample misclassification and reduces the performance loss caused by parameter estimation bias. Moreover, most of the performance gain is achieved after only one iteration, indicating rapid convergence with minimal additional computational overhead. Under different numbers of sampling points, BER improvements of approximately 0.5, 1.6, and 1.1 orders of magnitude are achieved under complex Gaussian, PSK, and QAM sources, respectively (Fig. 3). These results indicate that using the initial decisions for sample reclassification effectively reduces the estimation bias introduced by fixed equal partitioning, thereby improving detection performance under limited-sample conditions. Under different Relative Channel Difference (RCD) values, BER improvements of approximately 0.56 and 0.7 orders of magnitude are achieved under complex Gaussian and QAM sources, respectively (Fig. 4). As the RCD increases, the separation between the received signal power distributions becomes more pronounced, improving the accuracy of the initial decisions and enabling more reliable sample reclassification. This positive feedback process further refines the parameter estimates and improves detection performance.  Conclusions  An iterative threshold parameter estimation method is proposed to address the sample misclassification introduced by conventional sorting-based grouping in ED. The proposed method uses the initial decisions to reclassify the received signal samples and iteratively update the threshold parameters. In addition, closed-form expressions for the detection threshold and BER under QAM sources are derived. Simulation results demonstrate that the proposed method effectively reduces parameter estimation bias with only one iteration while maintaining robust performance under limited-sample and varying channel conditions. Significant BER improvements are achieved with minimal additional computational overhead, making the proposed method well suited for practical, high-reliability AmBC systems.
LLMA-GCN: A Semantically Enhanced Hierarchical Spatiotemporal Graph Convolutional Network for Skeleton-based Action Recognition
JIA Guimin, ZHOU Xilong
Available online  , doi: 10.11999/JEIT260154
Abstract:
  Objective  Current methods that introduce the Large Language Model (LLM) into skeleton-based action recognition face three main limitations. Semantic guidance remains decoupled from spatial topology learning, temporal modeling lacks hierarchical semantic support, and traditional classification paradigms have limited generalization. To address these issues, this paper proposes LLMA-GCN, a semantically enhanced graph convolutional framework. The proposed framework integrates LLM-derived semantic prior knowledge with graph convolution to improve spatiotemporal feature learning and action classification.  Methods  LLMA-GCN uses visual skeleton data and semantic inputs in a dual-branch framework. Frozen LLMs and prompt engineering are used to precompute the joint semantic adjacency matrix and action text prototypes. The framework consists of three main components: an LLM-based hybrid graph topology learning strategy, a hierarchical visual sequence encoder based on the LLM Refinement Block (LRB), and an action text-prototype-guided decision learning mechanism. These components enable semantic guidance in graph topology learning, hierarchical spatiotemporal feature extraction, and visual-text alignment.  Results and Discussions  Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD I show that LLMA-GCN achieves competitive or superior performance compared with state-of-the-art methods. Ablation studies confirm the key roles of the hybrid graph topology, the LRB, and the action text-prototype-guided decision learning mechanism. Model-complexity analysis further indicates the potential of the proposed framework for practical application.  Conclusions  By fusing the joint semantic adjacency matrix and the physical adjacency matrix, LLMA-GCN enables action semantics to directly guide graph convolution and improves semantic perception. The LRB embeds semantic information into hierarchical spatiotemporal feature extraction, which strengthens the modeling of complex actions. The action text-prototype-guided decision learning mechanism further shifts skeleton-based action recognition from purely visual classification to text-prototype-guided visual-text alignment. Overall, LLMA-GCN provides a robust and generalizable framework for skeleton-based action recognition through deep fusion of visual and semantic features.
Low-Complexity Phase Ambiguity Resolution DOA EstimationAlgorithm for Composite Hierarchical Receiving Array Structure
CHEN Yiwen, DONG Yangze, CHEN Xiahua, LING Wenchang, XIONG Yiwen
Available online  , doi: 10.11999/JEIT260447
Abstract:
  Objective  Direction Of Arrival (DOA) estimation is a key technique for sonar target localization. As the demand for high-precision DOA estimation in complex environments continues to increase, the number of array elements used for estimation is steadily growing, leading to massive arrays. Although larger arrays improve DOA estimation accuracy and resolution, they also impose a substantial computational burden on conventional DOA estimation algorithms. To address this issue, a low-complexity composite hierarchical receiving array structure is constructed, and two fast phase ambiguity resolution algorithms are proposed: Composite HierArchical Global Nearest-Neighbor Matching (CHA-GNNM) and Composite HierArchical Cross-Correlation Covariance Merging (CHA-CCM).  Methods  The CHA-GNNM algorithm constructs multiple candidate solution sets by exploiting the auto-covariance and cross-covariance relationships among the subarrays within each group. The true solution in each candidate solution set is identified through nearest-neighbor matching based on source consistency, and the final DOA estimate is obtained through multilevel coherent combining. This approach achieves phase ambiguity resolution and angle matching with relatively low computational cost. However, because the correlation information among all array elements is not fully exploited, some estimation performance is sacrificed. To improve DOA estimation performance, the CHA-CCM algorithm reorganizes the composite hierarchical structure into evenly partitioned groups, which are regarded as several large subarrays. Multiple large candidate solution sets are first constructed from the cross-correlation relationships among these groups. Each group is then divided into multiple small subarrays, from which additional candidate solution sets are generated using the corresponding auto-covariance and cross-covariance relationships. A coarse DOA estimate is obtained through coprime clustering, followed by a more accurate initial DOA estimate derived from the small candidate solution sets. This initial DOA estimate is subsequently used to eliminate spurious solutions from the large candidate solution sets, yielding the final DOA estimate. Combined with a low-complexity covariance block-processing strategy, this approach avoids computationally expensive operations while improving DOA estimation accuracy.  Results and Discussions  Simulation results demonstrate that both proposed algorithms substantially reduce the computational burden as the number of array elements increases, while effectively achieving phase ambiguity resolution through the proposed composite hierarchical receiving array structure (Fig. 5). Compared with the conventional Root-MUSIC algorithm, CHA-GNNM achieves coarse DOA estimation with nearly four orders of magnitude lower computational complexity (Fig. 7), making it suitable for applications with stringent real-time requirements. In contrast, CHA-CCM requires only a modest increase in computational cost (Fig. 7) while achieving DOA estimation performance close to the Cramér-Rao Lower Bound (CRLB) above a certain signal-to-noise ratio threshold (Fig. 6). Therefore, a favorable balance is achieved between DOA estimation accuracy and computational complexity.  Conclusions  To address the rapid increase in computational complexity associated with massive arrays, a composite hierarchical receiving array structure is constructed for efficient DOA estimation. By hierarchically grouping the array elements, the proposed structure provides a new framework for low-complexity DOA estimation. Based on this structure, two fast DOA estimation algorithms are developed. Both algorithms achieve effective phase ambiguity resolution with low computational complexity by exploiting the structural differences among array groups and the consistency of observations from the same source across different groups, thereby enabling rapid DOA estimation. CHA-GNNM primarily exploits the phase relationships among subarrays to perform phase ambiguity resolution and angle matching through a simple computational procedure, making it suitable for applications requiring high computational efficiency and real-time processing. Because the cross-correlation information among all array elements is not fully exploited, some estimation performance is reduced under challenging signal conditions. To overcome this limitation, CHA-CCM reorganizes the composite hierarchical receiving array into evenly partitioned groups while preserving the low-complexity advantage of the hierarchical structure. Group-level cross-correlation information is further exploited so that the intrinsic relationships among different groups are more fully utilized. In addition, the signal processing procedure is simplified by eliminating unnecessary computational steps, thereby improving the robustness and accuracy of DOA estimation while maintaining manageable computational complexity. Compared with CHA-GNNM, CHA-CCM incurs only a small increase in computational cost and achieves a better balance between computational complexity and DOA estimation performance. Overall, the proposed composite hierarchical receiving array structure and the two fast DOA estimation algorithms provide an effective solution for efficient DOA estimation in massive arrays. CHA-GNNM is more suitable for applications with stringent real-time requirements, whereas CHA-CCM is better suited for applications requiring higher DOA estimation accuracy and robustness. The proposed structure achieves efficient phase ambiguity resolution and accurate DOA estimation and provides both theoretical significance and practical value for engineering applications of massive array signal processing.
Off-grid Blind Near-Field Integrated Sensing And Communication: Algorithm Design and Lower Bound
YUAN Zhengdao, GUO Qinghua, HUANG Chongwen, GAO Dawei, MEI Fengtong, LIAO Guisheng
Available online  , doi: 10.11999/JEIT260404
Abstract:
  Objective  With the widespread deployment of extra-large-scale antenna arrays in 6G networks, user terminals are increasingly located in the near-field region. Existing Near-Field Integrated Sensing And Communication (NF-ISAC) algorithms face key challenges, including off-grid power leakage, severe model mismatch, and strong pilot dependence. These limitations make them unsuitable for low-overhead, high-performance 6G transmission. This paper aims to design an off-grid blind NF-ISAC algorithm and derive the theoretical performance bound for near-field sensing.  Methods  To overcome the limitations of analytical geometric steering vectors and adapt to more accurate electromagnetic propagation characteristics without closed-form expressions, an amplitude-phase separation method is first proposed. This method decomposes the nonlinear near-field steering vector into amplitude and phase terms, enabling high-precision characterization of the steering vector using a single-hidden-layer neural network. Second, the NF-ISAC problem is formulated as a constrained matrix factorization problem, and a corresponding factor graph model is constructed. The trained neural network is embedded into the factor graph as a function node. Message passing through the embedded neural network is then achieved, enabling joint blind coordinate sensing, channel estimation, and signal detection in a pilot-free manner. Finally, the Cramér-Rao Lower Bound (CRLB) for multi-user near-field joint distance and angle sensing in polar coordinates is derived based on the neural-network-fitted steering vector.  Results and Discussions  Extensive Monte Carlo simulations are conducted to evaluate the performance of the proposed algorithm. The simulation results show that the proposed algorithm achieves millimeter-level position sensing. Compared with existing mainstream algorithms, it improves both communication Bit Error Rate (BER) and sensing accuracy. The proposed algorithm achieves a 2~3 dB gain in sensing accuracy over the state-of-the-art near-field off-grid algorithm, and its performance is closest to the derived theoretical CRLB. These results indicate that the proposed algorithm effectively mitigates off-grid power leakage and model mismatch.  Conclusions  The proposed off-grid blind NF-ISAC algorithm overcomes the pilot dependence and model mismatch of existing NF-ISAC schemes. It achieves integrated high-precision sensing and reliable communication for near-field users in a pilot-free manner. The derived CRLB provides a theoretical benchmark for evaluating the sensing performance of NF-ISAC systems. This work provides technical support for the design of 6G NF-ISAC systems.
A Lightweight Spatial-Spectral Dual-Branch Transformer Network for Classifying Polarized White Blood Cell Hyperspectral Images
YANG Yushi, YAN Jiaxuan, XIE Yi, QIU Lijia, HUANG Danfei
Available online  , doi: 10.11999/JEIT260124
Abstract:
  Objective  White Blood Cell (WBC) classification and morphology are essential in routine blood analysis and provide important information for disease diagnosis and health assessment. Current WBC classification mainly relies on hematology analyzers and manual microscopic examination. Hematology analyzers cannot acquire cellular images, limiting classification accuracy when abnormal cellular characteristics are present, whereas manual microscopic examination depends on operator experience and is susceptible to human error. Although automated WBC classification based on deep learning and computer vision has attracted considerable attention, methods using stained images are sensitive to staining conditions and image quality. In addition, conventional hyperspectral imaging has limited ability to distinguish WBC subtypes with highly similar morphological and spectral characteristics. To address these limitations, this study combines Polarized Hyperspectral Imaging (PHSI) with deep learning and proposes a lightweight classification framework for polarized hyperspectral WBC images, providing an efficient and reliable approach for clinical decision support.  Methods  A polarized hyperspectral microscopic imaging system is established to acquire images at multiple polarization angles. Based on Stokes Vector Theory, polarization parameters are calculated to generate multidimensional data cubes containing both light intensity and polarization-state information. Degree of Linear Polarization (DOLP) images are then computed to construct a PHSI dataset of WBCs. To exploit the multidimensional characteristics of PHSI, a Lightweight Spatial-Spectral Dual-Branch Transformer Network (LSDBT) is proposed. After preprocessing, the input data are fed into a dual-branch feature extraction module that separately extracts local spatial features and joint spatial-spectral features. An adaptive scaling factor is introduced to fuse the two feature streams and balance their contributions, enabling effective utilization of the multidimensional information contained in PHSI. A lightweight Swin Transformer backbone performs effective global feature modeling while reducing computational complexity. Global average pooling and a fully connected layer are used for classification. Model performance is evaluated using Overall Accuracy (OA), Precision, Recall, Specificity, and F1-Score. Ablation studies, comparative experiments, and feature visualization are conducted to validate the proposed method.  Results and Discussions  The DOLP spectra and PHSI visualizations of monocytes, lymphocytes, and neutrophils (Figures. 5 and 6) demonstrate different polarization characteristics that reflect the selective absorption and scattering of light by their internal structures. Compared with conventional intensity images, PHSI improves image contrast and provides additional polarization information that enhances discrimination among WBC types. The proposed LSDBT achieves an OA of 99.29% on the test set, with consistently high classification performance across all cell categories (Table 1). Analysis of the adaptive scaling factor shows that classification performance first improves and then decreases slightly as the scaling factor increases, with the optimal value of 3 providing the best balance between spatial and spatial-spectral features (Figure. 8). Ablation experiments (Tables 2 and 3) demonstrate that the dual-branch feature extraction module substantially improves classification performance, whereas the lightweight design greatly reduces computational complexity and model parameters with only a marginal reduction in accuracy. Compared with conventional hyperspectral imaging, the PHSI dataset achieves higher classification accuracy with all evaluated classifiers, indicating that polarization information provides complementary physical features that improve discrimination among WBC types (Table 4). Comparisons with representative methods show that LSDBT achieves the best overall classification performance across multiple evaluation metrics (Table 5). Furthermore, t-SNE visualization (Figure. 9) shows compact intra-class distributions and clear separation among different cell types, confirming the strong discriminative capability of the learned features. Although LSDBT does not have the lowest computational cost among the compared methods, it achieves the best balance between classification performance, model size, and computational efficiency (Table 6).  Conclusions  This study proposes a lightweight dual-branch Transformer network for polarized hyperspectral WBC classification. To the best of our knowledge, this is the first study to combine PHSI with deep learning for WBC classification. Comparative experiments with conventional hyperspectral imaging validate the superiority of PHSI for WBC classification. The proposed LSDBT integrates spatial and spatial-spectral information through a dual-branch feature extraction module and performs efficient global feature modeling using a lightweight Swin Transformer backbone. The network maintains high classification performance while substantially reducing computational complexity and model parameters. These results demonstrate that LSDBT provides an accurate and computationally efficient solution for automated WBC classification and supports the application of PHSI in cellular microscopic analysis and clinical auxiliary diagnosis.
Research on Adaptive Hybrid Beamforming Method for Massive MIMO LEO Satellite Communication Systems
XIAN Yongju, HUANG Xiaolong, XING Zhitong, LI Yun
Available online  , doi: 10.11999/JEIT260458
Abstract:
  Objective  With the growing demand for high-capacity and high-spectral-efficiency transmission in LEO satellite communications, massive MIMO has become a promising enabling technology. However, conventional fully digital beamforming is difficult to implement in practice due to the strict constraints on power consumption, hardware complexity, and payload cost of satellite platforms. Although partially connected hybrid beamforming can reduce hardware complexity, the conventional fixed subarray structure lacks flexibility and suffers from performance degradation, especially under low-resolution PSs constraints. To address this issue, this paper investigates adaptive antenna-RF chain mapping for hybrid beamforming design in LEO satellite massive MIMO systems.  Methods  This paper first establishes a system model for LEO satellite multi-user downlink massive MIMO hybrid beamforming and formulates a joint optimization problem with the objective of maximizing spectral efficiency. Considering the constant-modulus discrete phase constraints of low-resolution PSs, antenna-RF chain connection constraints, and transmit power constraint, the resulting problem is highly non-convex. To solve it, the original problem is transformed into an equivalent WMMSE formulation, and auxiliary variables are introduced to decouple the coupled variables. Based on this reformulation, a double-layer iterative optimization framework is developed by combining the BCD method and the PDD method. For adaptive antenna-RF chain mapping, the mapping problem is reformulated as a capacity-constrained linear assignment problem, and an optimal adaptive mapping method based on the Hungarian algorithm is proposed. Furthermore, to reduce the computational burden in large-scale antenna array scenarios, a low-complexity adaptive mapping method based on antenna priority sorting is developed.  Results and Discussions  Simulation results show that the proposed methods exhibit good convergence behavior. Specifically, the spectral efficiency increases rapidly in the initial iterations and then gradually converges, while the constraint violation decreases continuously, confirming the effectiveness of the proposed iterative optimization framework (Fig. 3). In terms of spectral efficiency, the proposed adaptive mapping methods consistently outperform the conventional fixed subarray and the existing greedy dynamic subarray scheme over different transmit powers and antenna scales. Among them, the Hungarian-based method achieving better spectral efficiency, whereas the antenna priority sorting-based method attains near-optimal performance with significantly reduced computational complexity (Figs. 4 and 5). As the number of PSs quantization bits increases, the system performance gradually approaches that of the continuous-PSs case, demonstrating the effectiveness of the proposed low-resolution PSs-based design (Fig. 6). In terms of energy efficiency, the proposed methods also outperform the conventional fully digital beamforming, fully connected, fixed subarray, and greedy dynamic subarray hybrid beamforming under different transmit powers and antenna scales (Figs. 7 and 8).  Conclusions  This paper proposes an adaptive hybrid beamforming design for LEO satellite massive MIMO systems under low-resolution PS constraints. By combining WMMSE reformulation with a PDD-BCD based optimization framework, joint design of digital precoding, analog precoding, and adaptive antenna-RF chain mapping is achieved. Simulation results demonstrate that the proposed methods provide superior performance in both spectral efficiency and energy efficiency. In particular, the Hungarian-based method provides better system performance, while the antenna priority sorting-based method achieves a favorable trade-off between performance and computational complexity. The proposed design provides an effective solution for high-performance hybrid beamforming in hardware-constrained LEO satellite massive MIMO systems.
FedFACO: Personalized Federated Learning Method Based on Fisher Information Matrix for Adaptive Aggregation and Client Collaborative Optimization
JIANG Wei-Jin, LIU Zhi-Hua, CUI Xin-Yu, XU Yu-Sheng, CHEN Shen-You, HU Jia-Long
Available online  , doi: 10.11999/JEIT260344
Abstract:
  Objective  Non-IID data heterogeneity remains one of the major challenges in personalized federated learning, as it often leads to inconsistent local optimization directions, insufficient global knowledge transfer, and degraded model personalization performance. To address these issues, this paper proposes FedFACO, a personalized federated learning method based on Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The proposed method aims to improve the adaptability of federated models to heterogeneous client distributions while maintaining effective knowledge sharing across clients. By introducing an adaptive aggregation mechanism and a collaborative optimization strategy, FedFACO provides a more principled way to balance global generalization and local personalization, which is particularly important in complex Non-IID federated environments.  Methods  FedFACO consists of two key components. First, an adaptive aggregation (AA) mechanism is employed to dynamically adjust fusion weights between global and local models based on client-specific states, generating personalized initializations aligned with local data. Second, a collaborative optimization (CO) mechanism is introduced, combining feature alignment with FIM-based client weighting. This enhances useful global knowledge transfer and suppresses low-quality updates. The FIM is utilized to quantify the information contribution of each update, ensuring reliable aggregation. The method is evaluated on MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet under Non-IID settings, and compared with representative baselines. Convergence behavior, dropout robustness, and sensitivity to low-quality updates are also examined.  Results and Discussions  Experimental results demonstrate that FedFACO consistently outperforms competitive baseline methods across all four benchmark datasets, achieving an average accuracy improvement of approximately 3.1% over mainstream approaches (Fig. 1, Table 2). On the more challenging Tiny-ImageNet dataset, FedFACO reduces the total training time required to reach convergence by approximately 4.8% compared with the best baseline (Table 3). Ablation studies confirm that performance is substantially improved by both AA and CO mechanisms, with their joint application yielding optimal accuracy (Table 6). Furthermore, FIM-guided weighting is shown to accurately quantify contribution quality in client dropout and asynchronous scenarios, significantly enhancing aggregation reliability (Fig. 2Fig. 3). Superior robustness is also demonstrated in malicious client scenarios (Fig. 4).  Conclusions  This paper presents FedFACO, a personalized federated learning method for Non-IID environments via Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The method effectively balances global knowledge sharing and local personalization while enhancing training stability and robustness under heterogeneous client participation. Experimental results validate its effectiveness and superiority in accuracy, convergence efficiency, and robustness. Future work will focus on lightweight Fisher information approximations and adaptive triggering strategies to reduce computational overhead, as well as integration with privacy-preserving and security-defense mechanisms for deployment in resource-constrained and high-security environments.
A High-Parallelism Simulated Adiabatic Bifurcation Processor for Combinatorial Optimization Problems
LI Renlong, HAO Xin, CHEN Zhuojun, DING Ding
Available online  , doi: 10.11999/JEIT260780
Abstract:
  Objective  Combinatorial optimization problems (COPs) are widely encountered in fields such as network optimization, autonomous driving path planning, VLSI design, and computational biology. The solution space of these problems grows exponentially with problem size, making it impossible for traditional von Neumann architectures (e.g., CPUs) to find high-quality solutions within a feasible time. Quantum-inspired Ising machines have emerged as a promising computing paradigm for accelerating COP solving. However, existing CMOS-based Ising machines still face significant challenges in simultaneously achieving high speed, high energy efficiency, and high accuracy. Discrete-time Ising machines often require thousands of iterations and rely heavily on random number generators, leading to large chip area and long solving times. Continuous-time Ising machines suffer from poor solution quality, frequently getting trapped in local minima. To address these limitations, this paper designs and fabricates a high-parallelism simulated adiabatic bifurcation application-specific processor for combinatorial optimization problems in 65 nm CMOS technology.  Methods  The proposed processor adopts the simulated adiabatic bifurcation (SAB) algorithm, which is inspired by quantum adiabatic optimization. Unlike simulated annealing, SAB does not require Gibbs sampling or random number generators to escape local minima. The algorithm models each spin as a nonlinear oscillator and solves a set of ordinary differential equations to simulate the adiabatic evolution of a classical nonlinear Hamiltonian system exhibiting bifurcation phenomena. The processor builds a hardware architecture that supports fully connected spin topologies and leverages the inherent fully parallel spin update characteristic of the SAB algorithm. To achieve bubble-free iterative computation, a three-stage pipelined spin update strategy is proposed, dividing the update process into coupling coefficient access, momentum update, and position update. To reduce the storage overhead introduced by the fully connected coupling matrix, a folded coupling coefficient storage array is designed, exploiting matrix symmetry to eliminate redundant storage. The momentum update unit and position update unit are implemented using 8-bit fixed-point arithmetic (2 bits for integer, 6 bits for fractional part) to achieve SAB evolution with low hardware overhead. The chip is fabricated in 65 nm CMOS technology, occupying an area of 0.47 × 1.25 mm2 and operating at a 200 MHz clock frequency and 1 V supply voltage.  Results and Discussions  The chip achieves a total power consumption of only 14 mW, with the spin evolution module consuming 58% of the total power. The folded coupling coefficient storage array reduces storage area by 52% compared to full matrix storage, and the rectangular restructuring avoids irregular shapes in physical layout, reducing routing congestion and layout voids (Fig.6). The parallel loading access mechanism allows all coupling coefficients to be read and distributed within a single clock cycle, eliminating the memory access bottleneck inherent in serial reading. For predefined Max-Cut problems configured as 8×8 grid structures (64 nodes) with coupling coefficients quantized to 2-bit precision, the chip converges to the global optimum within only 5 computing cycles, achieving a final Ising energy of –4032 (Fig.9). This energy follows the analytical expression (n4−n2)(n4−n2), confirming that the chip solves Max-Cut problems with the shortest solving time. For larger extended Max-Cut problems (108×108 nodes), the simulated adiabatic bifurcation algorithm achieves 100% accuracy relative to the theoretical ground state (Fig.10). Monte Carlo simulations over 1,000 independent trials on randomly generated Max-Cut problems demonstrate that simulated adiabatic bifurcation achieves an average Hamiltonian of –3617.42, significantly outperforming simulated annealing which achieves only –3352.12 (Fig.11). For 3-SAT problems with 40 clauses and 8 variables, the chip solves instances with clause-to-variable ratios of 3, 4, and 5 in approximately 6 μs, 10 μs, and 16 μs, respectively (Fig.12). These results align perfectly with theoretical phase transition predictions, confirming the processor's effectiveness across varying problem complexities.  Conclusions  This work proposes an simulated adiabatic bifurcation machine that enables fully parallel updates without duplicating spin copies. To improve throughput, a three-stage pipeline strategy is designed that integrates coupling-coefficient access, momentum update, and position update, achieving bubble-free parallel updating and low-latency solving. For sparse coupling coefficients, a folded storage scheme is adopted to significantly reduce memory area overhead. Both momentum and position variables are represented in 8-bit fixed-point format, ensuring sufficient computational accuracy while balancing resource efficiency. Compared with previous fully connected Ising machines, the proposed bifurcation machine achieves 100% solving accuracy, along with higher energy efficiency and lower hardware overhead, demonstrating substantial application prospects in edge-side combinatorial optimization.
A State Prediction Method for Long-Endurance Fixed-Wing UAV Propulsion Systems
LI Sicheng, WANG Lianqing, LI Zhiyong, WANG Guochang, GE Kaihua, CHEN Junfeng, TAN Rongqing
Available online  , doi: 10.11999/JEIT260188
Abstract:
  Objective  Accurate single-step prediction of key propulsion-system states is essential for early fault warning and autonomous health management of long-endurance fixed-wing unmanned aerial vehicles (UAVs). During high-altitude missions lasting more than 24 h, electrical and thermal variables in the propulsion system exhibit strong coupling, multi-time-constant dynamics, and pronounced day-night regime shifts. These characteristics cause short-term disturbances and long-term drifts to coexist, and hinder adaptive feature weighting under time-varying variable sensitivities. General-purpose time-series predictors may therefore fail to meet the accuracy and robustness requirements of multivariate propulsion-state prediction. To address these challenges, a Grouped Squeeze-and-Excitation Multi-scale Temporal Convolutional Network (GEMS-TCN) is developed by enhancing a modern pure-convolution forecasting backbone with multi-scale embedding and grouped channel attention. The aim is to obtain accurate 10 s-ahead single-step forecasts for 16 key propulsion states from 72-dimensional flight telemetry while satisfying the real-time inference requirement of the 1 Hz telemetry cycle.  Methods  Real flight telemetry from a representative long-endurance fixed-wing UAV is used for model construction and evaluation. The data are sampled at 1 Hz for 9 consecutive days, yielding 806,629 time steps and 72 variables. (Fig.2) Sixteen propulsion-related key states, including the bus voltage, control-unit temperature, winding temperature, and power-device temperature of four motors, are selected as prediction targets, and all 72 variables are used as inputs. (Table 1) The raw data are processed through time indexing, interquartile range (IQR)-based anomaly handling, and interpolation, and are then chronologically divided into training, validation, and test sets at a ratio of 8:1:1 to avoid information leakage. (Fig.4) (Fig.5) GEMS-TCN uses a multi-scale embedding layer with parallel one-dimensional convolutions to extract temporal patterns at different receptive-field scales. Stacked GEMS-TCN blocks combine depthwise temporal convolution, grouped convolutional feed-forward networks, and Grouped Squeeze-and-Excitation (GroupSE) modules to recalibrate intra-variable and cross-variable channel responses hierarchically. (Fig.1) The models are trained with Adam and mean squared error (MSE) loss, and are evaluated using mean absolute error (MAE), MSE, and symmetric mean absolute percentage error (SMAPE). PatchTST, ModernTCN, FEDformer, DLinear, and TimesNet are used as comparison models, with ablation and robustness experiments conducted for further verification.  Results and Discussions  On the full 16-dimensional target set, GEMS-TCN achieves a test-set MAE of 0.171 and an MSE of 0.069. (Table 4) Compared with TimesNet, the strongest baseline in overall trend tracking, GEMS-TCN reduces MSE by 28.1% while maintaining comparable MAE and SMAPE, indicating stronger suppression of large prediction deviations. (Table 4) Stable accuracy is obtained in both daytime and nighttime segments, with MAE/MSE values of 0.189/0.085 and 0.150/0.051, respectively, demonstrating robustness to diurnal operating-condition changes. (Table 4) The prediction trajectories show tighter alignment at thrust-transition points, reduced overshoot, and fewer spurious spikes, while low-bias tracking is preserved for slowly varying nighttime temperature profiles. (Fig.7) Ablation results show that multi-scale embedding and GroupSE provide complementary improvements, and their combination achieves the best overall performance among the ablation settings. (Table 6) Under 1% and 5% random dropouts and 60 s continuous missing intervals, the MSE remains below 0.07, indicating tolerance to practical data-loss scenarios. (Table 5) In addition, GEMS-TCN contains 27.88 M parameters and achieves an inference latency of 0.90 ms per sample, which is well below the 1 s sampling interval.  Conclusions  GEMS-TCN provides a practical convolution-based solution for multivariate propulsion-state prediction in long-endurance fixed-wing UAVs. By integrating multi-scale temporal embedding with hierarchical group-wise channel recalibration, the proposed method jointly represents rapid fluctuations and slow evolution, and better captures multi-time-constant dynamics and multivariate coupling in propulsion telemetry. Real-flight experiments demonstrate stable prediction performance across diurnal regimes, state categories, and data-missing scenarios. Ablation results further confirm the effectiveness and complementarity of multi-scale embedding and GroupSE. These findings indicate that structure-aware modeling tailored to propulsion-state evolution can support health monitoring, early fault warning, and autonomous health management of long-endurance fixed-wing UAVs.
Customized Preparation of Single Crystal Tungsten Tips and Surface Reconstruction Mechanism
GUO Jiamei, YIN Shengyi, ZHANG Yongqing, SUN Wanzhong
Available online  , doi: 10.11999/JEIT260328
Abstract:
  Objective  Refractory metal tungsten, particularly single crystal tungsten, serves as a critical material for high-performance field emission cathodes, which are core components in advanced electron microscopes and electron beam lithography systems. The fabrication of single crystal tungsten microtips with precisely controlled geometry and surface cleanliness remains a significant challenge, especially for applications requiring tip radii ranging from sub-100nm for cold field emission to 0.3–1.0 μm for Schottky-type thermal field emission. Currently, the domestic development of high-end electron microscopes in China faces a major bottleneck due to heavy reliance on imported field emission cathodes. Electrochemical corrosion, the mainstream method for tip preparation, suffers from limitations such as difficulty in cutoff timing control, susceptibility to tip bending or passivation, residual surface impurities, and low yield rates. Moreover, the anisotropic nature of single crystal tungsten introduces additional complexity in morphology control. This study aims to establish a controllable fabrication method for single crystal tungsten tips, enabling tailored geometry and surface quality to meet demanding requirements of different field emission applications while significantly improving fabrication yield.  Methods  Single crystal tungsten wires (diameter 0.12 mm, purity 99.95%, (100) orientation) were used as the starting material. The tips were first pre-shaped by electrochemical corrosion in 1 mol/L NaOH solution under a 10 V DC voltage with pulsed control (6 kHz, 100 μs pulse width) for 10 min. Following corrosion, the tips underwent surface cleaning by sequential immersion in ultrapure water and anhydrous ethanol, followed by drying with high-purity nitrogen. Subsequently, the samples were placed in an ultra-high vacuum system (base pressure < 1 × 10–6 Pa) and subjected to high-temperature oxygen treatment. Process parameters were systematically varied, including temperature (15001600 °C), oxygen flow rate (0.015–1.0 sccm), and treatment duration (5–1440 min). During treatment, high-purity oxygen was introduced while maintaining vacuum levels better than 10–4 Pa. Tip morphologies were characterized by scanning electron microscopy, and surface compositions were analyzed by energy dispersive X-ray spectroscopy. Geometric parameters, including half-cone angle and tip radius, were measured following standardized protocols.  Results and Discussions  The combination of electrochemical corrosion and high-temperature oxygen treatment enabled both effective surface purification and controlled morphological reconstruction. Scanning electron microscopy characterization revealed distinct evolutionary pathways depending on oxygen partial pressure (Fig. 3, Fig. 4). Under oxygen-rich conditions (1.0 sccm, 15001600 °C), the tips underwent “blunting reconstruction,” evolving from an initial inverted cone into a characteristic “cylindrical segment + hemispherical cap” structure (Samples 2–5, Table 2). For Sample 3 treated at 1500 °C for 10 min, the tip radius increased to 327 nm with a half-cone angle reduction of 21%; for Sample 4 treated for 15 min, further evolution occurred with tip radius decreasing to 275 nm and cylindrical segment height increasing substantially. With increasing temperature from 1500 °C (Sample 2) to 1600 °C (Sample 5), the half-cone angle change rate increased from 11% to 25%, while tip radius decreased from 380 nm to 172 nm, indicating accelerated kinetics. This anisotropic behavior is attributed to orientation-dependent surface energies of tungsten and preferential oxidation along specific crystallographic planes. Under oxygen-lean conditions (0.015 sccm, 1500 °C, 1440 min), “sharpening reconstruction” occurred, with tip radius decreasing dramatically from 143 nm to 28 nm and half-cone angle reducing from 8° to 1° (Sample 6). In contrast, treatment in vacuum at 1500 °C primarily removed surface contaminants without altering tip geometry (Sample 1). EDS analysis confirmed that treated surfaces were free from impurities other than trace carbon contamination (Table 3), demonstrating the dual functionality of the process in achieving both purification and reconstruction. The findings reveal that oxygen partial pressure serves as the key determinant of reconstruction direction, with oxygen-rich conditions favoring blunting and oxygen-lean conditions promoting sharpening.  Conclusions  A combined process of electrochemical corrosion followed by high-temperature oxygen treatment was successfully developed for the controllable fabrication of single crystal tungsten field emitter tips. The process achieves dual functionality: effective surface purification through high-temperature vacuum treatment, which removes residual surface contaminants from electrochemical corrosion, and controlled morphological reconstruction enabled by the introduction of oxygen under precisely regulated conditions. By systematically adjusting process parameters including treatment temperature, oxygen flow rate, and treatment duration, tip morphology and dimensions can be tailored to meet specific application requirements. The oxygen partial pressure plays a decisive role in determining the reconstruction pathway. Under oxygen-rich conditions (e.g., 1500 °C, 1.0 sccm), blunting reconstruction yields a “cylindrical segment + hemispherical cap” structure ideally suited for Schottky-type thermal field emission cathodes requiring tip radii between 0.3 μm and 1.0 μm. Under oxygen-lean conditions (e.g., 1500 °C, 0.015 sccm), sharpening reconstruction produces nanoscale sharp tips with tip radius as low as 28 nm, suitable for cold field emission cathodes and scanning tunneling microscope probes. This anisotropic reconstruction mechanism is explained by the selective reaction of oxygen atoms with different tungsten crystal planes under high-temperature conditions, where the formation of volatile oxides modifies local surface energy distribution and promotes the exposure of specific crystallographic planes. Control experiments confirmed that such morphological reconstruction does not occur under oxygen-lean or oxygen-free conditions at the same temperatures, further demonstrating the critical role of oxygen in triggering this process. Importantly, this approach provides an effective method to correct imperfections from the initial electrochemical corrosion step, significantly improving fabrication yield and process robustness, thereby offering a viable technical pathway for the domestic fabrication of high-performance field emission cathodes.
Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding
ZHOU Tong, CHEN Hongzhi, XU Jialong, SUN Peng, JIANG Dajie, LIU Jiankang
Available online  , doi: 10.11999/JEIT260379
Abstract:
Currently, the 3rd Generation Partnership Project (3GPP) Technical Specification Group Service and System Aspects Working Group 4 (TSG SA Working Group 4, SA4) has begun researching Artificial Intelligence (AI)-based audio codecs. Based on the Descript Audio Codec (DAC) being studied by SA4, this paper presents a joint optimization scheme for DAC source-channel coding and modulation with total power constraints, a joint optimization scheme for DAC source-channel coding and modulation with constant modulus constraints, and a DAC joint source-channel coding scheme. Furthermore, simulation evaluation is conducted using the typical physical layer parameter configuration of 3GPP Geostationary Earth Orbit (GEO) voice scenario. Finally, preliminary research suggestions are made, with the hope of providing useful insights for the subsequent research on GEO voice.  Objective  In recent years, Joint Source-Channel Coding (JSCC) has gained widespread attention in academia and industry. In academia, the sources studied for JSCC include video, image, and audio. However, the industry has not focused on these sources as in academia, but rather on JSCC where channel state information (CSI) is considered as the source. The reason the industry has not considered JSCC for video and images is due to two main factors. First, for video and images, 3GPP generally does not conduct independent research, but instead directly reuses source compression standards developed by external organizations such as the Moving Picture Experts Group (MPEG) and the Joint Photographic Experts Group (JPEG). Second, it is constrained by the challenges of the source-channel coding architecture. For example, in joint source-channel coding at the application layer, the channel information obtained at the application layer is often outdated, which limits the potential gains of source-channel coding. However, the situation is quite different for audio. Audio coding and decoding are designed independently by 3GPP, and currently, 3GPP SA4 is researching AI-based speech coding and decoding. Moreover, in GEO speech scenarios, where satellites remain largely stationary relative to ground stations and the channel changes slowly, the physical layer CSI can be transmitted to the application layer, overcoming the issue of outdated channel information. This makes the research on JSCC for audio more promising within 3GPP, compared to JSCC for video or image sources.  Methods  This paper firstly uses the typical physical layer parameter configurations of 3GPP GEO voice and the DAC encoder, currently being studied by SA4, as an evaluation baseline, ensuring the fairness of the proposed scheme’s gain evaluation. Based on this, the paper proposes the integration of DAC with channel coding and modulation to improve audio quality in low-bitrate scenarios. First, the total power constrained (TPC) DAC-JSCCM scheme is considered, which offers the maximum potential gain due to the joint optimization of source coding, channel coding, and modulation. Then, considering the high PAPR (Peak-to-Average Power Ratio) impact on the power amplifier, the paper proposes a constant-envelope constrained (CEC) DAC-JSCCM scheme. Finally, since the modulation symbols included in RTP payload require significant changes to the existing protocol, a DAC-JSCC scheme that is more compatible with current protocols is proposed, where bit sequences are included in RTP payload.  Results and Discussions  In the evaluation of baseline scheme 1, the link-level simulation at the physical layer is first conducted to obtain the relationship between SNR and BLER for an input length of 64 bits using a 1/3 Turbo code. Then, assuming that the error rate of an audio frame is the same as the BLER, the transmission error rate of an RTP packet containing L audio frames is calculated. According to the existing processing logic, if one audio frame in an RTP packet is transmitted with an error, the entire RTP packet will be discarded. In real-time communication, VoIP, or speech enhancement scenarios, a PESQ score of 2.5 is generally considered the minimum usable threshold, and voice quality below this score is regarded as unsuitable for professional or commercial services. Therefore, in GEO voice calls, we choose PESQ = 2.5 as the reference point. When PESQ = 2.5, the TPC DAC-JSCCM and CEC DAC-JSCCM provide a coverage gain of approximately 4 dB compared to baseline scheme 1 (Fig.4). The DAC-JSCC scheme offers a coverage gain of 2.5 dB compared to baseline scheme 1 (Fig.4). The coverage gains of DAC-JSCCM and DAC-JSCC come from the application layer performing joint source-channel coding, which avoids the cliff effect caused by using Turbo coding at the physical layer. When the Complementary Cumulative Distribution Function (CCDF) of Peak-to-Average Power Ratio (PAPR) is \begin{document}$ {10}^{-3} $\end{document}, the TPC DAC-JSCCM is 4 dB higher than the QPSK modulation, with sharper signal peaks, higher requirements for power amplifier linearity, and is more prone to distortion. In contrast, the CEC DAC-JSCCM is very close to the QPSK modulated PAPR.  Conclusions  This paper proposes the Total Power Constrained DAC-JSCCM, Constant Modulus Constrained DAC-JSCCM, and DAC-JSCCM schemes, based on the DAC codec currently being researched by SA4. Simulations are conducted using the typical physical layer configuration of the existing 3GPP GEO voice scenario. The simulation results show that, when the PESQ is 2.5 (the minimum usable threshold), compared to DAC combined with existing physical layer transmission technologies, the proposed DAC-JSCCM provides a coverage gain of 4 dB, while the proposed DAC-JSCC scheme provides a coverage gain of 2.5 dB. SA4 is currently researching more advanced GEO voice codecs beyond DAC, and we will extend our research by combining better GEO voice codecs in the future.
Lightweight Semantic Communication System Driven by User Personalization in UAV Networks
WEI Yuxuan, CHEN Xiao, CHEN Qiuyu, JIANG Hao, YANG Zhaohui
Available online  , doi: 10.11999/JEIT260370
Abstract:
  Objective  With the rapid development of the low-altitude economy and 6G intelligent networks, Unmanned Aerial Vehicle (UAV) image communication shows strong potential in target reconnaissance, emergency communication, and intelligent inspection. However, conventional pixel-level transmission cannot meet the requirements of efficient, low-latency, and intelligent communication because UAVs are constrained by limited bandwidth, payload capacity, and onboard computational resources. Semantic communication, which transmits only task-relevant information, provides an effective solution for improving communication efficiency in resource-constrained scenarios. However, current studies on UAV image transmission face several challenges. First, fixed network architectures use unified semantic encoding and transmission strategies for all users and cannot adapt to different personalized requirements. Second, new user access usually requires interest pre-training or model fine-tuning, which increases deployment overhead. Third, most models have high computational complexity. To address these issues, this paper proposes the Lightweight Personalized UAV Semantic Communication (LPUSC) system to balance computational cost, transmission bandwidth, and personalized requirements. The system enables personalized transmission through low-overhead semantic index interaction and a lightweight semantic extraction module, without pre-training for new users. A dual-branch end-to-end network is also designed. In this network, the semantic index transmission network works with the semantic image transmission network trained by a weighted hybrid loss function, thereby supporting high-precision and high-quality transmission of personalized semantic images.  Methods  The proposed LPUSC system adopts a dual-branch architecture for accurate task-driven semantic content transmission. In the semantic index interaction branch, the lightweight object detection model YOLO11s is used to perform semantic perception on UAV-captured visual scenes. Complex image information is compressed into low-dimensional semantic index vectors, which reduces transmission redundancy and communication overhead. On this basis, an end-to-end semantic index transmission network is designed to improve the robustness of semantic index transmission under complex wireless channel conditions. Through the semantic index interaction mechanism, the system accurately identifies targets of user interest and provides prior guidance for subsequent semantic content extraction. In the semantic image transmission branch, the lightweight and high-precision MobileSAM model is adopted for semantic region extraction. This branch uses the target bounding boxes returned by the semantic index interaction branch as box-prompt inputs, enabling pixel-accurate segmentation and extraction of specific semantic targets. To further improve semantic image reconstruction quality, a weighted hybrid loss function is designed. This function integrates Mean Squared Error (MSE), L1-norm loss, Structural Similarity Index Measure (SSIM) loss, gradient loss, perceptual loss, and background suppression loss. These losses jointly optimize pixel accuracy, structural preservation, and fine-detail restoration. Through the joint constraints of multiple loss terms, the proposed system improves semantic region reconstruction and achieves high-quality semantic image transmission.  Results and Discussions  Simulation results validate the proposed LPUSC system in semantic extraction and end-to-end transmission. For semantic extraction, three schemes are compared: YOLO11s-seg, YOLO11s + Segment Anything Model (SAM), and YOLO11s + MobileSAM (Fig. 4). The results show that the detection-segmentation decoupled architecture achieves better semantic boundary localization accuracy. Combined with the quantitative analysis in Table 1, the YOLO11s + MobileSAM scheme reduces resource use while maintaining high extraction accuracy. This confirms its suitability for resource-constrained UAV platforms. For end-to-end transmission, the semantic index vector transmission results (Fig. 5) show that the Bit Error Rate (BER) decreases monotonically as the Signal-to-Noise Ratio (SNR) increases in all three channel environments. The rural environment achieves the best performance, followed by the suburban and urban environments. These differences are mainly caused by variations in scatterer density and link blockage across environments. The proposed transmission network maintains stable BER under different Doppler frequencies, demonstrating its robustness under dynamic channel conditions. For semantic image transmission, the proposed weighted hybrid loss function shows good training stability (Fig. 6), and LPUSC consistently outperforms the Deep Joint Source-Channel Coding (DeepJSCC) and JPEG + Low-Density Parity-Check (LDPC) baselines across the full SNR range (Fig. 7). Specifically, LPUSC achieves SSIM and Peak Signal-to-Noise Ratio (PSNR) gains of 1.3% and 4.8% over DeepJSCC, respectively, and gains of 43% and 79.5% over JPEG + LDPC, respectively. These results indicate that the proposed personalized semantic image transmission network achieves high-quality reconstruction and remains robust to channel variations.  Conclusions  To improve the efficiency and flexibility of UAV image communication, this paper proposes LPUSC, a lightweight personalized semantic communication system. The system uses a dual-branch transmission architecture that integrates lightweight, high-precision object detection and semantic segmentation models. It enables personalized content transmission without interest pre-training. This design satisfies personalized user requirements while maintaining low computational and communication overhead. Simulation results show that the LPUSC system achieves stable and reliable semantic index interaction and outperforms the DeepJSCC and JPEG + LDPC baselines in semantic region reconstruction. The proposed system provides a useful reference for efficient UAV image semantic communication in 6G low-altitude intelligent networks.
Design of Lightweight Gated Recurrent Unit Network Model Based on Memristor
HUA Honghu, XU Jia, ZHANG Bohao, WANG Wei, LI Zhiwei, LIU Haijun
Available online  , doi: 10.11999/JEIT260152
Abstract:
  Objective  With the slowdown of Complementary Metal-Oxide-Semiconductor (CMOS) technology scaling and the inherent memory-computation separation of von Neumann architectures, conventional computing systems face increasing challenges in processing large-scale sequential data. Memristors provide a promising solution because of their high integration density, fast switching speed, and synaptic plasticity. Memristor crossbar arrays naturally support Vector-Matrix Multiplication (VMM) in the analog domain, enabling energy-efficient in-memory computing. As a representative recurrent neural network, the Gated Recurrent Unit (GRU) has achieved excellent performance in sequential tasks such as trajectory prediction and urban sound classification. However, conventional hardware implementations of GRU networks require frequent data transfer between memory and processing units, resulting in high energy consumption and limited throughput. Although memristor-based GRU implementations improve computational efficiency, their large parameter size and high weight precision require substantial hardware resources and reduce deployment reliability on resource-constrained memristor arrays. In addition, device non-idealities, such as conductance fluctuations, further reduce inference accuracy. Existing memristor-based GRU methods generally treat weights and activations using the same quantization strategy without considering their different hardware implementation characteristics, and they provide limited robustness against device variations. This paper addresses these issues through a hardware-algorithm co-design strategy.  Methods  This paper proposes a lightweight memristor-based GRU network model. A 1T1R (one-transistor-one-resistor) memristor crossbar array is adopted for weight mapping and analog Multiply-Accumulate (MAC) operations. Signed weights are represented by differential pairs of positive and negative conductance matrices because memristor conductance values are inherently non-negative. A linear transformation is used to map trained network weights to memristor conductance values. To account for the different hardware implementation paths of weights and activations, a device-aware fusion quantization method based on performance analysis is proposed. Symmetric quantization is applied to weights stored in the memristor array because the zero-centered quantization range eliminates zero-point storage and simplifies write-driver circuit design. In contrast, asymmetric quantization is applied to activation values computed in peripheral circuits, thereby preserving the dynamic range and reducing quantization error. To improve robustness against memristor conductance fluctuations, weight noise training is incorporated into Quantization-Aware Training (QAT). Gaussian noise with an intensity determined by the device variation parameter is injected into quantized weights during each forward pass. This strategy acts as a regularizer that guides the model toward flatter loss minima and improves tolerance to weight perturbations. During backpropagation, the straight-through estimator updates the full-precision floating-point weights, whereas noise is dynamically resampled in every forward pass.  Results and Discussions  On the public UrbanSound8K dataset, the proposed full-precision lightweight memristor-based GRU network model achieves a classification accuracy of 93.94%. After applying the device-aware fusion quantization method, the 6-bit quantized model achieves 92.68% accuracy, corresponding to only a 1.26% decrease while reducing weight precision by 81.25% (Table 1). The proposed model outperforms Dilated Convolution (78.00%), LM-MFCC+GRU (92.00%), TFFS-DNN (88.74%), TFCNN (93.10%), and CL-Transformer (92.95%) under their full-precision settings (Table 2). Under noisy input conditions with Signal-to-Noise Ratios (SNRs) ranging from −10 dB to 10 dB, the 6-bit quantized model exhibits robustness comparable to or better than that of the full-precision model, demonstrating the effectiveness of the proposed device-aware fusion quantization strategy (Table 3). From the perspectives of storage, hardware resources, and device feasibility, 6-bit quantization reduces weight storage from 5.6 MB to 1.05 MB, corresponding to a compression ratio of 81.2%, while requiring only 2.8 million memristor cells under the 1T1R mapping scheme. Weight noise training also substantially improves robustness against device non-idealities. When the conductance variation reaches 14%, the classification accuracy increases from 82.97% to 91.14%. At the maximum simulated variation of 28%, the accuracy increases from 54.23% to 87.01% (Fig. 7), demonstrating improved tolerance to memristor device variations. On a self-constructed true-false trajectory dataset, the lightweight memristor-based GRU network model achieves 97.35% accuracy at full precision and 96.51% after 6-bit quantization, with only a 0.84% decrease, outperforming the Dilated Convolution baseline (Table 4). To further verify its applicability to different sequential tasks, the lightweight memristor-based GRU network model is evaluated on lithium-ion battery State-of-Charge (SOC) estimation using a public dataset. The 6-bit quantized model achieves Root Mean Square Errors (RMSEs) of 1.48%, 0.79%, and 0.74% at 0 °C, 25 °C, and 45 °C, respectively, outperforming the existing memristor-based GRU implementation. The proposed model also achieves lower RMSEs than the comparison method at all evaluated quantization precisions of 6 bits and above (Table 5).  Conclusions  This paper presents a lightweight memristor-based GRU network model for hardware deployment. By combining device-aware fusion quantization with weight noise training integrated into Quantization-Aware Training (QAT), the model achieves substantial memory compression while maintaining high classification accuracy and improving robustness to memristor device non-idealities. Experimental results on multiple datasets and sequential tasks demonstrate that the 6-bit quantized model preserves competitive accuracy and stable performance, providing an effective solution for deploying GRU networks on resource-constrained memristor-based edge computing platforms.
A Survey of Cooperative Mission Planning for Imaging Satellites Observing Moving Targets
XU Zhuo, FAN Shenghua, YUE Haitao, QU Tao, WANG Dingwen, SUN Shilei
Available online  , doi: 10.11999/JEIT260133
Abstract:
  Significance   Cooperative mission planning for imaging satellites observing moving targets is a key technique that supports the transition of space-based Earth observation systems from static regional coverage to a dynamic closed-loop paradigm consisting of wide-area search, dynamic tracking, and feedback-guided supplementary search. It plays an important role in emergency response, maritime monitoring, wide-area situational awareness, and persistent observation of high-value moving targets. The primary challenge arises from the conflict between uncertainty in future target states and the reliance of conventional mission planning models on deterministic inputs. Unlike static targets, moving targets have neither fixed locations nor fixed visibility windows. Their future states are generally represented by probability distributions, confidence regions, or grid-based target existence probabilities. Effective mission planning therefore requires not only accurate target motion prediction but also systematic integration of uncertainty into planning objectives, constraints, and replanning triggers. A comprehensive review from an uncertainty-driven perspective is therefore needed.  Progress   This survey reviews cooperative mission planning for imaging satellites observing moving targets from an uncertainty-driven perspective. Typical moving targets are classified into maritime moving targets, highly time-sensitive aerospace targets, and ground moving targets according to their operating environments, dynamic characteristics, and observation requirements. Although these target categories differ in maneuverability, prior constraints, and observation windows, they share a common planning challenge: coupling uncertain target motion with deterministic satellite observation actions under stringent platform and resource constraints. Methods for target motion prediction and spatiotemporal uncertainty representation are first reviewed. Physics-based methods characterize target state evolution using kinematic constraints, dynamic models, covariance propagation, reachable sets, and Markov state transition models. Data-driven methods learn motion patterns from historical trajectories, Automatic Identification System (AIS) data, remote sensing observations, meteorological information, and geographic constraints. From the perspective of mission planning, the utility of these methods depends on whether outputs such as covariance, target existence probability, confidence regions, and information gain can be directly incorporated into planning models. Observation task modeling, cooperative planning architectures, optimization algorithms, and closed-loop replanning mechanisms are then analyzed. Deterministic task models simplify uncertainty into trajectory points, visibility windows, or fixed geographic regions, while probabilistic task models incorporate target existence probability, belief states, and information gain into objective functions, constraints, and state transition models. Centralized, distributed, and hybrid planning architectures are compared with respect to global optimization capability, onboard autonomy, communication overhead, and response timeliness. Exact optimization methods, heuristic methods, metaheuristic algorithms, Deep Reinforcement Learning (DRL), and Large Language Model (LLM)-assisted solution strategies and algorithm design are also reviewed. Finally, state-triggered replanning, Receding Horizon Optimization (RHO), and Model Predictive Control (MPC) are summarized as representative approaches for closed-loop dynamic scheduling.  Conclusions  The reviewed studies indicate that cooperative mission planning is evolving from open-loop static scheduling to closed-loop dynamic planning. Nevertheless, several challenges remain. First, uncertainty information generated during target prediction is not fully exploited in planning decisions. Rich probabilistic information is frequently reduced to deterministic time windows, discrete trajectory points, or geometric regions, thereby limiting risk-aware task allocation. Second, distributed cooperation lacks reliable belief-state consistency. Differences in local observations may lead satellites to maintain inconsistent estimates of the same target state, resulting in redundant observations, task conflicts, and inefficient resource utilization. Third, dynamic replanning lacks unified benefit-cost criteria for determining replanning triggers. Excessively frequent replanning increases attitude maneuver time, energy consumption, and onboard storage resource use, whereas delayed replanning may fail to respond to actual target maneuvers. Fourth, LLMs have demonstrated potential for task requirement parsing, constraint modeling, heuristic generation, and algorithm design for satellite scheduling, but their application to cooperative mission planning for moving targets remains limited.  Prospects   Future research should focus on developing a more robust closed-loop planning framework. Prediction uncertainty should be incorporated directly into planning models through chance-constrained planning, belief-state planning, or Partially Observable Markov Decision Processes (POMDPs), enabling covariance, target existence probability, and information entropy to be integrated into planning objectives, constraints, and replanning triggers. Bayesian updating or sequential Bayesian filtering should use both successful detections and missed detections to continuously refine the prediction layer. Distributed cooperation requires lightweight state synchronization and belief fusion supported by compact state-sharing descriptors and event-triggered communication. Replanning decisions should be guided by information gain and benefit-cost evaluation. In addition, LLMs should be developed as verifiable auxiliary tools rather than direct replacements for optimization solvers. They can assist with task requirement structuring, constraint modeling, heuristic generation, and algorithm component design, whereas feasibility verification, solution refinement, and performance evaluation should remain the responsibility of formal verification methods, conventional optimization algorithms, and simulation environments. These research directions are expected to improve the robustness and uncertainty awareness of mission planning for satellite observation of moving targets.
Heterogeneous Task Cooperative Scheduling Architecture for Networked Radar in Saturation Attack Air Defense Early Warning
YE Juhang, FANG Yuyuan, WEI Shaopeng, DUAN Jia, ZHANG Lei
Available online  , doi: 10.11999/JEIT260373
Abstract:
  Objective  To address the severe challenges posed by Unmanned Aerial Vehicle (UAV) swarms and intelligent loitering munitions, which generate massive, sudden, and heterogeneous early warning tasks during saturation attacks, a scalable networked-radar cooperative scheduling architecture is proposed. Existing architectures suffer from rigid dynamic coordination and insufficient capability for heterogeneous task scheduling. The non-convex cooperative scheduling problem is therefore decoupled into a multi-stage decision process consisting of multidimensional dynamic resource coordination and adaptive heterogeneous task scheduling. By incorporating a dispatch mechanism and Hierarchical Reinforcement Learning (HRL), a hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling is developed. The architecture is implemented through an execution-redispatch cognitive closed loop, together with a Target Dispatch Algorithm (TDA) and a hierarchical command-and-scheduling method, to address multi-radar, multi-target, and multi-task scheduling in saturation attack scenarios.  Methods  The cooperative scheduling problem is first decoupled into dispatch-oriented network-level multidimensional resource coordination and hierarchical-command-based node-level adaptive heterogeneous task scheduling. An environment perception layer establishes a dynamic uncertainty model based on the Bayesian Cramér-Rao Lower Bound (BCRLB) and radar detection probability to jointly characterize target threat levels and radar operating states. The network coordination layer then adopts an adaptive weighted dispatch model based on comprehensive combat effectiveness to achieve dynamic target allocation while constructing a scalable execution-redispatch cognitive closed loop. To solve the resulting large-scale, nonlinear, multi-constraint generalized bipartite graph matching problem, the proposed TDA, an MMAS-based algorithm incorporating feasibility-rule-based constraint handling, is developed as a constructive solution. At the node level, Hierarchical Q-Learning (HQL) is implemented through a serial dual-Q-table implementation for distributed heterogeneous task scheduling. Using task proportion control as the hierarchical subgoal, the proposed method transforms upper-level operational intent into lower-level beam dwell scheduling, enabling adaptive, long-term, and interpretable execution of heterogeneous tasks with complex dependency relationships.  Results and Discussions  A simulated point-defense scenario against UAV swarm and loitering munition saturation attacks is established using three networked homogeneous S-band medium-range phased-array radars to counter 200 high-speed maneuvering targets. For network-level coordination, the proposed TDA replaces conventional penalty functions with a hierarchical solution framework based on feasibility-rule constraint handling. By exploiting prior model information, TDA achieves higher solution quality and faster convergence than the Max-Min Ant System (MMAS), Artificial Bee Colony (ABC), and Genetic Algorithm (GA) (Fig. 3). Although computational complexity increases, the execution time remains well within the dispatch cycle, improving solution quality with only millisecond-level computational overhead while satisfying real-time operational requirements (Fig. 4). For node-level scheduling, HQL employs hierarchical macro- and micro-level decisions to ensure policy consistency. Supported by an internal dense transfer-reward mechanism, HQL achieves higher learning efficiency and better long-term policy quality than Q-Learning (QL) and the Priority-Based Method (PBM) (Fig. 5). The hierarchical serial dual-Q-table framework maintains balanced task proportions, maximizes comprehensive combat effectiveness, and improves resource utilization (Fig. 6). Furthermore, comparisons of the target track-loss rate and mean tracking error show that HQL achieves the lowest mean tracking error by prioritizing high-quality tracking tasks, despite a moderately higher target track-loss rate, demonstrating superior long-term scheduling capability (Fig. 7).  Conclusions  The proposed hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling effectively addresses the limitations of rigid dynamic coordination and insufficient heterogeneous task scheduling capability in existing networked radar systems. The proposed framework enables real-time cooperative scheduling of search, confirmation, and tracking tasks during large-scale, high-speed saturation attacks, thereby improving the operational capability of the air defense early warning system. Simulation results demonstrate improvements in applicable processing scale, environmental adaptability, long-term scheduling capability, scalability, and interpretability. Future work will explore online learning to reduce the discrepancy between offline training and online deployment and will further extend the architecture by integrating weapon-target assignment to support unified early warning and fire-control systems.
Analysis of Age upon Decisions and Distortion at Decisions in IoT Status Update Systems with Batch Arrivals
LIU Lei, JIN Wenkai, ZHANG Qingqing, LI Yuzhou, JIANG Fan
Available online  , doi: 10.11999/JEIT260359
Abstract:
  Objective  The rapid development of the Internet of Things (IoT) makes the timely transmission and processing of status updates essential for modern systems, where information freshness at decision epochs plays a critical role. In many IoT applications, such as smart grid fault detection and Industrial Internet of Things (IIoT) cluster monitoring, status updates typically arrive in batches rather than individually. However, most existing studies on Age of Information (AoI) assume single-update arrivals and therefore cannot accurately characterize the queueing dynamics caused by batch arrivals. Besides information freshness, distortion at decision epochs is another key factor because it directly affects decision quality. A fundamental tradeoff therefore exists between information freshness and distortion. Waiting for more complete status update information allows more completed status updates to be incorporated into joint estimation, but increases queueing and transmission delays, thereby reducing information freshness. In contrast, triggering decisions earlier reduces delay but increases distortion because fewer completed status updates are available for joint estimation. To address this problem, this paper investigates the tradeoff between information freshness and distortion in an IoT status update system with batch arrivals by adopting Age upon Decisions (AuD) and Distortion at Decisions (DaD) as performance metrics. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. Furthermore, for the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. The results provide theoretical insight and practical guidance for the design of IoT status update systems with batch arrivals.  Methods  Information freshness and distortion at decision epochs are analyzed for an IoT status update system with batch arrivals. AuD and DaD are adopted to quantify information freshness and distortion, respectively. Based on queueing theory, analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. A typical case with geometrically distributed batch sizes is then investigated. An alternating iterative optimization algorithm is further developed to jointly optimize the batch arrival rate, average batch size, and decision threshold to minimize the weighted sum of the average AuD and average DaD.  Results and Discussions  Simulation results validate the theoretical analysis. The average AuD exhibits a nonmonotonic trend as the batch arrival rate increases, first decreasing and then increasing. In addition, the Batch-size Coefficient Of Variation (BCOV) has a significant effect on the average AuD, with a smaller BCOV providing better information freshness performance. Under high-load conditions, queue backlogs become more severe, and stochastic fluctuations in batch arrivals have a greater effect on the queueing process. This increases service-time variability and amplifies the effect of BCOV on the average AuD (Fig. 2). As the average batch size increases, the system queue length and queueing delay increase, leading to a larger average AuD. At the same time, the decision control unit can utilize more completed status updates for joint estimation, thereby reducing the average DaD (Fig. 3). Moreover, the average DaD decreases as the decision threshold increases because more completed status updates are incorporated into the joint estimation process, improving estimation accuracy. A larger BCOV also increases the number of completed status updates available for joint estimation and therefore further reduces the average DaD (Fig. 4). The optimization results show that the solutions obtained by the proposed algorithm lie on the Pareto frontier, demonstrating its effectiveness. By comparison, fixed batch arrival rates and decision thresholds produce performance that is considerably farther from the Pareto frontier, demonstrating the advantage of jointly optimizing system parameters (Fig. 5).  Conclusions  This paper investigates an IoT status update system with batch arrivals by adopting AuD and DaD to quantify information freshness and distortion, respectively. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. For the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. Simulation results verify the theoretical analysis and reveal the effects of the batch arrival rate, average batch size, and decision threshold on the average AuD and average DaD. The results also demonstrate that the proposed low-complexity algorithm effectively identifies Pareto-optimal solutions for the AuD-DaD tradeoff. This study considers only the batch arrival characteristics of status updates. Future work will incorporate batch service mechanisms to further examine their effects on the tradeoff between AuD and DaD. Flexible decision mechanisms can also be developed to achieve adaptive AuD-DaD tradeoffs according to the heterogeneous requirements for information freshness and distortion across applications with different batch characteristics.
Kolmogorov-Arnold Nonlinear Enhancement Method for Aerial-Ground Person Re-IDentification
CHEN Yijun, ZENG Xianxian, LIU Shun, WANG Leijun
Available online  , doi: 10.11999/JEIT260430
Abstract:
  Objective  Aerial-Ground Person Re-IDentification (AG-PReID) aims to match the same person across Unmanned Aerial Vehicle (UAV) and ground-camera views. Compared with conventional same-platform person re-identification, this task faces larger cross-view appearance variation and more severe cross-domain distribution shifts. Under these conditions, identity-consistent cues are often weakened by strong viewpoint asymmetry and cross-domain appearance distortion. Existing methods mainly focus on feature extraction and cross-view representation alignment. However, the classification supervision branch still relies heavily on linear feature transformation, which limits its ability to model complex nonlinear discriminative relationships in high-dimensional feature spaces. A stronger nonlinear supervision mapping is therefore needed to better exploit high-order feature interactions and local discriminative variations. To address this issue, this paper proposes a Kolmogorov-Arnold Nonlinear Enhancement Module (KANEM). KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier. It uses learnable nonlinear mappings to adaptively enhance features for more discriminative cross-view representation learning.  Methods  The backbone follows the View-Decoupled Transformer (VDT), which introduces an additional view token and performs layer-wise view decoupling. This design separates view-related factors from identity features and reduces representation bias between aerial and ground domains. Based on this framework, KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier, thereby providing adaptive nonlinear mappings for feature enhancement. Specifically, KANEM consists of a base activation branch and a spline branch, which are stacked into cascaded function-mapping layers. This design enables more flexible nonlinear modeling than conventional linear or MultiLayer Perceptron (MLP)-based transformations. It allows the model to capture local nonlinear variations and complex correlations among feature dimensions. To improve discriminability and further separate identity and view information, the network is jointly optimized using identity classification loss, view classification loss, triplet loss, and orthogonality loss. KANEM is used only during training and is removed during inference, so no extra inference cost is introduced.  Results and Discussions  Comprehensive evaluations are conducted on the CARGO and AG-ReID datasets. The results show that the proposed method consistently performs better than the baseline model and existing state-of-the-art methods. On CARGO, the proposed method achieves 70.19%/63.16%/51.34% in Rank-1 accuracy, Mean Average Precision (mAP), and Mean Inverse Negative Penalty (mINP), respectively, under the overall ALL retrieval protocol. It also achieves 58.75%/53.27%/41.11% under the most challenging aerial-ground (A↔G) cross-view retrieval protocol (Table 1). On AG-ReID, the proposed method achieves the best performance under both retrieval protocols. It reaches 84.41%/76.21%/53.05% in Rank-1/mAP/mINP for aerial-to-ground (A→G) retrieval and 86.69%/77.99%/52.28% for ground-to-aerial (G→A) retrieval (Table 2). Ablation studies on CARGO further verify the effectiveness of KANEM. They show that KANEM achieves better overall performance than conventional linear transformation and MLP-based alternatives. This result indicates that the proposed nonlinear enhancement strategy is more suitable for supervision mapping in CARGO (Tables 3 and 4). In addition, integrating KANEM into other person re-identification tasks further demonstrates its potential generalization ability across different scenarios (Table 5). Parameter analysis shows that setting λ to 0.001 enables the model to better balance the complexity difference between view classification and identity classification (Fig. 2(a)). When G and P are set to 5 and 3, respectively, the model effectively fits nonlinear variations in the feature space while preserving the smoothness and continuity of spline functions. This setting achieves effective nonlinear feature enhancement (Fig. 2(b)(d)). The two-dimensional t-distributed Stochastic Neighbor Embedding (t-SNE) visualization shows that the enhanced features have higher intra-class compactness and better inter-class separability (Fig. 3). The top-5 retrieval comparisons further provide qualitative evidence that the proposed method improves ranking quality and retrieval robustness under all four retrieval protocols on CARGO. It promotes correct matches to higher positions and returns more relevant samples among the top-ranked results (Fig. 4).  Conclusions  This paper presents KANEM for AG-PReID. The proposed module is motivated by the large discrepancy between UAV and ground-camera views and by the limited capacity of linear feature transformation in the classification branch to capture complex nonlinear discriminative relationships. By replacing the conventional fully connected feature transformation between backbone output features and the linear classifier, KANEM provides a more flexible nonlinear supervision mechanism for cross-view representation learning. Through adaptive nonlinear enhancement, it better models complex feature interactions in high-dimensional spaces and strengthens the representation of cross-view consistency and fine-grained discriminative cues. Experimental results on CARGO and AG-ReID demonstrate the effectiveness of the proposed method, particularly in challenging scenarios with large view discrepancies. Future work will further refine the nonlinear mapping mechanism of KANEM and explore its use in more complex cross-view settings to improve model discriminability and generalization performance.
A Task Prediction-augmented Hierarchical Offloading Method for Space-Air-Ground Integrated Networks
ZHANG Linghao, XU Bo, SUN Jinlong, LAI Haiguang, ZHAO Haitao
Available online  , doi: 10.11999/JEIT260217
Abstract:
  Objective  Space-Air-Ground Integrated Networks (SAGIN) have become key infrastructure for future 6G communications. They support wide-area coverage and flexible deployment through the coordinated operation of Low Earth Orbit (LEO) satellites, Unmanned Aerial Vehicles (UAVs), and Ground Users (GUs). With the rapid growth of Internet of Things (IoT), Internet of Vehicles (IoV), and smart city applications, terminal devices generate increasingly diverse computation-intensive tasks. These tasks impose high requirements on real-time computing and resource scheduling. Mobile Edge Computing (MEC) has been integrated into SAGIN architectures to provide near-user computing services by using UAVs and satellites as edge nodes, thereby reducing task completion latency. However, efficient task offloading remains challenging when average task completion latency and UAV flight energy consumption must be jointly reduced. This difficulty is caused by the strong coupling among UAV trajectory planning, task offloading, and computational resource allocation. It is further intensified by the dynamic and partially observable nature of SAGIN environments. Existing Multi-Agent Reinforcement Learning (MARL) methods mainly rely on reactive decisions based on instantaneous observations. They lack awareness of future task workload changes, which leads to decision lag and limited adaptability under bursty traffic. To address these issues, a task prediction-augmented MARL method is proposed to support forward-looking decisions in dynamic SAGIN environments.  Methods  A three-layer SAGIN-MEC architecture is considered, including one LEO satellite, multiple UAVs, and GUs. Tasks can be processed locally, offloaded to UAVs through Ground-to-Air (G2A) links, or further relayed to the LEO satellite through Air-to-Satellite (A2S) links under a partial offloading mechanism. The joint optimization of UAV trajectory, user association, offloading ratios, and computational resource allocation is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem. The objective is to minimize the weighted sum of average task completion latency and UAV flight energy consumption. Owing to the nonconvexity and high dimensionality of this problem, it is reformulated as a DECentralized Partially Observable Markov Decision Process (DEC-POMDP). A Prediction-Augmented Multi-Agent Proximal Policy Optimization (PA-MAPPO) algorithm is then developed. A lightweight Exponential Smoothing-Autoregressive (ES-AR) prediction module is used to generate multi-step workload forecasts, which are incorporated into the state space of each agent. The algorithm adopts a bilevel structure. In the outer layer, Centralized Training and Decentralized Execution (CTDE)-based PA-MAPPO generates UAV trajectory actions. In the inner layer, Block Coordinate Descent (BCD)-based convex optimization solves the resource allocation and offloading subproblems, and closed-form resource allocation solutions are obtained through Lagrangian analysis. Generalized Advantage Estimation (GAE) and the PPO-Clip objective are used to improve training stability and convergence.  Results and Discussions  Simulations are conducted with one LEO satellite, five UAVs, and 50 GUs in a 1×1 km2 area. PA-MAPPO is compared with MAPPO without prediction and Prediction-Augmented Multi-Agent Deep Deterministic Policy Gradient (PA-MADDPG). The training curves show that PA-MAPPO converges within 500~700 episodes, with the highest average reward and the smallest variance, indicating better stability (Fig. 3). As the number of GUs increases from 20 to 80, PA-MAPPO consistently achieves the lowest system cost. Compared with MAPPO and PA-MADDPG, it reduces the average cost by approximately 12.4% and 18.7%, respectively (Fig. 4). Experiments with different UAV numbers show a U-shaped cost curve for all algorithms. The best configuration is obtained when U=5, where PA-MAPPO achieves the minimum cost (Fig. 5). Sensitivity analysis of the latency-energy tradeoff weight ω confirms that PA-MAPPO remains robust under different optimization preferences (Fig. 6). The prediction horizon H has a nonmonotonic effect on performance. When H=5, PA-MAPPO obtains the best result and reduces the cost by approximately 14.9% compared with the no-prediction case. Longer horizons degrade performance because prediction errors accumulate (Fig. 7).  Conclusions  The PA-MAPPO algorithm is proposed to solve the joint optimization of UAV trajectory planning, user association, task offloading, and computational resource allocation in dynamic SAGIN environments. By integrating a lightweight ES-AR task workload prediction module into the MARL process, PA-MAPPO enables UAV agents to account for future task dynamics. This design reduces the decision lag caused by purely reactive methods. The inner BCD-based convex optimization converges to a Karush-Kuhn-Tucker (KKT)-stationary point, while the outer CTDE-based PPO mechanism improves training stability and scalability. Simulation results show that PA-MAPPO outperforms baseline methods in average task completion latency, UAV flight energy consumption, and overall system cost. It also maintains strong scalability and robustness under different system configurations. Future work will study online prediction and decision co-optimization in multi-satellite cooperative scenarios and examine the effect of dynamic network topology changes on algorithm performance.
A Dual-polarized Magnetoelectric Dipole Antenna Array with Differential Feeding
TANG Li, WANG Zhihui, ZHAO Luyu
Available online  , doi: 10.11999/JEIT260505
Abstract:
  Objective  This study addresses key challenges in Fifth-Generation (5G) millimeter-wave terminal antennas by designing a compact, high-performance dual-polarized array. Existing designs often face trade-offs among bandwidth, beam-scanning range, and integration complexity. To address these limitations, this paper proposes a differentially fed magnetoelectric dipole array. A stacked stripline-slot-stripline balun is used to enable efficient single-ended-to-differential conversion, and the array design is optimized. The objective is to realize an integrated solution with wideband operation, low cross-polarization, wide-angle beam scanning, and high integration density for practical 5G millimeter-wave applications.  Methods  A structured design method is adopted. First, a stacked differential balun based on a stripline-slot-stripline configuration is developed to achieve efficient single-ended-to-differential conversion. A single-polarized magnetoelectric dipole antenna element is then designed and integrated with the balun, and its performance is characterized. The design is further extended by orthogonally integrating two elements to form a dual-polarized unit, which is used to construct a 1×4 linear array. Iterative full-wave electromagnetic simulation and optimization are conducted to balance wideband impedance matching, high port isolation, stable wide-angle beam scanning, grating-lobe suppression, and mutual-coupling reduction.  Results and Discussions  The optimized 1×4 dual-polarized differentially fed magnetoelectric dipole antenna array uses an element spacing of 4.6 mm, corresponding to 0.4 free-space wavelength at 26 GHz. This spacing achieves a favorable balance between grating-lobe suppression and inter-element mutual-coupling reduction. The measured –10 dB reflection coefficient bandwidths are 25~29.4 GHz for the +45° polarization port and 25~27.7 GHz for the –45° polarization port (Fig. 20). The slight matching difference is attributed to the incomplete structural symmetry of the baluns under the two polarization modes (Fig. 13). At 26 GHz, both polarization modes provide a peak gain of 10.7~11 dBi and support ±60° wide-angle beam scanning, with main-lobe gain attenuation no greater than 3 dB (Fig. 21). The measured radiation performance agrees well with the simulated results. Minor deviations are mainly caused by the high dimensional sensitivity of millimeter-wave structures and small errors in fabrication and test assembly. The array also maintains stable low cross-polarization and high port isolation across the operating band. These results are achieved through equal-length feed lines, symmetric layout, ground-pad shielding, and metallized-via electromagnetic isolation (Fig. 16), which suppress mutual coupling and parasitic radiation and ensure consistent dual-polarized radiation performance.  Conclusions  This paper presents a dual-polarized magnetoelectric dipole antenna array with differential feeding for 5G millimeter-wave applications. By using a stacked stripline-slot-stripline balun and optimizing the radiating structure and array layout, the design achieves wide bandwidth, high gain, low cross-polarization, and wide-angle beam scanning. The differential balun enables efficient single-ended-to-differential conversion with good amplitude and phase balance across the target band. The implemented 1×4 array, with an optimized element spacing of 4.6 mm, achieves a simulated peak gain of 11 dBi at 26 GHz and supports ±60° beam scanning, with gain variation below 3 dB. The overall design verifies the feasibility of a differentially fed magnetoelectric dipole architecture for compact, high-performance 5G millimeter-wave terminal antenna modules. Future work may focus on larger array configurations and further integration with BeamForming Integrated Circuits (BFICs).
A Behavioral Economics-Based Game Model for Side-Channel Security Attack and Defense Strategies
CAI Juesong, YAN Yingjian, WANG Jindong
Available online  , doi: 10.11999/JEIT260121
Abstract:
  Objective  The field of side-channel security currently lacks a systematic, quantifiable, and reproducible model for guiding the selection of attack and defense strategies, particularly in real-world engineering contexts where resource constraints necessitate informed cost-benefit trade-offs. The absence of such a framework impedes the practical realization of the “appropriate security” principle, often leading to either over-protection or under-protection of cryptographic modules. Traditional approaches to evaluating attack and defense costs rely heavily on subjective expert judgments, which are inherently arbitrary, difficult to replicate, and lack a structured multi-dimensional assessment. To bridge this critical gap, this research proposes an interdisciplinary model that integrates game theory, behavioral economics, and the Analytic Hierarchy Process (AHP). The primary objective is to establish a holistic decision-support system that not only quantifies the multi-faceted costs of various side-channel strategies but also incorporates the psychological dimensions of decision-making under risk, thereby enabling dynamic and economically rational security strategy selection tailored to specific asset values and security levels.  Methods  This study constructs a multi-layered modeling framework based on a static non-cooperative game with incomplete information. First, the side-channel analyst and the defense designer are formally defined as rational players, each possessing a finite set of strategies: the attacker may choose from non-modeling attacks such as DPA/CPA, modeling-based attacks like template attacks, or emerging deep learning-based side-channel analysis; the defender may adopt countermeasures including time hiding, amplitude hiding, or masking/blinding techniques. To systematically quantify the often-overlooked cost dimension, an AHP-based structured cost model is introduced. Through pairwise comparison matrices, the model decomposes costs into multiple criteria—such as time, data storage, computational resources, expertise, and hardware overhead—and assigns objective weights to each criterion, thereby replacing subjective cost estimates with a reproducible, hierarchical evaluation system. Furthermore, to reflect real-world decision-making behavior, key concepts from behavioral economics are integrated: Prospect Theory models how gains and losses are perceived relative to a reference point, while risk aversion coefficients capture players’ tolerance for uncertainty. These behavioral parameters are explicitly linked to the security level of the cryptographic module, allowing the model to adapt to different operational contexts. The resulting behavioral-augmented Bayesian game is then solved using the concept of Bayes-Nash Equilibrium, wherein each player’s optimal mixed strategy is derived based on their private type (behavioral profile) and beliefs about the opponent. To ensure engineering relevance, the As Low As Reasonably Practicable principle is incorporated as a constraint, enforcing that any selected defense strategy must be justifiable in terms of risk reduction versus cost incurred. Numerical solutions are obtained via a customized sequential quadratic programming algorithm implemented in Python.  Results and Discussions  A comprehensive experimental evaluation was conducted to validate the proposed model’s consistency, sensitivity, and practical utility. The AHP-based cost quantification demonstrated strong internal consistency, with all consistency ratios below the 0.1 threshold, confirming the reliability of the judgment matrices. The derived weight distributions revealed intuitive priorities: attackers placed greater emphasis on technical barriers and computational cost, whereas defenders prioritized design complexity and performance overhead. The behavioral adjustment layer successfully modulated perceived costs according to security levels: under low-security conditions (high risk aversion), costs were perceptually inflated, leading to conservative strategy choices; under high-security conditions, decision-making aligned more closely with objectively quantified costs. Equilibrium analysis across varying asset values and security levels yielded interpretable and rational strategy profiles. For low-value assets, both players exhibited a strong tendency toward low-cost or “no action” strategies, adhering to the lower bound of the ALARP region. As asset value increased, a clear threshold effect was observed, triggering a shift toward high-cost, high-efficacy strategies such as deep learning-based attacks and masking-based defenses. Sensitivity analysis further confirmed that defense strategy probabilities increased monotonically with asset value, validating the model’s ability to capture the non-linear relationship between protection intensity and asset criticality. These findings underscore the model’s capacity to support context-aware, adaptive security decision-making that balances risk, cost, and psychological factors.  Conclusions  This research presents a novel, behaviorally informed game-theoretic model for side-channel security strategy selection, addressing a significant void in existing literature regarding structured cost-benefit assessment. By integrating AHP-based objective cost quantification, behaviorally adjusted subjective valuations, and ALARP-driven engineering constraints, the proposed framework offers a multi-dimensional, reproducible, and context-sensitive tool for analyzing attack-defense interactions. The model advances the field by explicitly linking security levels to behavioral parameters, enabling dynamic strategy adaptation in response to both asset value and decision-makers’ risk perceptions. Although the current implementation relies partially on expert-defined parameters and operates within a static game setting, it establishes a critical foundation for transitioning from heuristic-based security decisions to quantitatively grounded, interdisciplinary analysis. Future work will focus on parameter calibration using real-world attack/defense datasets, extension to multi-stage dynamic games to capture strategic evolution over time, and empirical validation in industrial cryptographic evaluation scenarios. This study contributes the first systematic methodology for cost-aware, behaviorally realistic strategy optimization in side-channel security, offering both theoretical insights and practical guidance toward achieving “appropriate security” in cryptographic engineering.
An Overview of Key Technologies for 6G-Enabled Communication-Computing Integration and Energy-Efficiency Optimization
LIU Guangyi, CAI Qing, WANG Xinyao, CHEN Tianjiao, JIN Jing, XUE Yahui, WANG Ailing, WANG Hanning
Available online  , doi: 10.11999/JEIT260399
Abstract:
  Significance   Constrained by size, power consumption, and cost, emerging intelligent terminals often face excessive energy consumption and limited battery life. These limitations have become major bottlenecks to large-scale deployment. Compared with Fifth-Generation (5G) wireless networks, Sixth-Generation (6G) wireless networks are expected to enhance the Radio Access Network (RAN) architecture and move computing capability toward the RAN side. High-energy and compute-intensive Artificial Intelligence (AI) tasks that are originally executed by end devices can therefore be processed by the network. Through End-Edge Collaboration, emerging intelligent terminals can be upgraded toward lightweight design, low cost, and long battery life, thereby supporting the large-scale deployment of ubiquitous intelligence in 6G networks.  Progress   Current progress in Terminal Energy Consumption Optimization through 6G End-Edge Collaboration is reviewed, with emphasis on local execution, full offloading, and partial offloading. In local execution, User Equipment (UE) processes all tasks locally, which leads to high computing energy consumption. In full offloading, all tasks are transferred to the RAN. This reduces terminal-side computing energy consumption but can increase transmission energy consumption, especially under poor channel conditions. Partial offloading combines the benefits of both modes and optimizes energy consumption according to real-time network conditions. For partial offloading, four representative optimization techniques are summarized. (1) Feature Extraction and Filtering. Semantic encoding and information extraction are performed at the UE, and only task-relevant data are transmitted to the RAN. This reduces redundant data transmission and lowers transmission energy consumption. (2) Split Offloading. A large Deep Neural Network (DNN) is divided into layers according to its structure. Simpler shallow layers are processed at the UE, whereas more complex deep layers are offloaded to the RAN. This method balances terminal-side and RAN-side computational loads through End-Edge Collaborative Inference. (3) Model Lightweighting. Model complexity is reduced through pruning, quantization, and knowledge distillation, which lowers computational overhead while maintaining task performance. (4) Incremental Inference. Only changed data or updated features are processed, while historical computations are reused. This reduces redundant computation. Together, these techniques improve terminal performance and energy efficiency within the 6G End-Edge Collaboration framework.  Conclusions  This paper systematically reviews Terminal Energy Consumption Optimization for 6G End-Edge Collaboration. It summarizes the functional evolution of enhanced RAN, constructs an End-Edge Collaborative service framework for Communication-Computing Integration, and establishes a theoretical model of terminal computing energy consumption and transmission energy consumption. The composition and influencing factors of energy consumption under different offloading modes are clarified. Key energy optimization technologies, including Feature Extraction and Filtering, Split Offloading, Model Lightweighting, and Incremental Inference, are then discussed. To address energy consumption fluctuations caused by dynamic wireless channels, the paper proposes energy optimization mechanisms based on Adaptive Semantic Compression, Dynamic Split Offloading, Adaptive Model Pruning, and Incremental Inference. These mechanisms maintain a dynamic balance between energy optimization and task performance. Using embodied intelligent robot video understanding as a typical application scenario, a test platform is developed to verify the effectiveness of the proposed mechanisms. Current challenges and future research directions are also analyzed.  Prospects   Although End-Edge Collaborative energy-saving technologies have achieved initial progress, practical deployment still faces challenges in real network environments, dynamic wireless channels, and large-scale user access. Future research should examine the trade-off between optimization overhead and system robustness. It should also study dynamic communication-computing resource substitution modeling in stochastic resource environments, multi-user collaboration strategies, and global energy-efficiency optimization. As the technology matures, standardization and engineering implementation of End-Edge Collaborative energy-saving frameworks will become critical to the large-scale adoption of 6G applications. Future studies should further integrate algorithm design with network architecture, enabling practical deployment of low-power and high-efficiency intelligent communication systems.
Gating Adaptive Repeat Query Framework for Reliable Collaborative Inference with Edge Heterogeneous LLMs
WANG Tengsheng, YU Tao, LI Jihong, ZHENG Guhan, ZHANG Shunqing
Available online  , doi: 10.11999/JEIT260218
Abstract:
  Objective  Reliable inference at the network edge is indispensable for 6G-enabled ubiquitous AI, yet the deployment of large language models (LLMs) in such environments remains a cornerstone challenge. In resource-constrained edge settings, single-LLM inference often proves unreliable due to knowledge limitations and inherent biases, severely hampering real-world deployment. Collaborative inference leveraging multiple heterogeneous LLMs emerges as a promising remedy to boost robustness, but it introduces nontrivial hurdles under stringent latency and energy budgets, especially when wireless channel conditions and query content vary unpredictably. These challenges include the need for dynamic sequential decision-making for LLM selection and resource allocation, the fundamental paradigm mismatch between bit-level reliability protocols and semantic-level error correction, and the lack of adaptive mechanisms to align and fuse disparate LLM outputs effectively. To fill these critical gaps, this paper presents a novel framework that fundamentally reinterprets collaborative inference as a semantic-driven, closed-loop process, thereby transitioning from conventional bit-retransmission to semantic-retransmission and offering a practical path toward reliable 6G edge intelligence.  Methods  In response to these critical challenges, we propose the Gating Adaptive Repeat Query (G-ARQ) framework. Its core innovation is the Semantic-Space Alignment and Error-Guided Retransmission (SEMAR) mechanism. SEMAR first aligns the token-level probability distributions from heterogeneous LLMs into a unified semantic space using relative representation, enabling comparable outputs. It then models the collaborative process probabilistically, explicitly capturing error dependencies among models, and uses an Expectation-Maximization (EM) algorithm to infer a latent error direction, which guides the selection of the next LLM for query retransmission, steering it towards outputs orthogonal to previous errors. To jointly optimize the LLM gating and uplink power allocation under communication constraints without requiring explicit system dynamics—often unavailable in practice—we design a black-box trajectory optimizer. This optimizer formulates the sequential decision problem as sampling from a target distribution that encodes dynamic feasibility, optimality, and constraints. It employs a diffusion-based sampling process with a model-guided prior and Monte Carlo estimation to generate near-optimal policy trajectories that satisfy hard latency and energy limits.  Results and Discussions  To evaluate the practical viability of G-ARQ under realistic edge conditions, simulations are conducted in a scenario with five base stations hosting five heterogeneous 7B-parameter LLMs (Mistral-7B, Vicuna-7B, Nous-Capybala-7B, Gemma-7B, and Llama-2-7B). The user equipment (UE) performs a question-answering task evaluated on a mixed SQuAD and TriviaQA dataset. Component-level evaluations, each designed to isolate the contribution of a single innovation, validate the effectiveness of every key element. The error-guided gating of SEMAR, compared to a Top-k gating baseline, improves accuracy by 0.23 % on average, and its dynamic weight ensemble contributes an additional 0.7 % gain (Fig. 3). The black-box trajectory optimizer, which operates without any explicit channel model, achieves accuracy close to that of the unconstrained model-greedy strategy while ensuring strict latency constraints (Fig. 4). The convergence of the optimizer is verified by tracking the evolution of \begin{document}$ J(\boldsymbol{S}) $\end{document} and selection probability over diffusion steps (Fig. 5). System-level performance under varying latency and energy constraints demonstrates that G-ARQ consistently surpasses two baselines: one combining model-greedy selection with Proximal Policy Optimization (PPO) for power optimization, and another combining model-greedy selection with Simulated Annealing for power optimization, both using average output weights. The accuracy improvement is most significant under the most stringent resource limits, reaching up to 2.2 % for \begin{document}$ {K}_{\max }=1 $\end{document} and 1.9 % for \begin{document}$ {K}_{\max }=2 $\end{document} (Fig. 6, Fig. 7). The framework successfully establishes a Pareto boundary that characterizes the inherent trade-off between inference accuracy and communication latency, providing a valuable design guideline for resource-constrained edge systems and offering actionable insights for real-world deployment. The GARQ-S variant is noted to outperform GARQ-E by avoiding the integration of outputs from previously erroneous models.  Conclusions  This paper proposed G-ARQ framework, an innovative closed-loop framework that transforms collaborative edge inference into a semantics-guided retransmission process. By introducing SEMAR for error-based alignment and selection of heterogeneous LLMs, and employing a black-box trajectory optimizer for joint model selection and power allocation, the framework achieves up to a 2.2% accuracy improvement under strict resource constraints. The results validate G-ARQ as an effective and practical approach toward reliable and efficient 6G edge intelligence.
Queue Stability Constrained Robust Secure Beamforming for Low-Altitude UAV-ISAC Systems
OUYANG Jian, REN Wei, XU Ba, LIU Xiaoyu, JIANG Wanmu
Available online  , doi: 10.11999/JEIT260275
Abstract:
  Objective  To address the challenges of antenna array angle errors caused by UAV jitter, transmission instability induced by random data arrivals, and secure transmission guarantee in multi-eavesdropper scenarios for low-altitude UAV-ISAC systems, this paper proposes a robust secure beamforming algorithm based on queue stability constraints. The proposed algorithms aims to minimize long-term transmit power consumption while maintaining data queue stability while enhancing beamforming robustness against UAV jitter.  Methods  To guarantee the stability of the wireless transmission in UAV-ISAC systems, this paper formulates an optimization problem aimed at minimizing the long-term average transmit power, subject to constraints on data queue stability, secure communication rate, sensing performance, and maximum transmit power. Since this long-term optimization problem is intractable, the Lyapunov optimization framework is employed to transform it into a sequence of short-term subproblems. To handle the antenna array angle errors caused by UAV jitter within each short slot, we jointly adopt the second-order Taylor series expansion and the S-Procedure method to approximate the short-term subproblem into a convex form. Consequently, a robust secure beamforming algorithm based on penalty successive convex approximation optimization is proposed.  Results and Discussions  Simulation results demonstrate the impact of the number of antennas, secrecy rate threshold, beam gain threshold, UAV jitter error, and Lyapunov weight factor on the system transmit power. As illustrated by the beam gain pattern in Fig. 3, the communication beamformer facilitates cooperative sensing toward the sensing area, while the sensing beamformer enhances system security by directing interference toward potential eavesdroppers. This validates the effectiveness of the proposed algorithm in simultaneously improving sensing and secure communication performance. Furthermore, leveraging the dual-function characteristics of the communication and sensing beamformers, the proposed integrated sensing and communication scheme achieves significantly higher resource utilization efficiency than the communication-only and sensing-only schemes, as shown in Fig. 5. Additionally, Fig. 6 indicates that, compared with the non-robust scheme, the proposed robust scheme strictly satisfies security requirements under various angle errors. Finally, Fig. 8 shows that the proposed queue-aware scheme can effectively suppress the transmit power fluctuations caused by random data arrivals, exhibiting superior stability compared to the queue-free baseline.  Conclusions  This paper investigates a robust secure beamforming method for UAV-ISAC systems subject to queue stability constraints. First, based on the Lyapunov optimization framework, the long-term stochastic optimization problem is transformed into a sequence of short-term subproblems. Second, to address the issue of UAV jitter, the second-order Taylor series expansion and the S-Procedure method are jointly employed to approximate the non-convex constraints into tractable convex forms. Finally, a robust secure BF optimization algorithm based on penalty successive convex approximation is proposed to efficiently solve the deterministic short-term subproblems. Simulation results demonstrate that the proposed scheme can effectively tackle the challenges posed by random data arrivals, UAV jitter, and eavesdropping threats, thereby ensuring the stability and security of downlink data transmission in low-altitude UAV-ISAC systems.
VT2R: Video and Text-driven Method for Generating Large-scale Millimeter-wave Radar Data
DENG Kaikai, LING Yue, XING Ling, WU Honghai, ZHAO Dong, MA Huahong
Available online  , doi: 10.11999/JEIT260240
Abstract:
  Objective  The lack of large-scale training data impedes progress in developing robust and generalized deep learning models. However, existing millimeter-wave radar data generation methods are ineffective due to a lack of sufficient data sources. To address this gap, this paper proposes a video and text-driven radar data generation method, VT2R, which utilizes video or text data to generate large-scale, realistic radar data, solving the key problem of constructing the mapping relationship between video and text and radar data.  Methods  The proposed method consists of three main components: video feature encoding network, text feature encoding network, radar feature encoding network and data fitting and decoding network. Video feature encoding networks and text feature encoding networks extract temporally consistent visual representations and alignable semantic features, respectively, while the radar encoding network learns the structure and dynamic information of point clouds through hierarchical spatiotemporal modeling. In the data fitting and decoding network based on Variational AutoEncoder (VAE), multi-modal features are mapped to a unified latent distribution space and decoded into radar data through reparameterized sampling. During training, reconstruction loss, Kullback-Leibler (KL) divergence loss, and cross-modal similarity loss are jointly optimized.  Results and Discussions  This paper constructs the first radar point cloud dataset for reclining gesture recognition (Figs. 6 and 7), covering 5 gesture categories, 32 participants, and a total of 14,400 samples. Experimental results based on this dataset show that VT2R achieves a recognition accuracy of 89.2% when trained using only generated radar data, a 33.88% improvement over the representative RFGen (Figs. 9 and 10). When combined with a small amount of real radar data for joint training, the accuracy further improves to 97.62%, a 21.48% improvement over RFGen (Figs. 9 and 11). Furthermore, VT2R still achieves average recognition accuracies of 89.35% and 97.21% under different scenarios and factors (Figs. 16-18). In addition, this paper also verifies the accuracy of VT2R under different postures, achieving average accuracies of 89.98% and 97.55% in the first and third settings, respectively (Fig. 19), which is basically consistent with the result obtained when lying down, demonstrating its robustness under cross-posture conditions.  Conclusions  This paper proposes a radar data generation system, VT2R, which addresses the severe lack of realistic radar training data when users are performing gestures in a lying position. Through a video feature encoding network built on a vision-language pre-trained model, a text encoding network incorporating cue templates, a hierarchical radar encoding network for sparse point clouds, and a VAE-based data fitting and decoding network, these components collaboratively generate large-scale, realistic radar data. It also supports augmented reconstruction based on limited real radar data, providing rich data support for radar perception tasks. Future work will focus on solving multi-modal data generation for more complex gesture scenes, providing better data support for emerging large-scale models.
Multi-task Lightning Nowcasting with Spatio-temporal Focal Perception and Synergistic Weighted Loss
TANG Zhihao, HAN Yuanpeng, ZHANG Hui, SONG Lin, ZHANG Qilin, LIU Yi
Available online  , doi: 10.11999/JEIT260234
Abstract:
  Objective  Lightning nowcasting is essential for early warning systems and for protecting critical infrastructure, including aviation, power grids, and transportation systems. Traditional numerical weather prediction models depend strongly on parameterization schemes and require high computational costs, which limits their use in rapid-update nowcasting. Although deep learning methods have advanced, they still have difficulty handling extreme data sparsity, suffer from serial-computation bottlenecks in recurrent architectures, and mainly focus on binary occurrence prediction rather than the joint optimization of lightning-frequency prediction and regional localization. Moreover, conventional loss functions are easily dominated by extensive non-lightning areas, which biases predictions toward zero or causes excessive false alarms. To address these limitations, a Spatio-Temporal Focal perception and synergistic weighted loss Network (STF-Net) is proposed as a multi-task lightning nowcasting model that jointly predicts lightning frequency and occurrence regions. It integrates three key components: a Lightning Adaptive Attention Module (LAAM) for explicit spatio-temporal dependency modeling, a Spatio-Temporally Weighted Hybrid Loss for data sparsity and imbalance, and a spatio-temporal dual-branch Generative Adversarial Network (GAN) to improve prediction fidelity and temporal coherence.  Methods  STF-Net is built on the SimVP video prediction architecture and adopts an encoder-translator-decoder paradigm. LAAM uses a three-dimensional decoupled attention mechanism along the height, width, and channel dimensions, enabling adaptive focus on convectively sensitive regions while maintaining computational efficiency. The Spatio-Temporally Weighted Hybrid Loss combines Temporally Weighted Mean Squared Error (TW-MSE) for frequency regression and Dual-Weighted Cross-Entropy loss (DWCE) for regional localization. Time-increasing weights are incorporated to improve medium- to long-term forecast robustness (Fig. 5). DWCE integrates static class weights with dynamic grid weights, thereby balancing global class proportions and local lightning-frequency heterogeneity. A spatio-temporal dual-branch GAN, consisting of a spatial PatchGAN discriminator and a temporal three-dimensional convolutional discriminator, is used to improve the textural fidelity and temporal coherence of predicted lightning-frequency fields. The model uses six consecutive historical lightning-frequency frames at 256×256 resolution and 10-min intervals to predict the next six frames, corresponding to a 1-h forecast window. Experiments are conducted on a high-resolution Very Low Frequency Lightning Location Network (VLF-LLN) dataset containing 11,748 images that cover different seasonal and weather conditions. The dataset is split at a ratio of 7:3 for training and testing.  Results and Discussions  Comprehensive evaluation metrics are used, including Mean Squared Error (MSE) and Mean Absolute Error (MAE), which are computed only on lightning pixels; Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) for image fidelity; and Probability Of Detection (POD), False Alarm Rate (FAR), and Critical Success Index (CSI) for regional detection. STF-Net achieves a CSI of 0.663 within the 1-h forecast window, representing a 14.5% improvement over the SimVP baseline (0.579). It also reduces FAR from 0.351 to 0.216, corresponding to a relative reduction of 38.5% (Table 1). Ablation studies validate each component. Adding GAN improves CSI to 0.624 and reduces MSE to 0.109. Further incorporation of LAAM increases CSI to 0.629 and yields the highest POD of 0.894. The complete STF-Net with the hybrid loss achieves the best performance, with a CSI of 0.663 and an MSE of 0.105 (Table 1, Fig. 5). The joint prediction of frequency and region is supported by simultaneous improvements in regression metrics (MSE and MAE) and detection metrics (CSI and FAR). Time-step analysis shows that LAAM reduces long-term performance degradation, with STF-Net maintaining the highest CSI compared with SimVP+GAN and SimVP (Fig. 6). Comparative experiments with ConvLSTM and PredRNN further demonstrate the superiority of STF-Net across all lead times. STF-Net consistently achieves higher CSI and lower FAR, and its advantage becomes more evident as the forecast horizon increases (Fig. 7). Its consistent gains in PSNR and SSIM further indicate that the spatio-temporal GAN helps generate coherent, detail-rich predictions. Visualization results show that STF-Net produces structurally clear and continuous lightning-activity regions centered on high-frequency areas. It accurately tracks dynamic evolution patterns, including movement, merging, and splitting, while generating minimal noise in non-lightning regions. These results demonstrate effective collaborative prediction of both frequency magnitude and spatial distribution.  Conclusions  STF-Net is presented as a deep learning model for joint lightning-frequency prediction and regional localization. It explicitly models long-range spatio-temporal dependencies, focuses on critical convective zones, addresses extreme data sparsity and class imbalance, and jointly optimizes frequency regression and regional localization while suppressing false alarms. Its spatio-temporal dual-branch GAN further improves the spatial structural consistency and temporal coherence of predictions. Experimental results show that STF-Net outperforms baseline and state-of-the-art models, achieving a CSI of 0.663, an FAR of 0.216, and the best MSE and MAE values within a 1-h forecast window. The model effectively reduces long-term performance degradation, captures the evolution trends of lightning regions, and generates physically plausible predictions with minimal background noise. This study provides an efficient end-to-end solution for operational lightning nowcasting systems and offers guidance for model design in sparse meteorological spatio-temporal sequence prediction.
Intelligent Privacy-Aware Computation Offloading Method against Multi-server Joint Inference Attacks
MIN Minghui, LIU Mingcheng, ZHANG Peng, DUAN Jincheng, LI Shiyin, ZHANG Hongliang
Available online  , doi: 10.11999/JEIT260249
Abstract:
  Objective  With the rapid development of the low-altitude economy, services such as intelligent transportation, smart healthcare, and low-altitude logistics have become increasingly common. Their efficient operation depends on the real-time processing of massive sensing data. Mobile Edge Computing (MEC) improves task execution efficiency and reduces device computational burdens by offloading tasks to nearby servers. However, user privacy and security risks have become increasingly severe. In dynamic scenarios where multiple MEC servers jointly process tasks, information sharing can enable multi-server joint inference attacks and greatly increase the risk of user location privacy leakage. Although existing studies have used Differential Privacy (DP) to protect user location privacy, current DP-based solutions remain limited. These methods inject noise into offloading decisions, but unconstrained noise may reduce task allocation accuracy. In addition, user mobility causes continuous changes in channel states during dynamic computation offloading. Privacy leakage risks and attacker behaviors are also uncertain. Traditional optimization methods based on static system models are therefore unsuitable for such dynamic environments. To address these challenges, this paper proposes an Asynchronous Advantage Actor-Critic (A3C)-based Intelligent Privacy-Aware Computation Offloading (AIPCO) scheme against multi-server joint inference attacks. The proposed scheme protects user location privacy while maximizing the overall utility of the MEC system.  Methods  This paper proposes a DP-based task offloading rate perturbation mechanism. By adding controlled noise, the mechanism increases the randomness of user task offloading toward multiple MEC servers. A truncated Laplace mechanism is used to constrain the perturbed offloading rates within valid boundaries. This design satisfies the mathematical guarantees of DP and reduces the accuracy of multi-server joint inference attacks on sensitive user locations. Privacy entropy is then introduced to dynamically evaluate the real-time privacy protection level. Finally, the AIPCO scheme is constructed. Through a multi-threaded asynchronous training mechanism, the scheme interacts with the environment through iterative trial and error and efficiently learns the optimal real-time offloading policy online. The proposed scheme dynamically protects user privacy, reduces computational cost, and maximizes comprehensive system utility.  Results and Discussions  The AIPCO scheme jointly optimizes user privacy and task offloading cost by incorporating multidimensional performance variables into the reinforcement learning reward function. A comprehensive performance analysis (Fig. 4) shows that, when the number of continuous learning iterations reaches 200, the privacy protection level of AIPCO increases by 2.52%, 3.56%, and 22.90% compared with RCLM, JODRL, and DODA-DT, respectively. This advantage is mainly attributed to the DP-based task offloading rate perturbation method, which uses the truncated Laplace mechanism to increase data randomness while strictly constraining the perturbation range. By contrast, RCLM perturbs the task offloading rate through range-limited DP without using the truncated Laplace mechanism. JODRL increases randomness only through network policy optimization, resulting in a lower privacy protection level. DODA-DT focuses on balancing energy consumption and system latency without optimizing user privacy. For the privacy weight parameter (Fig. 5), increasing $\omega$ improves privacy protection. For example, the privacy protection level increases by 5.64% when $\omega$ rises from 0.2 to 0.7, with a clear performance gain at 0.7. As the system agent reduces its focus on computational cost, user utility remains optimal despite increased cost. When the physical distance between users and the MEC server is adjusted (Fig. 6), AIPCO shows stronger privacy protection in long-distance scenarios. A greater distance reduces the number of tasks offloaded to the server. Therefore, attackers obtain less information, and privacy protection improves. Although computational cost increases with distance, AIPCO consistently outperforms competing schemes. These results confirm that AIPCO achieves optimal MEC system utility while protecting user privacy.  Conclusions  To mitigate multi-server joint inference attacks caused by information sharing among collaborative MEC servers, this paper proposes an AIPCO method. A DP-based task offloading rate perturbation scheme is designed to increase randomness, and a truncated Laplace mechanism is used to constrain perturbed rates within reasonable boundaries. The scheme is proven to satisfy strict DP mathematical guarantees, and privacy entropy is introduced to quantitatively evaluate the privacy protection level. In addition, the AIPCO scheme uses a multi-threaded asynchronous training mode, enabling the agent to efficiently learn the optimal perturbed offloading policy in a continuous space and maximize overall system utility. Simulation results show that the proposed scheme outperforms the baselines in both dynamic and average performance. It achieves optimal system utility while protecting user privacy.
Bearing Fault Diagnosis of Roadheader via Cross-modal Kernel Fusion-sphere Space Learning
SU Shuzhi, GUI Yang, MA Tianbing, ZHU Yanmin, WU Kanghui
Available online  , doi: 10.11999/JEIT260494
Abstract:
  Objective  Traditional roadheader bearing fault diagnosis methods often struggle with high-dimensional and nonlinear multi-sensor data. They also fail to effectively perceive cross-modal, multi-scale fault information or integrate local and global structural features. To address these limitations, this paper proposes a Cross-modal Kernel Fusion-sphere Space Learning (CKFSL) method. By perceiving cross-modal multi-scale fault information, CKFSL extracts highly discriminative features from roadheader bearing cross-modal fault samples and improves diagnostic accuracy.  Methods  CKFSL first maps roadheader bearing cross-modal fault samples into a high-dimensional kernel space through implicit transformation. Dual extremal point anchoring and polar neighbor allocation mechanisms are then used to capture fault sample clusters with similar isomorphic information, forming kernel fusion-spheres. An adaptive binary partitioning strategy is designed according to the geometric span of internal fault samples. This strategy tightens isomorphic boundaries, constructs micro-neighbor kernel fusion-spheres, and achieves highly isomorphic manifold aggregation at the microscopic scale. A micro-neighbor kernel fusion-sphere space is further formed to re-evaluate local isomorphism (Fig. 1). To characterize wide-area topological correlations, a wide-area topological isomorphism constraint is proposed. This constraint constructs a wide-area dynamic isomorphism graph among micro-neighbor kernel fusion-spheres (Fig. 1). Finally, an objective optimization function is formulated within the space learning framework. It integrates local manifold isomorphism and wide-area topological correlations of roadheader bearing cross-modal fault samples, as shown in the CKFSL diagnostic flowchart (Fig. 2). The analytical solution for spatial projection is theoretically derived to obtain discriminative cross-modal kernel fusion-sphere space isomorphic features from roadheader bearing cross-modal fault samples.  Results and Discussions  CKFSL is first validated on the self-built AUST roadheader bearing cross-modal fault dataset, with the experimental platform shown in Fig. 3. The average recognition rates obtained with increasing numbers of training fault samples are shown in Fig. 4. On the AUST dataset, CKFSL achieves a recognition rate of 99.49% with only 70 training fault samples and reaches 100% as the number of training fault samples increases. Table 1 summarizes the standard deviations under different training fault sample sizes. The results show that CKFSL has the lowest standard deviation and stronger robustness than the other seven comparison algorithms. Three-dimensional fault feature distributions are shown in Fig. 5. The results confirm that CKFSL effectively separates highly overlapping fault samples into different clusters and reduces the boundary confusion observed in the comparison algorithms. To verify generalization capability, CKFSL is further evaluated on the public Paderborn dataset, with the experimental setup shown in Fig. 6. As shown in Fig. 7 and Fig. 8, CKFSL achieves a 100% average recognition rate across four complex fault categories. It also outperforms the comparison algorithms, which have difficulty exceeding an 85% recognition rate for the F4 fault category.  Conclusions  CKFSL effectively addresses the inability of traditional roadheader bearing fault diagnosis methods to perceive complex multi-scale fault information. By using the wide-area dynamic isomorphism graph learned in the micro-neighbor kernel fusion-sphere space, CKFSL integrates local manifold isomorphism with wide-area topological correlations of roadheader bearing cross-modal fault samples. This process enables CKFSL to extract highly discriminative cross-modal kernel fusion-sphere space isomorphic features. It improves the accuracy of roadheader bearing fault diagnosis and supports the reliability and continuous operation of roadheaders.
A General Evaluation Framework for Mission Planning Algorithms for Remote Sensing Satellite Constellations
LI Jinfei, YU Xiaogang, TIAN Jing, HE Haochen, XING Xiangwei, ZHANG Xiaohan
Available online  , doi: 10.11999/JEIT260335
Abstract:
  Objective  The rapid growth in remote sensing satellite constellations has shifted mission planning from single-satellite static scheduling to large-scale dynamic coordination across heterogeneous constellations. However, evaluation methods have not kept pace with algorithm development. Existing studies often rely on private datasets, simplified metrics centered on Completion Rate, and idealized simulations that ignore realistic constraints, such as attitude maneuvers, illumination conditions, and dynamic task insertion. These limitations prevent fair cross-paper comparison and slow engineering application. To address this gap, this paper proposes the Remote Sensing Constellation Mission Planning Benchmark (RSCMP-Bench), a general, open, and reproducible evaluation framework. It is designed as a unified benchmark for the community, similar to ImageNet in computer vision and General Language Understanding Evaluation (GLUE) in Natural Language Processing (NLP).  Methods  RSCMP-Bench consists of three components. First, the multi-scenario standard task library contains 300 standardized scenarios at three difficulty levels: Low, Medium, and High, with 100 scenarios per level. Satellite numbers range from 30 to 200, and task demands range from 56 to 560. All scenarios are generated from public Two-Line Element (TLE) data and explicitly model realistic constraints. Optical satellites require a minimum solar elevation angle, and Synthetic Aperture Radar (SAR) satellites require incidence angles within specified ranges. General constraints, such as per-orbit maximum on-time, minimum single-operation on-time, attitude maneuver time, and valid execution windows, are also modeled. The scenarios include point tasks, area tasks, static tasks, and dynamically inserted tasks. Second, the multi-dimensional effectiveness evaluation system includes a Basic Performance layer and a Dynamic Adaptability layer. The Basic Performance layer uses Completion Rate, Weighted Completion Rate, Average Response Delay, and Time Utilization. The Dynamic Adaptability layer uses multi-stage rolling evaluation with random dynamic task insertion. The Dynamic Adaptability Score measures the post-insertion Completion Rate relative to the baseline, and Dynamic Response Efficiency measures the performance gain per unit replanning time. A composite RSCMP-Bench Score is also provided. Third, the simulation and evaluation platform uses a client-server architecture. It integrates a Simplified General Perturbations 4 (SGP4) propagator, algorithm adapters, two-stage constraint verification, an intelligent scenario generator, and visualization tools. The platform has been deployed at https://www.tianzhibei.com and has supported a national competition with more than 80 research teams.  Results and Discussions  Baseline experiments comparing Random Scheduler and Priority Greedy validate the feasibility, reproducibility, and discriminative capacity of RSCMP-Bench. Random Scheduler yields very low Completion Rates of 7.3%, 3.8%, and 1.9% on the Low, Medium, and High levels, respectively. These results confirm the extreme sparsity of the feasible solution space. Priority Greedy achieves higher Completion Rates but still degrades as scenario difficulty increases, decreasing from 76.1% at the Low level to 63.7% at the Medium level and 49.2% at the High level. These findings indicate that high-difficulty scenarios remain challenging even for reasonable heuristic methods. They also show considerable room for more advanced algorithms. The dynamic adaptability protocol quantifies algorithm robustness under unexpected dynamic task insertion, which is not captured by static evaluations. The two-stage constraint verification module rejects infeasible plans and generates detailed error reports to support debugging.  Conclusions   RSCMP-Bench provides a unified, fair, and reproducible benchmark for remote sensing constellation mission planning. By combining a public library of 300 standardized scenarios, a multi-dimensional effectiveness evaluation system based on Basic Performance and Dynamic Adaptability, and a simulation and evaluation platform with realistic constraints and automated scenario generation, the framework addresses the long-standing lack of standardized evaluation in this field. Baseline results confirm its discriminative capacity and reveal clear performance bottlenecks in large-scale dynamic scenarios. Inspired by ImageNet and GLUE, RSCMP-Bench can support systematic community evaluation and fair competition. The framework has been deployed at https://www.tianzhibei.com, and its adoption can accelerate progress in intelligent mission planning for next-generation remote sensing constellations.
Dual-MPC-Driven Modeling and Spatiotemporal Evolution of Intelligent Connected Traffic Risk Fields
JIANG Linyuan, DING Fei, FAN Xuan, YANG Xuechao, SONG Aiguo, ZHANG Dengyin
Available online  , doi: 10.11999/JEIT260194
Abstract:
  Objective  With the deployment of intelligent connected vehicle-road-cloud cooperative systems, roadside infrastructure is evolving from traffic-state sensing units into intelligent decision-support platforms for multi-vehicle interaction analysis and dynamic risk inference. In highway and urban freeway scenarios, traffic operation is affected not only by the kinematic responses of individual vehicles but also by lane-changing intentions, car-following competition, and conflict propagation under local interactions. From a roadside perspective, a unified framework is therefore needed to continuously represent traffic risk, reveal its spatiotemporal evolution, and couple risk information with behavior decision-making and trajectory planning. Existing car-following models, such as the Optimal Velocity Model (OVM), Full Velocity Difference (FVD) Model, and Intelligent Driver Model (IDM), can describe speed-spacing evolution. However, these models mainly focus on longitudinal interactions and usually embed risk implicitly in safety-distance or acceleration constraints. They cannot explicitly characterize the coupling between longitudinal following and lateral lane changing, nor can they provide a continuous risk representation suitable for regional traffic assessment. Although Artificial Potential Field (APF) methods and Model Predictive Control (MPC) methods can improve trajectory safety, existing studies still lack a unified mechanism that links risk assessment, behavior decision-making, and motion planning. In addition, discrete behavior choices and continuous control actions are difficult to process efficiently within a single optimization framework.  Methods  An intelligent connected traffic risk-field model oriented toward vehicle-road cooperation is first established. The model integrates vehicle-interaction risk, lane-marking constraint risk, and road-boundary repulsive risk (Fig. 1). In the vehicle-interaction layer, motion-state-induced risk is formulated by considering the relative speed and relative orientation between the ego vehicle and surrounding vehicles. Distance-induced risk is modeled to reflect attenuation as separation distance increases (Fig. 2(a)(b)). To represent the stronger influence of forward hazards than lateral and rear hazards, a directional non-uniformity coefficient is used. This coefficient adjusts the angular attenuation of field strength and enables anisotropic spatial risk representation around the vehicle. In the road-constraint layer, lane markings and road boundaries are modeled separately. The total driving risk field is obtained by weighting and combining the lane-marking field, road-boundary field, and multi-vehicle interaction field (Fig. 2(e)(f)). Based on this representation, a dual-MPC hierarchical decision and motion-planning architecture is designed (Fig. 3). In each control cycle, the upper layer evaluates candidate behavior modes according to the vehicle state, surrounding traffic state, and dynamic risk field, and then outputs a unique behavior mode. The lower layer activates the corresponding control branch. When lane keeping is selected, longitudinal speed-planning MPC is used. When lane changing is selected, lane-change trajectory-planning MPC is activated under road-boundary, lane-marking, and safe-gap constraints.  Results and Discussions  The proposed framework reconstructs microscopic traffic risk evolution under different datasets and time scales. In the HighD highway scenario, when the slicing interval is 0.4 s, the local evolution of following and lane-changing interactions is captured in detail. This includes the process in which the ego vehicle initially follows a preceding vehicle and then starts changing to the adjacent lane (Fig. 4(a)(d)). When the interval is increased to 1.0 s, a wider spatiotemporal interaction range becomes visible, and lane-change completion and the subsequent return maneuver are identified more clearly (Fig. 4(e)(h)). In the NGSIM scenario, a smaller interval provides a finer description of the lane-change disturbance process. By contrast, a larger interval reveals the wider reconstruction of interaction relationships among the original lane, target lane, and surrounding vehicles (Fig. 4(i)(p)). These results indicate that the proposed roadside-oriented risk field can describe both local interaction details and larger-scale evolution trends, depending on the selected reconstruction interval. Sensitivity experiments further confirm the role of the directional non-uniformity coefficient. As this coefficient increases, forward risk concentration becomes stronger, local peak risk increases, and the coverage of high-risk regions decreases (Table 2). This finding shows that the coefficient effectively regulates anisotropic field distribution. Comparative experiments with IDM, OVM, FVD, and APF show that the proposed method performs better in most representative scenarios and error metrics (Fig. 5, Table 3). In the lateral cut-in scenario, its advantage lies in the early representation of lateral intrusion risk, which enables the behavior decision layer to anticipate conflict and the motion-planning layer to generate continuous evasive actions. In congested scenarios, the superposition of forward congestion risk, lateral neighboring-vehicle influence, and road-boundary constraints allows the dual-MPC controller to evaluate safety and feasibility simultaneously in local space.  Conclusions  A unified framework for roadside-oriented traffic-risk modeling and behavior-driven trajectory planning is developed. By integrating multi-vehicle interaction risk, lane-marking constraint risk, and road-boundary repulsive risk into a continuously evolving dynamic risk field, the spatial quantification of multi-vehicle interaction risk is realized. The directional non-uniformity coefficient further enables asymmetric risk perception modeling in forward, lateral, and rear directions. On this basis, a dual-MPC hierarchical architecture is constructed to couple behavior decision-making with motion planning, so that lane-keeping and lane-changing behaviors can be adaptively selected and optimized under a unified risk-driven mechanism. Experiments based on HighD and NGSIM datasets show that the proposed method can effectively characterize the spatiotemporal evolution of traffic-risk fields and outperform representative comparison models in most typical scenarios and error metrics.
MGM-3DUNet: A Multi-scale Edge Semantic-guided GraphConvolutional Sequence Method for Brain Tumor Segmentation
ZHUANG Jianjun, LI Xiang, JING Shenghua, LÜ Zhenglong
Available online  , doi: 10.11999/JEIT260128
Abstract:
  Objective  Feature fusion in U-Net and its 3D variants mainly relies on simple single-scale concatenation, which limits the use of encoder features and weakens fine-grained segmentation of Tumor Core (TC) and Enhancing Tumor (ET) regions. Recent methods such as VM-UNet improve sequence modeling efficiency, but they mainly focus on global information modeling. Local detail preservation and edge enhancement remain insufficient. Therefore, current methods still have limitations in segmentation accuracy and clinical utility. To address these problems, this paper proposes MGM-3DUNet for brain tumor segmentation.  Methods  The Multi-Scale Edge semantic Guidance Module (MEGM) is designed to improve tumor boundary segmentation through learnable edge detection. The Graph Convolutional Sequence Module (GCSM) combines the local aggregation ability of graph convolution with efficient long-range modeling based on a Mamba-like structure. This design improves semantic consistency while preserving small tumor structures with fewer parameters. The Multi-scale Context Perception Module (MCPM) is introduced to strengthen feature complementarity across different tumor scales through dual-scale fusion.  Results and Discussions   Experiments show that the proposed method achieves better average Dice similarity coefficient (Dice) and 95th percentile Hausdorff Distance (HD95) than the comparison methods. With only 2.3M parameters, MGM-3DUNet achieves Dice values of 91.2%, 90.4%, and 89.2% for Whole Tumor (WT), TC, and ET, respectively. The visualization results (Fig. 9, Fig. 10) further show that MEGM improves boundary localization. Overall, the proposed method shows improved sensitivity to edge details and contextual correlations while maintaining a low parameter count.  Conclusions   This method improves tumor boundary prediction by introducing shallow-layer edge enhancement to emphasize tumor contours. Local and global semantic information is fused in the bottleneck layer, and multi-scale contextual features are integrated during decoding. The proposed design achieves accurate segmentation with low computational cost and is suitable for deployment on resource-constrained platforms.
A Low-latency Synchronization Header Detection Algorithm and Circuit for the JESD204C Interface
YIN Peng, ZHANG Chao, LEI Changan, HOU Weizhou, SHU Zhou, LIU Shubin, ZHU Zhangming
Available online  , doi: 10.11999/JEIT260163
Abstract:
  Objective  With rapid advances in high-speed electronics, front-end Analog-to-Digital Converters and Digital-to-Analog Converters (ADCs/DACs) continue to increase in sampling rate and resolution. Back-end Field-Programmable Gate Arrays and Application-Specific Integrated Circuits (FPGAs/ASICs) also provide stronger computing capability. These trends impose strict requirements on high-speed data interfaces, including high bandwidth, low latency, low power consumption, and reliable synchronization. As a mainstream high-speed Serializer/Deserializer (SerDes) interface, the JESD204C interface still suffers from long link initialization latency and high synchronization power consumption. These limitations restrict system real-time performance and energy efficiency. To address these issues, this study optimizes the link-layer design of the JESD204C receiver and proposes an efficient Synchronization Header (SH) detection method. The method implements exponential compression of the search set through global observation and iterative convergence. Detection efficiency is improved, fast and accurate SH positioning is achieved, link synchronization latency is reduced, and synchronization stability and energy efficiency are enhanced.  Methods  A typical JESD204C interface uses serial sliding detection for SH detection, which causes high link initialization latency and large delay jitter. To solve these problems, an Iterative Set Screening (ISS)-based SH detection algorithm is proposed. The SH detection task is modeled as the rapid localization of a deterministic pattern in a binary random sequence. A theoretical model based on information theory and stochastic processes is constructed. Expected space utilization and Bit Error Rate (BER) are introduced to support quantitative performance evaluation. In this model, SH candidate positions are defined as a dynamic set. Based on the inherent polarity inversion characteristic of the SH and global observations in each clock cycle, multilevel XOR logic is used to verify all candidate hypotheses in parallel. Non-inverting candidate positions are eliminated, and the search space is dynamically compressed. This design improves synchronization speed and position robustness, providing a low-latency and reliable initialization solution for high-speed SerDes links.  Results and Discussions  The proposed ISS-based SH detection algorithm is validated under harsh conditions, including SH crossing block boundaries and loss of lock caused by burst errors. The results demonstrate strong robustness, with rapid SH locking and link resynchronization under all test conditions (Figures 1116). To evaluate performance, four representative schemes are reproduced: a single-bit serial locking circuit, a 66-bit serial locking architecture, a register-intensive block synchronization method, and a parallel search circuit. A systematic comparison is then conducted between these schemes and the proposed design. The results show that the normalized locking time of the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method varies substantially with SH position (Figure 17(a)), especially at block boundaries (Figure 17(b)). When the SH is located at the Most Significant Bit (MSB), typical sliding detection requires about 1.8 times the time needed at the Least Significant Bit (LSB), indicating strong sensitivity to the starting position and search path. In contrast, the proposed ISS scheme maintains a stable normalized locking time within 1.0 ± 0.05 across all positions, with the standard deviation reduced by more than 70%. By evaluating all candidate positions equally through parallel filtering, the scheme eliminates position dependence. Synchronization can be completed within tens of clock cycles whether the SH is located at the LSB, the MSB, or any other position in the block. The experimental results verify that the ISS algorithm improves synchronization robustness and predictability while accelerating link initialization. Table 3 summarizes the performance metrics. The average locking time is only 24.4 clock cycles, representing an overall improvement of more than 70% compared with the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method. The standard deviation of locking time is only 4.7, indicating a more stable synchronization process. In terms of resource utilization, the design consumes 509 Look-Up Tables (LUTs) and only 2.0 mW, much lower than the 3 503 LUTs and 94.1 mW required by the register-intensive scheme. Its energy efficiency reaches 0.03 mW/bit, which is better than those of the three conventional methods. Compared with the parallel search circuit, the average locking time is reduced by 6.11%, power consumption is reduced by 50.3%, and energy efficiency is improved by 53.8%. Therefore, the proposed JESD204C receiver link shows advantages in SH detection speed, stability, power consumption, and energy efficiency.  Conclusions  An ISS-based SH detection algorithm is proposed for the JESD204C receiver. By screening the data stream in parallel through multilevel XOR logic, dynamically compressing the search space, and efficiently eliminating non-inverting candidate positions, the algorithm converges to the true SH position. This approach improves the conventional serial detection mechanism. The design is verified on the Xilinx KC705 FPGA platform. A Pseudorandom Binary Sequence 31 (PRBS31) is used to emulate the random distribution of polarity transitions, and a high-speed SubMiniature version A (SMA) cable is used for data loopback transmission. The results show that the algorithm achieves an average locking time of only 24.4 clock cycles, with a standard deviation as low as 4.7. High robustness is maintained for the SH at any position within the 66-bit block, and the energy efficiency reaches 0.03 mW/bit. The algorithm is superior to existing typical schemes in locking speed, delay stability, and energy efficiency. It provides a low-latency, reliable, and energy-efficient synchronization initialization approach for high-speed SerDes links.
Survey on Intelligent Semantic Covert Communication
FENG Zhaoxin, XU Yifan, XING Chengwen, XU Yuhua, ZHAO Nan, WANG Jinlong
Available online  , doi: 10.11999/JEIT260184
Abstract:
  Significance   As the Sixth-Generation mobile communication network (6G) evolves from the Internet of Everything to the Intelligent Internet of Everything, the communication paradigm is shifting from reliable bit transmission to effective semantic transmission. Semantic communication extracts and compresses task-related semantics to reduce redundancy and resource use. However, because semantic information is highly structured and task-specific, it is vulnerable to eavesdropping, inference, and attacks. Covert communication addresses this risk by hiding transmission behavior from unauthorized monitoring. With support from Artificial Intelligence (AI), covert communication can use reinforcement learning to adjust power and resource allocation in dynamic environments. Generative models can also conceal transmitted signals by learning and reproducing environmental patterns. However, strict covertness constraints limit the achievable transmission rate and make large-scale information transmission difficult. Intelligent semantic covert communication integrates semantic extraction with covert transmission, providing a reliable approach to secure and efficient 6G communications.  Progress   With the development of AI, especially deep learning for complex feature modeling, semantic communication can support efficient semantic extraction and nonlinear compression of multimodal data. Research on semantic communication has also shifted from Separate Source-Channel Coding (SSCC) to Joint Source-Channel Coding (JSCC), which supports end-to-end training and improved transmission performance. For image transmission, Convolutional Neural Networks (CNNs) use local receptive fields to capture spatial correlations. For sequential data transmission, Long Short-Term Memory (LSTM) networks use gating mechanisms to maintain temporal coherence. In covert communication, Generative Adversarial Networks (GANs) and diffusion models can learn the statistical patterns of environmental noise in the time, frequency, and spatial domains, thereby concealing transmitted signals. These methods reduce the effectiveness of unauthorized monitoring and detection, and improve system adaptability in dynamic environments. AI also improves autonomous decision-making in dynamic covert communication. By modeling covert transmission as a Markov Decision Process (MDP), Deep Reinforcement Learning (DRL) can learn resource allocation strategies through interaction with the environment. This approach reduces computational complexity compared with traditional convex optimization methods. By integrating semantic extraction and covert transmission, intelligent semantic covert communication further supports semantic-driven covert transmission. Large Language Models (LLMs) can evaluate semantic sensitivity and contextual risks, enabling selective covert transmission of sensitive semantic information.  Conclusions  Research on intelligent semantic covert communication shows the advantages of coordinated semantic perception and physical-layer covert mechanisms. AI improves semantic extraction efficiency and strengthens adaptation to dynamic and complex environments. By integrating semantic understanding with covert transmission strategies, intelligent semantic covert communication supports both efficiency and security for ubiquitous 6G services.  Prospects   Future research on intelligent semantic covert communication should address several key challenges, including AI-enabled detection, unified semantic metrics, lightweight model design, multimodal semantic alignment, system interpretability, and semantic hallucination. Active threat detection and adaptive defense strategies are needed to counter AI-driven surveillance. Causal reasoning in Large Multimodal Models (LMMs) can help mitigate semantic hallucination and improve data transmission reliability. Advances in model compression and cloud-edge collaboration are also needed to deploy high-complexity AI models on resource-limited terminals. With the rapid development of AI, intelligent semantic covert communication is expected to provide core support for intelligent connectivity of everything and help build more secure, efficient, and reliable 6G networks.
An Inverse-Hybrid-Modeling Digital Twin System for Natural Gas Energy Metrology
LIU Bin, ZHONG Lu, FENG Quanyuan, CHEN Yihong
Available online  , doi: 10.11999/JEIT260289
Abstract:
  Objective  Global natural gas consumption continues to increase at an average annual rate of 3.2%. A 0.1% reduction in energy measurement error can reduce trade disputes by approximately $750 million per year. Traditional studies mainly use indirect methods for energy measurement. Among these methods, chromatographic analysis and acoustic velocity correlation are the most widely used, but both have clear application limits. Chromatographic analysis has a low interference error, but it shows delayed dynamic response at high flow rates and limited dynamic calibration capability. It also has poor adaptability to multi-gas-source switching, requires manual calibration, and has high operation and maintenance costs. The lack of interoperability standards for energy networks further increases the difficulty of system integration. Acoustic velocity correlation provides a low-latency dynamic response for flow measurement, but it has a high interference error. This error may increase when the content of a single component changes, such as when the hydrogen content increases from 5% to 10%. The method may even fail under complex operating conditions, such as multi-gas-source mixing and dynamic pressure fluctuations. To address these issues, new mechanism-modeling-oriented methods have been developed. The two most representative directions are mechanism-modeling-driven methods and hybrid-modeling methods. Both methods combine multi-source data fusion with virtual-physical interaction to establish mechanism models that link flow rate, other parameters, and energy. These methods provide a new approach for accurate energy measurement, but new challenges remain. Mechanism-modeling-driven methods are usually based on static flow modeling using Computational Fluid Dynamics (CFD). However, their dynamic parameter updates are slow, with delays of more than 30 s. They also have difficulty adapting to real-time operating-condition changes, rely on large labeled datasets, and have limited interpretability. Hybrid-modeling methods still face unresolved problems in collaborative optimization across multiple modules. In addition, existing studies lack support from industrial-grade verification platforms. These limits restrict their ability to solve the dynamic response delay, parameter identification difficulty, excessive physical simplification, and weak interference resistance of traditional natural gas energy metrology methods under complex conditions. Based on recent progress in mechanism-modeling-driven and hybrid-modeling methods, this study proposes an inverse-hybrid-modeling-driven digital twin system. The system introduces a Variational AutoEncoder (VAE)-based operating-condition feature extraction algorithm and a Dynamic Bayesian Network (DBN)-based parameter calibration mechanism. It also uses a Variational Expectation-Maximization (VEM) algorithm for offline calibration. The proposed system aims to improve the accuracy, adaptability, and interference resistance of natural gas energy metrology under complex operating conditions.  Methods   A natural gas energy metrology digital twin system based on inverse hybrid modeling is proposed. The system is built on a three-tier “algorithm-system-scenario” architecture. It integrates calorific value, flow, and energy mechanism models with multi-source real-time data streams. The VAE is used for unsupervised mining of operating-condition features. A parameter self-correction loop is then constructed by combining the DBN with VEM-based system calibration. Industrial-grade devices, including ultrasonic flowmeters and gas chromatographs, are integrated to ensure real-time data transmission and closed-loop control. The system covers key operating conditions, including dynamic pressure fluctuations, hydrogen-blended gas mixtures, and multi-gas-source switching. This design ensures strong adaptability between the model and practical applications. The system was continuously verified for 25 weeks on a full-scale industrial-grade experimental platform. The results show an operational delay of ≤3.8 s, data transmission jitter of ≤0.5 s, average daily energy consumption per device of ≤1.2 kW·h, Mean Time Between Failures (MTBF) of ≥4 100 h, energy measurement error of ≤0.25%, calorific value error of ≤0.12%, and flow indication error of ≤0.2%. The system also meets security requirements through industrial Ethernet encryption and hierarchical access control. It provides engineering support for intelligent pipeline-network optimization and standardized integration.  Results and Discussions  First, a multi-level hybrid modeling framework is established. Modular hybrid modeling is achieved through the algorithm-system-scenario three-tier architecture. Numerical methods combined with data are more flexible than purely analytical models and can represent complex multiphysics systems with fewer lumped physical parameters. These parameters may change during energy measurement under mechanical, energy, and hydrodynamic effects. The VAE and DBN are used to deeply integrate mechanism models with real-time data. This reduces the parameter synchronization delay to 3.8 s and supports fluid-acoustic co-simulation and rapid response under complex operating conditions, such as hydrogen-blended natural gas. Second, an integrated algorithm for inverse hybrid modeling and system calibration is proposed. By incorporating the VAE, DBN, and VEM algorithm, the inverse hybrid modeling algorithm forms a self-supervised, adaptive intelligent system with an internal closed-loop operation. The VAE encoder compresses high-dimensional operating-condition data into low-dimensional feature vectors. This enables unsupervised feature extraction without large labeled datasets. Based on the learned internal data distribution, the VAE can also generate perturbed data similar to the input data. These data are used to simulate abnormal operating conditions and verify interference resistance. The DBN constructs a continuous “prior-evidence-posterior” iterative cycle to support system self-correction and adaptive response to operating-condition changes. The VEM algorithm compensates for systematic errors that are difficult for the DBN to capture, thereby overcoming the limits of traditional static models.  Conclusions  This study describes and validates a hybrid digital twin system that combines experimental data-driven methods with physical models. The system successfully simulates the physical characteristics of natural gas energy metrology. A full-scale test platform was constructed, and the main system parameters were validated using experimental measurement data and compared with industry benchmarks. Each independent module in the algorithm-system-scenario three-tier hybrid modeling architecture, including calorific value measurement, flow calculation, and energy conversion, was continuously verified for 25 weeks. The results confirm strong consistency between model predictions and actual measurements. On the natural gas energy metrology digital twin experimental platform, systematic validation was performed for three core functions: flow measurement under dynamic conditions, multi-component calorific value determination, and energy accumulation. The results show that the output of the digital twin model matches the physical device measurement data with an accuracy of more than 99.5%. Under complex operating conditions, such as pressure pulsations and hydrogen-blended gas mixtures, the system maintains the measurement error within 0.5%. This performance is better than that of traditional methods and meets the Class A accuracy requirements for natural gas measurement. By introducing a multi-tier hybrid modeling framework, this study addresses the parameter identification difficulty and excessive physical simplification of traditional natural gas energy metrology methods. The integration of the VAE, DBN, and VEM algorithm enables unsupervised feature extraction under complex operating conditions and adaptive calibration of model parameters. This reduces dependence on prior physical knowledge and large labeled datasets. The experimental results show that the proposed method maintains high precision and strong stability under complex scenarios, including pressure pulsations and hydrogen-blended gas mixtures, where traditional models have difficulty providing accurate descriptions.
Non-Terrestrial Network Architecture and Key Technologies for Civil Aviation
LIU Xiangnan, QIU Yu, HUANG Zhipeng, ZHANG Haijun
Available online  , doi: 10.11999/JEIT260348
Abstract:
  Significance   Civil aviation communication systems are entering a new stage of development driven by the rapid growth of global air transportation, the increasing demand for intelligent air traffic management, and the continuous expansion of in-flight connectivity services. Traditional civil aviation communication systems mainly rely on high frequency radio, high frequency radio, terrestrial air-to-ground links, and conventional satellite communication systems. These technologies have supported aircraft operation, air traffic control, airline operational communication, and low-rate data transmission for a long time. However, they still face limitations when applied to future civil aviation scenarios characterized by global coverage, high-speed mobility, low latency, high reliability, and service diversification. Particularly, terrestrial networks are difficult to deploy in transoceanic routes, polar regions, deserts, mountains, and remote airspace, while traditional geostationary satellite systems suffer from large propagation delay and limited capacity. Current systems cannot fully meet the requirements of continuous aircraft access, real-time flight monitoring, engine health data transmission, aviation safety communication, and passenger broadband services. Non-Terrestrial Networks (NTNs) provide a promising technical path for overcoming these limitations. By integrating GEOstationary satellites (GEO), Medium Earth Orbit satellites (MEO), low Earth orbit satellites (LEO), Very Low Earth Orbit satellites (VLEO), High-Altitude Platform Stations (HAPS), Unmanned Aerial vehicles (UAV), electric Vertical Take Off and Landing (eVTOL), and terrestrial infrastructures, NTN can construct a multi-layer air-space-ground integrated communication system. Such a system is able to provide continuous coverage, flexible deployment, resilient connectivity, and differentiated service support for civil aviation. NTN is becoming an important enabling technology for future civil aviation communication systems and for the digital and intelligent transformation of the aviation industry.  Progress   This paper reviews the development of NTN technologies for civil aviation and summarizes key research progress from three aspects: network architecture, access and mobility management, and resource management and scheduling. (1) We propose an aviation-oriented NTN networking framework composed of three layers: the satellite edge layer, the airborne core layer, and the terrestrial assistance layer. The satellite edge layer includes GEO, MEO, LEO, and VLEO satellites connected through inter-satellite links. GEO satellites are suitable for wide-area broadcasting and non-real-time services, MEO satellites can support navigation and intermediate-delay services, LEO satellites are suitable for low-latency and high-capacity broadband access, and VLEO satellites can further reduce propagation delay for future near-real-time aviation applications. The airborne core layer includes civil aircraft, HAPS, UAVs, and eVTOL platforms. HAPS can act as a regional relay, edge computing node, or software-defined control carrier, while UAVs and eVTOL platforms can provide flexible low-altitude coverage, emergency communication, and local access support. The terrestrial assistance layer consists of terrestrial base stations and gateway stations, which support air-to-ground communication and satellite-terrestrial interconnection. Civil aviation services can be divided into air traffic control and air traffic management services, airline operational control services, and airline passenger communication or in-flight entertainment services. Through network slicing, these heterogeneous services can be logically isolated and managed over a shared air-space-ground infrastructure. In congestion, rain attenuation, or shortened visibility-window scenarios, safety slices should be protected with the highest priority, while passenger service slices can be rate-limited, buffered, or degraded. (2) We analyze the characteristics of NR-NTN access and air-to-ground direct access in civil aviation. NR-NTN can provide continuous coverage for oceanic, polar, desert, and remote flight routes through satellites or HAPS, while air-to-ground direct access can provide low-latency and high-rate links in areas where terrestrial base stations can be deployed. However, aircraft differ significantly from ordinary terrestrial terminals because their flight trajectory, altitude, speed, and route are highly predictable. Therefore, the key issue in aviation NTN access is not only how to execute random access, but how to predict the access window, timing compensation, frequency offset, and target access node before the aircraft enters the coverage area. By using satellite ephemeris, Global Navigation Satellite System information, aircraft trajectory, and velocity parameters, civil aircraft can predict satellite visibility and pre-compute timing advance, scheduling offset, and Doppler compensation before initiating access. This transforms random access from a passive response process into a proactive and predictive access process, thereby improving access certainty and synchronization stability in highly dynamic aviation scenarios. For mobility management, a signaling interaction process for aircraft handover is designed. Based on trajectory prediction and satellite visibility prediction, the network can select a target satellite or gateway with longer residence time and better service capability. Before the aircraft reaches the handover boundary, the source and target network sides can complete context preparation, user-plane path preparation, radio resource reservation, and protocol data unit session update. When the handover condition is triggered, the aircraft performs random access to the target satellite or beam and then switches the user-plane path. This “prediction–preparation–fast handover” mechanism can reduce service interruption and maintain session continuity. For safety-critical traffic, priority and isolation policies should remain consistent and auditable throughout session preparation, handover execution, and path switching. (3)We discuss computing and caching resource management in civil aviation NTN. As onboard computing capability is limited and aviation applications generate increasing computing demands, NTN can provide mobile edge computing and caching services through LEO satellites, HAPS, UAVs, and inter-satellite cooperation. The paper introduces several computing offloading modes, including on-orbit satellite collaborative offloading, network-level integrated offloading, and cloud-edge-terminal hybrid offloading. These mechanisms can support tasks such as aviation monitoring, trajectory analysis, intelligent inference, and in-flight service optimization. In addition, caching mechanisms such as onboard satellite caching, inter-satellite cooperative caching, and named-data-networking-based content caching can improve content delivery efficiency and service continuity. Cache placement should consider content popularity, regional demand prediction, visibility windows, cache prefetching, and cooperative cache sharing among different satellite layers.  Conclusions   NTN can effectively complement traditional civil aviation communication systems by filling coverage gaps in remote and oceanic airspace, enhancing service continuity, and supporting differentiated aviation services. The proposed aviation-oriented NTN architecture integrates multi-orbit satellites, HAPS, UAVs, civil aircraft, and terrestrial infrastructures into a unified framework. The on-demand isolated slicing mechanism can provide differentiated protection for ATC/ATM, AOC, and APC/IFE services. Ephemeris-map-assisted access and predictive mobility management can improve access reliability and reduce handover interruption in high-speed aviation scenarios. Computing offloading and cooperative caching further enhance the ability of NTN to support intelligent and data-intensive aviation applications.  Prospects   Future civil aviation NTN should evolve toward deeper integration of low-altitude networks, space networks, and terrestrial networks. Cross-domain topology visualization, link-state sharing, policy distribution, and programmable logical networks are essential for improving controllability and scalability. In mobility management, integrated cross-domain handover mechanisms should be developed to cope with satellite beam switching, terrestrial cell handover, and air-to-air relay reconstruction. In resource management, communication, navigation, computing, and caching resources should be jointly scheduled and transformed according to aviation service requirements. With continuous advances in NTN architecture, network slicing, predictive access, mobility management, computing offloading, and caching, NTN is expected to provide more efficient, stable, and intelligent communication support for civil aviation and to promote the digital transformation of future air transportation systems.
Research on Secure and Covert Transmission for UAV-assisted Visible Light Communication Systems
WU Mengru, LIN Jiale, LU Weidang, LI Bo, GUO Lei
Available online  , doi: 10.11999/JEIT260239
Abstract:
  Objective  Unmanned Aerial Vehicles (UAVs) can serve as aerial base stations for Visible Light Communication (VLC) because of their mobility and on-demand coverage capabilities. However, air-ground communication links are exposed to open environments, which makes VLC vulnerable to data eavesdropping and malicious detection. To address this issue, this paper proposes a secure and covert transmission strategy for a UAV-assisted VLC system from the perspectives of Physical Layer Security (PLS) and Covert Communication. The proposed strategy jointly optimizes UAV transmit power and hovering altitude to maximize the system secrecy capacity. The optimization is subject to covert communication requirements, illumination requirements, and operational constraints on UAV transmit power and hovering altitude.  Methods  This paper investigates secure and covert communication in a UAV-assisted VLC system. A UAV-assisted VLC system model is first established. In this model, a mobile UAV equipped with a Light-Emitting Diode (LED) is used to establish a VLC link with a legitimate ground user in the presence of an eavesdropper (Eve) and a warden (Willie). An optimization problem is then formulated to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude. To solve this problem, a Two-Layer OPtimization (TLOP) algorithm is proposed. The transformed problem is decomposed into two subproblems: an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude design problem. A closed-form expression for the optimal transmit power is derived for the inner-layer problem. A Particle Swarm Optimization (PSO) algorithm is then developed to solve the outer-layer problem.  Results and Discussions  In the simulations, the proposed optimization scheme is compared with two baseline schemes. First, the convergence of the proposed TLOP algorithm is verified (Fig. 3). The results show that the algorithm converges rapidly within a limited number of iterations. Second, the optimal UAV hovering altitude with respect to the UAV horizontal coordinates is illustrated under the spatial distribution (Fig. 4). The results indicate that the optimal hovering altitude decreases as the UAV approaches the legitimate ground user. The secrecy capacity with respect to the UAV horizontal coordinates is then presented (Fig. 5). The secrecy capacity increases as the UAV approaches the legitimate ground user. This is because the legitimate VLC channel gain increases when the UAV is closer to the user. In contrast, when the UAV approaches Eve and Willie, the security and covertness constraints become stricter. The UAV is then forced to reduce its transmit power or increase its hovering altitude, which decreases the system secrecy capacity. Furthermore, the secrecy capacity of all schemes increases as ϵ increases (Fig. 6). This is because a larger ϵ relaxes the covertness requirement. The UAV can therefore adjust its hovering altitude and transmit power more flexibly to increase the system secrecy capacity. In addition, the secrecy capacity decreases as the number of symbols increases (Fig. 7). This occurs because more symbols provide Willie with more signal samples for detection, thereby improving Willie’s detection capability. Finally, the secrecy capacity of all schemes decreases as the uncertainty-region radius of illegal nodes increases (Fig. 8). This trend occurs because greater location uncertainty forces the UAV to address potential threats over a wider area. The UAV must therefore adopt a more conservative strategy under worst-case eavesdropping and detection conditions. Overall, the simulation results confirm that the proposed scheme improves the secrecy capacity of the UAV-assisted VLC system.  Conclusions  This paper investigates secure and covert communication in a UAV-assisted VLC system. The objective is to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude under covert communication, illumination, transmit power, and hovering altitude constraints. Because the formulated problem is highly non-convex, a PSO-based TLOP algorithm is designed to solve it. The proposed algorithm decomposes the problem into an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude optimization problem. Simulation results show that the proposed algorithm converges rapidly and improves the system secrecy capacity compared with the baseline schemes.
Energy-Efficient Trajectory Planning and Resource Optimization for UAV Relay Communications over Hybrid RF/FSO Links
LI Baolong, PAN Wenwei, JIANG Hao, FENG Simeng, WU Qihui
Available online  , doi: 10.11999/JEIT260139
Abstract:
  Objective  In low-altitude communication networks, hybrid Radio Frequency/Free-Space Optical (RF/FSO) Unmanned Aerial Vehicle (UAV) relaying can ease RF spectrum congestion and improve uplink data aggregation. However, in obstacle-rich urban environments, FSO backhaul links are vulnerable to blockage and intermittent outages. This creates a severe mismatch between the RF access-link rate and the FSO backhaul-link rate. UAV trajectory planning is also constrained by obstacle avoidance and flight dynamics. To address these coupled issues, this paper investigates an energy-efficiency maximization problem. Multiuser Non-Orthogonal Multiple Access (NOMA)-based RF access and the Three-Dimensional (3D) obstacle-avoiding UAV trajectory are jointly optimized, and buffer-assisted RF/FSO rate decoupling is incorporated.  Methods  A time-slotted UAV relaying model is considered, in which multiple ground users upload data to the UAV through an RF link using NOMA. The UAV decodes superposed signals by Successive Interference Cancellation (SIC), and the decoding order in each slot is determined according to the received-power ranking. The successfully received data are then forwarded to a Base Station (BS) through an FSO backhaul link. Urban blockage is modeled using 3D geometric obstacles. A visibility test is used to determine whether each relevant link is in Line-Of-Sight (LOS) or Non-Line-Of-Sight (NLOS), which captures the spatially correlated and time-varying RF access-link rate and intermittent FSO backhaul capacity. To suppress blockage-induced rate mismatch between the RF access link and the FSO backhaul link, an onboard finite-capacity buffer is deployed at the UAV. In each slot, the forwardable data amount is jointly limited by the instantaneous FSO backhaul capacity and the data available in the buffer, and buffer-capacity constraints are imposed to prevent overflow. System energy efficiency is defined as the ratio of cumulative data successfully delivered to the BS over the mission horizon to UAV propulsion energy consumption. Propulsion power is modeled as a function of UAV velocity and acceleration to reflect the effect of flight dynamics. Under 3D flight-region boundaries, prescribed start and end locations, discrete-time kinematic equations, maximum velocity and acceleration limits, and obstacle collision-avoidance constraints, a non-convex optimization problem is formulated. The decision variables are cross-slot multiuser transmit powers and the 3D UAV trajectory. An alternating optimization framework is then developed. For a fixed trajectory, propulsion energy is fixed, so maximizing energy efficiency is equivalent to increasing end-to-end successfully forwarded data. This yields a power-optimization subproblem. Because of NOMA coupling and logarithmic rate expressions, this subproblem remains non-convex and is solved by Successive Convex Approximation (SCA). For fixed transmit powers, Particle Swarm Optimization (PSO) is used to search candidate 3D trajectories in continuous space. To ensure feasibility under strict dynamics and safety constraints, Quadratic Programming (QP) projection is used to enforce velocity and acceleration constraints. Collision checks are performed for trajectory waypoints and inter-slot line segments to ensure obstacle-free flight. These two optimization procedures are performed alternately. The resulting joint design satisfies flight-dynamics feasibility and collision-avoidance requirements and improves energy efficiency.  Results and Discussion   Simulations are conducted in an urban airspace with multiple users, a BS, and dense 3D obstacles. Blockage causes frequent LOS/NLOS switching as the UAV moves. Fig. 2 and 3 compare the 3D trajectory and its planar projection, respectively. Compared with the initial trajectory, the optimized trajectory shows clear detours and necessary altitude adjustments. It achieves collision-free flight while satisfying velocity and acceleration constraints, thereby verifying the feasibility and safety of the proposed trajectory planning method. Fig. 4 shows the convergence of energy efficiency under different user transmit-power budgets. The proposed alternating optimization generally stabilizes within a small number of outer iterations. The converged energy efficiency increases with the power budget, indicating synergy between power control and trajectory adaptation. Fig. 5 shows buffer evolution over time. The buffer gradually accumulates data when the backhaul is blocked or experiences strong fading. It is quickly drained when the UAV enters regions with LOS backhaul and improved FSO capacity. To quantify buffering gain, Fig. 6 compares system energy efficiency between the proposed buffering mechanism and the no-buffer scheme. The proposed mechanism enables store-and-forward temporal smoothing during backhaul interruptions and improves system energy efficiency. Fig. 7 shows energy-efficiency convergence under different buffer capacities. As buffer capacity increases, the converged energy-efficiency level improves. A larger buffer enhances the UAV’s ability to temporarily store incoming data and reduces data accumulation and transmission blockage when RF access-link and FSO backhaul-link rates are mismatched or the backhaul link is constrained. Figure 8 compares four benchmark schemes, namely a non-optimized baseline, a power-optimization scheme, a trajectory-optimization scheme, and the proposed joint power-and-trajectory optimization scheme. The coordinated design of power allocation and obstacle-avoiding trajectory improves end-to-end energy efficiency. Trajectory optimization also plays a more dominant role under blockage-limited conditions.  Conclusion  This paper investigates a hybrid RF/FSO UAV relaying scheme with NOMA and an onboard buffering mechanism for low-altitude urban communication. Given dense obstacles, frequent blockage, FSO-link susceptibility, and strict flight-dynamics constraints, an energy-efficiency maximization problem is formulated for the joint optimization of multiuser NOMA power allocation and UAV trajectory. An SCA-based power-allocation method and an obstacle-avoiding trajectory design that combines PSO with QP projection are developed. The obtained trajectory satisfies flight-dynamics feasibility and collision-avoidance requirements and improves throughput per unit propulsion energy. Simulation results show that the planned trajectory can avoid obstacles, and that the onboard buffer provides an effective cushion between RF access and FSO backhaul to mitigate rate mismatch. The proposed method consistently outperforms benchmark schemes in energy efficiency. Trajectory optimization is also shown to be generally more effective than power allocation in improving overall system performance.
Joint Optimization Method for Pairwise Constrained Projection Clustering Integrating a Two-row Simultaneous Update Strategy
ZHU Jianyong, CHEN Kun, YANG Hui, NIE Feiping
Available online  , doi: 10.11999/JEIT260111
Abstract:
  Objective  As data structures become increasingly complex, conventional unsupervised clustering methods often fail to achieve satisfactory performance. Semi-supervised clustering has therefore attracted growing attention because it uses limited prior information to improve clustering quality. However, existing methods have two major limitations. First, traditional constrained projection clustering algorithms usually use a two-step independent strategy, in which the projection matrix is learned before k-means clustering is performed. This separation allows projection errors to be propagated directly to the clustering stage, causing accumulated learning errors. In addition, applying pairwise constraints only during projection deviates from the goal of using prior information to guide clustering. Second, many existing methods, including spectral clustering-based approaches, handle pairwise constraints implicitly, for example through eigen-decomposition of a modified similarity matrix. Such implicit processing may not strictly satisfy the constraints, especially Cannot-Link (CL) constraints, which are non-transitive, resulting in high constraint violation rates. To address these issues, this paper proposes a joint optimization method for pairwise constrained Projection Clustering Integrating a Two-row simultaneous Update Strategy (PCITUS). The objective is to unify dimensionality reduction and clustering within a single framework to reduce information loss, while designing an explicit optimization strategy that lowers constraint violations and improves computational efficiency.  Methods  The proposed PCITUS model integrates constrained projection and clustering into a unified objective function for collaborative optimization, with pairwise constraints optimized directly. First, the algorithm uses the transitive property of Must-Link (ML) constraints. Samples belonging to the same ML connected component are merged into a single hyper-point in the feature space. This preprocessing step ensures that all ML constraints are naturally satisfied. A trade-off parameter is then introduced to incorporate projection learning into the clustering framework as a regularization term, allowing both components to be jointly optimized under one objective. Prior information is further embedded into the clustering process by transforming pairwise constraints into row-wise constraints on the indicator matrix. An improved coordinate descent method is then used to optimize the discrete indicator matrix directly, which improves computational efficiency and produces better clustering results. A key feature of PCITUS is the two-row simultaneous update strategy for CL constraints. PCITUS explicitly checks CL conflicts by simultaneously evaluating objective function values obtained by moving conflicting rows to suboptimal classes and then selects the case with the higher value.  Results and Discussions  Extensive experiments are conducted on eight benchmark datasets and compared with nine state-of-the-art semi-supervised clustering algorithms. Quantitative results based on ACCuracy (ACC) and Normalized Mutual Information (NMI) demonstrate the superiority of PCITUS (Table 4 and Table 5). PCITUS achieves the best performance on most datasets. In particular, on the Mushroom dataset, NMI is improved by 7.29% compared with the second-best algorithm. The comparison with CNP, a two-step projection method, confirms that the unified framework effectively reduces error propagation and information loss. This effect is also supported by the mutual reinforcement between projection and clustering: a better projection space produces a clearer clustering structure, while a more reasonable clustering structure guides the formation of a more discriminative projection space. The effectiveness of explicit constraint handling is further illustrated (Fig. 1). PCITUS produces no ML constraint violations because of the hyper-point merging strategy. For CL constraints, the two-row simultaneous update strategy enables PCITUS to maintain an extremely low violation rate, such as 0.57% on Mushroom and 0.41% on Satimage, greatly outperforming methods that handle constraints implicitly. Additionally, the parameter sensitivity analysis (Fig. 2) shows that PCITUS remains stable across a wide range of trade-off parameter values. The noise sensitivity experiments (Fig. 3a and Fig. 3b) confirm its robustness. The convergence curves (Fig. 3c and Fig. 3d) and runtime comparisons (Table 7) further verify its computational efficiency, showing rapid convergence and a stable objective function value within approximately 10 iterations in most cases.  Conclusions  This paper presents PCITUS, a semi-supervised clustering framework that jointly optimizes pairwise constrained projection and clustering structures. The method addresses the difficulty of optimizing CL constraints and overcomes the limitations of traditional constrained projection clustering frameworks based on a two-step separation scheme. By integrating the projection objective into the clustering framework as a regularizer, the proposed method enables subspace learning and data partitioning to reinforce each other and jointly approach the global optimum. Pairwise constraints are used throughout the learning process, allowing prior knowledge to guide optimization more fully. The coordinate descent method with the two-row simultaneous update strategy directly and accurately allocates samples under CL constraints, significantly reducing constraint violations. Experimental results show that PCITUS outperforms existing algorithms in clustering performance.
Robust Optimization of Low-altitude Communication and Computation Resources in Uncertain Environments
GONG Yucheng, LI Bin, WANG Xinyi, FEI Zesong
Available online  , doi: 10.11999/JEIT260090
Abstract:
  Objective  Low-altitude edge computing networks provide flexible computing services and extended coverage for user equipment. However, quality of service is often degraded by uncertainty in task data size and by Unmanned Aerial Vehicle (UAV) position jitter caused by environmental disturbances. Existing robust methods commonly rely on deterministic uncertainty sets, which tend to be conservative and cannot accurately describe the stochastic distribution of task demands. To address these challenges, a robust energy minimization framework is proposed for multi-UAV-assisted Mobile Edge Computing (MEC) networks. The objective is to minimize the weighted sum of system energy consumption. This is achieved by developing a joint optimization model that coordinates UAV flight trajectories, task splitting decisions, and computation and communication resource allocation. The model explicitly accounts for the dual uncertainties of task data size and UAV trajectory.  Methods  To handle the nonconvexity and strong coupling among optimization variables, the problem is first modeled as a Markov Decision Process (MDP). A comprehensive state space is defined to characterize real-time system dynamics, and a continuous action space is designed for trajectory control and resource management. A Distributionally Robust Optimization Soft Actor-Critic (DRO-SAC) algorithm is then developed to solve the MDP. In this framework, an ambiguity set based on the L1-norm distance is constructed to characterize the distributional uncertainty of the task demand distribution. A maximum-entropy reinforcement learning mechanism is used to learn an optimal policy under the worst-case distribution within the ambiguity set. In this way, UAV trajectories, task splitting, and computation and communication resource allocation are jointly optimized to improve system robustness under dynamic environmental fluctuations.  Results and Discussions  The performance of the proposed DRO-SAC algorithm is evaluated through simulations. DRO-SAC achieves faster convergence and higher rewards than Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO) algorithms (Fig. 3). For energy consumption, the proposed method consistently achieves higher efficiency under different user densities (Fig. 4). The robustness of the system against position errors is also verified, with energy fluctuations kept at a low level (Fig. 5). Dynamic trajectory adjustment further confirms that the proposed method can provide effective user coverage while reducing system energy consumption (Fig. 6).  Conclusions  A DRO-SAC-based joint optimization framework is proposed to address uncertainty in task data size and UAV position jitter in multi-UAV-assisted MEC networks. By constructing an ambiguity set for the task demand distribution and optimizing the worst-case expected objective, the proposed method mitigates the limitations of traditional deterministic models in dynamic environments. Weighted system energy consumption is minimized while latency and safety constraints are satisfied. Simulation results demonstrate that the proposed scheme achieves stable convergence and high energy efficiency, even when communication and computation resources are limited and environmental parameters fluctuate strongly.
A Lightweight True Random Number Generator Based on Chain-Coupled Oscillation Rings
ZHANG Yuan, YING Haixuan, GAO Kai, YE Jin, WANG Shuang, ZHANG Jiliang
Available online  , doi: 10.11999/JEIT260377
Abstract:
  Objective  With the rapid growth of the Internet of Things, 5G/6G, and satellite Internet, resource-constrained devices increasingly require high-quality random numbers for key generation, authentication, masking, and other security functions. Although pseudo-random number generators are efficient, their outputs may be predictable once the seed or internal state is compromised. True random number generators (TRNGs) offer a hardware root of trust by extracting entropy from physical randomness, but many existing designs rely on multiple entropy sources or complex post-processing, leading to increased area and power consumption. To address this issue, this paper proposes a lightweight TRNG based on chain-coupled oscillation rings for high-quality randomness with very low FPGA overhead.  Methods  Starting from the state evolution of a Galois oscillation ring (GARO), this work demonstrates that ideal matched-delay conditions can result in periodic and predictable oscillation. However, in practical circuits, delay mismatch, jitter, and process variation disturb the ideal evolution and can be exploited as entropy sources. On this basis, a compact delay-feedback XOR ring is proposed to enhance state uncertainty, introduce feedback competition, and improve randomness through inter-stage delay differences. In addition, a second-order oscillation ring is incorporated to eliminate the all-zero stop state and provide continuous excitation. Multiple rings are then chain-coupled, enabling adjacent rings to mutually interfere with one another and thereby generate stronger irregular oscillations. The proposed design is modeled in MATLAB and implemented on a Xilinx Artix-7 FPGA. Finally, we evaluate its performance by NIST SP 800-22, NIST SP 800-90B, bias, autocorrelation, and voltage-temperature robustness tests.  Results and Discussions  Simulation confirms that the proposed structure avoids stable periodic locking and produces sustained irregular oscillation. Experimental results show that the TRNG passes all NIST SP 800-22 tests and achieves an average minimum entropy of 0.9936 in NIST SP 800-90B test, outperforming conventional RO and GARO-based TRNGs under similar conditions. The measured bias is only 0.0228%, and the autocorrelation remains well below the threshold, indicating excellent statistical independence. The design also maintains high entropy over temperatures from 0 °C to 80 °C and supply voltages from 0.9 V to 1.1 V. Implemented on Artix-7, our proposed TRNG achieves 200 Mbps throughput using only 11 LUTs and 4 DFFs, with 0.108 W power consumption.  Conclusions  This paper presents a lightweight chain-coupled oscillation-ring TRNG that exploits delay mismatch, phase disturbance, and feedback competition to generate high-quality physical randomness. The theoretical analysis clarifies how practical nonidealities transform ideal periodic oscillation into irregular oscillation, providing a design basis for compact oscillator-based entropy sources. By combining delay-feedback XOR rings with chain-coupled mutual disturbance and continuous excitation, the proposed design enhances entropy while avoiding excessive hardware overhead and complex post-processing. FPGA implementation and statistical evaluations verify high entropy, low bias, and high randomness under voltage and temperature variations. Therefore, the proposed TRNG achieves high randomness quality and high throughput while effectively reducing hardware overhead, making it suitable for resource-constrained security applications such as IoT terminals, lightweight cryptographic modules, and embedded authentication systems.
From Touch to Semantics: A Cross-Modal Framework for Zero-Shot Spiking Tactile Object Recognition
CHI Wei, XU Jin
Available online  , doi: 10.11999/JEIT260158
Abstract:
  Objective  Tactile perception enables robots to understand object properties and perform dexterous interactions. However, tactile data are costly to collect and difficult to scale, which limits conventional supervised learning in open-world scenarios. Zero-Shot Learning (ZSL) provides a promising solution by transferring knowledge from seen to unseen categories through semantic representations. Existing tactile ZSL methods either rely on auxiliary visual information or use manually designed attributes, which are often subjective and limited in generalization. Event-based spiking tactile signals are sparse and asynchronous, with rich spatiotemporal dynamics. These properties make semantic modeling more challenging. Systematic studies on zero-shot recognition for such data remain limited. To address these issues, this paper proposes a zero-shot object recognition framework for spiking tactile perception. The framework aims to bridge low-level tactile dynamics and high-level semantics in a scalable manner.  Methods  The proposed framework consists of three components (Fig. 1): spiking tactile feature extraction, semantic prototype construction, and cross-modal tactile-semantic alignment. First, a biomimetic Spiking Graph Neural Network (SGNN) is used to model raw event-based spiking tactile signals. By integrating Leaky Integrate-and-Fire (LIF) neurons with graph-based message passing, the SGNN captures temporal firing dynamics and spatial relationships among tactile sensing units. It then generates discriminative and biologically interpretable high-level tactile embeddings. Second, instead of using manually annotated attributes, a Large Language Model (LLM) is used to generate structured, fine-grained, and extensible tactile attribute descriptions for each object category. These textual descriptions are encoded as continuous semantic vectors to form class-level semantic prototypes with consistent dimensionality across categories. This strategy supports flexible semantic expansion and avoids labor-intensive attribute engineering. Third, a bidirectional tactile-semantic alignment mechanism is designed to improve generalization to unseen categories. A forward mapping projects tactile embeddings into the semantic space for classification, whereas a reverse mapping reconstructs tactile features from semantic representations. A cycle-consistency constraint is imposed between the two mappings to preserve structural coherence and semantic stability across modalities. The overall framework is trained only on seen categories. During zero-shot inference, tactile embeddings of unseen samples are matched with their corresponding semantic prototypes in the shared embedding space.  Results and Discussions  The proposed method is evaluated on the Ev-Object event-based tactile dataset under a strict zero-shot setting, with disjoint seen and unseen category sets. Performance is assessed using Mean Class Accuracy (MCA), Top-k accuracy, and the Semantic Alignment Score (SAS). The proposed framework consistently outperforms representative tactile ZSL baselines across all metrics. It achieves an MCA of 73.48%, a Top-1 accuracy of 62.68%, and a Top-2 accuracy of 88.75%. Ablation studies show that removing the LLM semantic module, bidirectional mapping, or cycle-consistency constraint reduces recognition performance and semantic alignment quality. Removing the LLM semantic module causes a substantial decrease in MCA, which confirms the role of structured LLM-generated tactile semantics in knowledge transfer. Removing the bidirectional mapping or the cycle-consistency constraint also reduces performance, indicating that both components help maintain stable cross-modal alignment. The t-SNE visualization further shows that cycle-consistent alignment yields more compact intra-class clusters and clearer inter-class separation for unseen categories. Semantic prototypes are also better located near the centers of tactile feature clusters. These results indicate that combining biologically inspired spiking models with LLM-generated tactile semantics provides an effective solution for open-world tactile perception.  Conclusions  This paper presents a zero-shot object recognition framework for spiking tactile perception by integrating SGNN-based tactile representation with semantic prototypes. The proposed method addresses key limitations of existing tactile ZSL approaches by avoiding visual data and manual attribute design while effectively modeling the spatiotemporal dynamics of event-based spiking tactile signals. Experimental results under strict zero-shot settings confirm the effectiveness and robustness of the proposed framework. This work provides a strong baseline for zero-shot spiking tactile recognition and offers a principled path toward open-world tactile cognition in robotic systems. Future work will explore generalized zero-shot tactile perception, multimodal extensions, and real-world robotic deployment under noisy and dynamic sensing conditions.
A Noise Reduction Strategy via Coprime-Spacing Subarrays for Biodiversity Acoustic Indices
CHEN Lei, XU Zhiyong, ZHAO Zhao
Available online  , doi: 10.11999/JEIT260237
Abstract:
  Objective  As a popular tool for rapid biodiversity assessment, acoustic indices have attracted increasing attention in the field of soundscape ecology in recent years. Nevertheless, most commonly used acoustic indices are susceptible to background noise. Traditional single-channel noise reduction strategies, including spectral subtraction, high-pass filtering, and threshold detection, have been widely adopted as preprocessing approaches to optimize the calculation of acoustic indices. However, when dealing with anthropogenic interference that overlaps with biotic signals in both time and frequency domains, the denoising capability of single-channel methods degrades severely. Although spatio-temporal adaptive whitening filtering based on microphone arrays provides a feasible approach for suppressing directional interference, it suffers from a non-uniform two-dimensional spatio-temporal amplitude response and the self-cancellation of target signal in the unconstrained interference cancellation. These disadvantages lead to distortion in the time-frequency distribution of target signals, causing acoustic index calculations to deviate from the ground truth. Therefore, this study aims to propose a noise reduction strategy via coprime-spacing subarrays for biodiversity acoustic indices. This method effectively suppresses directional interference while maximally preserving the time-frequency distribution structure of biotic signals.  Methods  The noise reduction strategy based on microphone array spatio-temporal adaptive whitening filtering is proposed, incorporating the Frequency-dependent Acoustic Diversity Index (FADI), which is insensitive to fluctuations in the array's two-dimensional spatio-temporal amplitude response. A noise-robust acoustic index method, termed Adaptive Interference Cancellation–Frequency-dependent Acoustic Diversity Index (AIC-FADI), is subsequently developed. Specifically, a non-uniform linear array is first constructed using three microphones to form two dual-element subarrays with coprime spacing. This design fully exploits the high spatial resolution of wide-spacing arrays to narrow the null width in the direction of interference. Meanwhile, it avoids the physical implementation difficulties and mutual coupling effects associated with small-spacing array designs caused by the ultra-wideband characteristics of target signals. The spatio-temporal adaptive whitening filtering is then performed on each coprime-spacing subarray separately, adaptively forming two-dimensional nulls within the interference support region, thereby suppressing directional anthropogenic interference in analytical data before index calculation. Next, a frequency-dependent threshold scheme is utilized to obtain the binary spectrogram for each coprime-spacing subarray output, abating the influence from gain differences along the frequency axis for a certain direction. Afterwards, by leveraging the high spatial resolution of wide-spacing arrays and the interleaved characteristics of spatial aliasing null positions between the spatio-temporal frequency responses of the two subarrays with coprime spacing, a pointwise maximum fusion is applied to the above two binary spectrograms. This process reconstructs the binary time-frequency distribution structure of target signals outside the interference support region, leading to a single binary spectrogram where biological sound components are preserved to a great extent and anthropogenic interference is considerably suppressed. Ultimately, from this single binary spectrogram, the proportions of non-zero time-frequency bins within each frequency band are calculated and forwarded to the entropy function, resulting in the final AIC-FADI result.  Results and Discussions  The simulation result indicates that the proposed AIC-FADI maintains numerical robustness across an SINR range down to –15 dB (the yellow line in Fig. 5), substantially outperforming the classical ADI version based on single-channel noise reduction algorithm (FADI) and other ADI versions based on single-array interference suppression processing mentioned in this paper (AIC-FADI-s, AIC-FADI1, and AIC-FADI2). The real-world experiment confirms that the proposed spatio-temporal adaptive whitening filtering effectively suppresses wideband interference signals in complex scenarios, thereby improving the SINR of the analyzed recording. This enables some weaker biotic signals to exceed their corresponding frequency-dependent adaptive thresholds, greatly reducing missed detection of the target signal. In addition, by performing pointwise maximum fusion of the binary spectrograms from the two coprime-spacing subarray outputs, AIC-FADI further alleviates the extent of target signal missed detection (Fig. 8). Nevertheless, the real-world experiments also reveal that the interference suppression performance of AIC-FADI degrades for highly time-varying interference components.  Conclusions  This paper addresses the challenge of calculating acoustic indices reliably in complex soundscapes where directional anthropogenic interference overlaps with biotic signals in both time and frequency domains. A noise reduction strategy using coprime-spacing subarrays is proposed, and a new noise-robust acoustic index (AIC-FADI) is then developed. The method is evaluated through simulations and real-world recordings, and the results show that: (1) By applying spatio-temporal adaptive whitening filtering on each coprime-spacing subarray followed by pointwise maximum fusion, the proposed method achieves both wideband interference suppression capability and target information fidelity in complex soundscapes containing strong interference. (2) As a result, the proposed AIC-FADI maintains numerical robustness down to –15 dB SINR, substantially outperforming the classical FADI algorithm and other ADI versions based on single-array interference suppression methods. (3) The proposed method provides a feasible technical solution for extending the practical application scenarios and spatio-temporal coverage of biodiversity acoustic indices in human-dominated areas. However, this study only considers directional interference that is relatively stable or slowly time-varying. Hence, the interference suppression performance degrades for highly time-varying or uncorrelated noise components. These challenges should be addressed in future work through more advanced signal processing techniques to further improve the robustness of acoustic indices in highly complex acoustic environments.
A Survey of Quantum Covert Communication Integration Schemes and Application Scenarios
SUN Yiheng, XU Yongjun, ZHANG Haibo, HUANG Zishan
Available online  , doi: 10.11999/JEIT260282
Abstract:
  Significance   With the growing demand for network communication security, research and development in covert communication and quantum communication have continued to evolve. However, current covert communication suffers from inherent security vulnerabilities; the transmission reliability of quantum communication has been limited by information eavesdropping and harmful interference. Therefore, quantum covert communication has become a research hotspot, integrating the advantages of both covert and quantum communication while addressing their respective security limitations. To this end, this paper provides a comprehensive survey of quantum covert communication integration schemes and application scenarios, including the principles of covert communication and typical enabling techniques; protocols for quantum communication and important quantum techniques; and three types of quantum covert communication integration schemes summarized by different application scenarios. This paper contributes to the design of advanced secure communication networks while offering guidance for the development of future quantum covert communication systems.  Progress   This paper presents a comprehensive survey of recent advances in quantum covert communication integration schemes and application scenarios, with an in-depth discussion of the principles of covert communication and key enabling techniques, such as Fluid Antenna (FA), Reconfigurable Intelligent Surface (RIS), and Unmanned Aerial Vehicle (UAV). FA actively reshapes wireless channel characteristics, particularly the spatial correlation of multipath components, by dynamically adjusting the transmitter physical configuration, thereby reducing information leakage. In Non-Line-of-Sight (NLoS) scenarios, RIS can dynamically alter the direction of reflected transmission of the incident signal, not only enhancing the Channel State Information (CSI) quality of the covert signal but also reducing signal leakage. In flexible or temporary communication networks, UAVs can increase CSI uncertainty, preventing unauthorized users from establishing a stable monitoring model and thereby complicating eavesdropping. Then, key protocols and significant techniques of quantum communication are introduced, including BB84, B92, and E91 for Quantum Key Distribution (QKD), and BF02, Two-Step for Quantum Secure Direct Communication (QSDC). Additionally, the quantum repeaters and Quantum Random Number Generator (QRNG) are reviewed. Based on different application scenarios, quantum covert communication integration schemes can be categorized into enabling, covert, and symbiotic integration schemes, depending on the integration mechanisms. To be specific, the enabling integration scheme leverages the unconditional security of quantum communication to address the security vulnerabilities in covert communication, the covert integration scheme utilizes enabling techniques in covert communication to reduce the detection probability of quantum communication, and the symbiotic integration scheme combines both advantages of covert communication and quantum communication to achieve mutual empowerment and deep symbiosis. Finally, critical challenges are highlighted, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles in large-scale applications. Promising directions for future research are also identified, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development.  Prospects   Despite remarkable progress in preliminary applications and specific scenarios, research on quantum covert communication remains in its infancy. As quantum covert communication scenarios become increasingly diverse and complex, future studies should prioritize challenges that restrict further development and large-scale application of quantum covert communication. The stringent hardware precision requirements are the primary challenge, limiting reliable transmission distance and stability. Low resource allocation efficiency is another challenge, as the quantum covert communication system that generates quantum entanglement over lossy channels remains subject to the Square Root Law (SRL) constraints, while signal transmission exhibits burstiness and dynamics. Additionally, high deployment costs and the lack of standardization present significant hurdles. To address the challenges mentioned, future directions should include R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development to facilitate the development of high-performance, large-scale, and multi-scenario quantum covert communication.  Conclusions  This paper provides a comprehensive survey of quantum covert communication with particular emphasis on integration schemes and application scenarios. The fundamentals and typical enabling techniques of covert communication are first reviewed, highlighting its Low Probability of Detection (LPD) secure paradigm and unique channel characteristics. The typical protocols and important techniques of quantum communication are then examined, including QKD, QSDC, quantum repeaters, and QRNG. Three types of quantum covert communication integration schemes have been further classified by different integration mechanisms and corresponding application scenarios. Finally, several existing challenges are identified, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles to large-scale applications. Relevant research directions are also outlined, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development. These directions are expected to serve as a valuable reference for advancing and standardizing quantum covert communication in future secure networks.
Full-Space Covert Integrated Sensing and Communications Assisted by Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface
XIE Wenwu, ZHANG Qinke, YANG Liang, WANG Ji, YU Chao, LIU Xinzhong, CUI Yaru
Available online  , doi: 10.11999/JEIT260145
Abstract:
  Objective  The evolution of Sixth Generation (6G) mobile communications toward higher frequencies and larger antenna arrays has made Integrated Sensing And Communication (ISAC) a key enabling technology. However, ISAC systems still face limited communication covertness and resource competition between sensing and communication. Covert communication and Reconfigurable Intelligent Surface (RIS) techniques provide promising solutions. However, most existing studies use reflective RISs with half-space coverage and assume far-field propagation. These assumptions limit deployment flexibility and fail to capture near-field spherical-wave characteristics. To address these issues, this paper proposes a near-field full-space ISAC framework assisted by an Extremely Large-Scale Simultaneously Transmitting And Reflecting Reconfigurable Intelligent Surface (XL-STAR-RIS). The objective is to jointly optimize active transmit beamforming and passive XL-STAR-RIS coefficient design to improve the covert communication rate while satisfying sensing performance and covertness requirements.  Methods  The detection capability of warden Willie is first analyzed, and a closed-form lower-bound expression for the minimum Detection Error Probability (DEP) is derived. A non-convex optimization problem is then formulated to maximize the covert communication rate under sensing Signal-to-Noise Ratio (SNR), covertness, and total transmit power constraints. Direct solution is difficult because the active transmit beamforming vectors and passive XL-STAR-RIS coefficients are strongly coupled. An Alternating Optimization (AO) framework is therefore adopted to decompose the original problem into two tractable subproblems. The active transmit beamforming subproblem is solved using SemiDefinite Relaxation (SDR) combined with a penalty-based successive convex approximation method. The passive XL-STAR-RIS coefficient design subproblem is solved using the Dinkelbach algorithm and a rank-one penalty method. The two subproblems are solved alternately until convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed framework. The algorithm converges within approximately 10 iterations and achieves a covert communication rate of about 11.5 bit/(s·Hz). This rate is higher than those of the passive-RIS scheme (9.8 bit/(s·Hz)) and the non-RIS scheme (8.0 bit/(s·Hz)). The performance gain becomes more evident as the transmit power increases, which indicates strong power adaptability. The proposed framework also maintains robust performance under strict operational constraints. When the sensing SNR threshold increases, it achieves a higher covert communication rate than the benchmark schemes. Under a stricter covertness requirement, it also preserves a higher communication rate. These results show that joint active transmit beamforming and passive XL-STAR-RIS coefficient design can effectively balance communication, sensing, and covertness in near-field ISAC systems.  Conclusions  This paper presents an XL-STAR-RIS-assisted covert communication framework for near-field ISAC systems. By jointly designing active transmit beamforming and passive XL-STAR-RIS coefficients through an efficient AO algorithm, the proposed framework balances communication rate, sensing performance, and communication covertness. Simulation results confirm its advantages over conventional passive-RIS and non-RIS schemes, especially under strict sensing and covertness constraints. The results also indicate the potential of XL-STAR-RIS for secure full-space 6G applications. Future work will consider imperfect Channel State Information (CSI), dynamic propagation environments, and multi-RIS collaboration to improve practical robustness.
Millimeter-Wave Air-to-Ground Channel Prediction Assisted by Visual Information of the Propagation Environment
CHENG Yuanxun, HU Qingsong, ZHANG Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260274
Abstract:
  Objective  Accurate prediction of air-to-ground (A2G) channel states is essential for adaptive transmission and resource optimization in unmanned aerial vehicle (UAV) communications. In urban millimeter-wave scenarios, however, A2G links are highly sensitive to blockage, reflection, scattering, and the rapidly changing geometric relationship among the transmitter, the receiver, and surrounding buildings. As a result, the channel exhibits strong spatial and temporal nonstationarity, and conventional pilot- or feedback-based acquisition methods may become ineffective because the obtained channel state information is easily outdated. Recent data-driven approaches have shown potential, but many of them rely heavily on historical channel observations or directly use raw images as network inputs, which may introduce redundant visual information and weaken physical interpretability. To address these limitations, this paper proposes a vision-assisted millimeter-wave A2G channel prediction method that extracts low-dimensional geometric features from the propagation environment instead of using raw visual data directly. The objective is to preserve the key structural information governing channel evolution while reducing irrelevant redundancy, thereby improving the prediction of channel.  Methods  A communication-and-sensing integrated dataset with strict spatial and temporal alignment is established for millimeter-wave UAV A2G channel prediction. On the sensing side, a high-fidelity three-dimensional urban scenario containing 23 buildings, roads, and intersections is constructed in Unreal Engine 4.27, where synchronized RGB and depth images are collected through AirSim using a multirotor UAV equipped with RGB and depth cameras. The UAV flies along 10 preset trajectories at a height of 55 m with a spatial sampling interval of 1 m, yielding 2160 valid visual samples (Fig. 1, Fig. 2). On the communication side, the same scene is reconstructed in Wireless InSite, and the transmitter-receiver positions are synchronously updated along the same trajectories to ensure frame-level alignment between visual and channel data (Fig. 3). To obtain compact and physically meaningful environmental representations, a cross-modal spatial feature extraction scheme is developed. Buildings are first detected from RGB images using YOLO-V8 (Fig. 4), and the detected regions are then registered with depth images to reconstruct three-dimensional point clouds. After Euclidean clustering and axis-aligned bounding-box fitting, key geometric attributes, including planar position, height, and volume, are extracted. These features are combined with the transmitter-receiver distance to form the spatial feature vector of each frame, and their relevance to path loss, received power, and RMS delay spread is evaluated through cosine-similarity-based correlation analysis (Fig. 6). Based on the extracted features, a hybrid Transformer-MLP network is designed for channel prediction (Fig. 5). Building features are first projected into a latent space, and a stacked Transformer encoder is employed to capture global interactions among buildings through masked multi-head self-attention. Masked average pooling is then used to aggregate building-level representations into a scene-level environmental descriptor, which is concatenated with the link distance feature and fed into a multilayer perceptron regressor to predict the three target channel parameters.  Results and Discussions  The results confirm the effectiveness of the proposed spatial feature representation. Correlation analysis shows that the extracted geometric features are consistently related to path loss, received power, and RMS delay spread under different aggregation strategies (Fig. 6), indicating that compact building descriptors can effectively characterize the propagation environment. Among them, building height exhibits the strongest correlation with all three channel parameters, highlighting its important role in blockage, attenuation, and multipath propagation in urban millimeter-wave A2G channels. In prediction experiments, the proposed method accurately tracks the variation trends of all three targets. It remains effective in deep-fading and sharp-fluctuation regions for path loss prediction (Fig. 7), achieves high consistency with the ground truth for RMS delay spread (Fig. 8), and follows rapid local fluctuations of received power with good fidelity (Fig. 9). In contrast, the benchmark model only captures the general trend and shows larger deviations in peaks, valleys, and abrupt-changing intervals. Residual analysis further demonstrates the superiority of the proposed method. Its errors are more concentrated around zero and fluctuate within narrower ranges than those of the benchmark model across all three tasks (Fig. 10). Quantitatively, both the mean absolute error and the root mean squared error are reduced (Fig. 11). In addition, the model maintains acceptable complexity, with about 5.5 M parameters and a single-frame inference delay of about 3.4 ms, indicating good potential for real-time deployment.  Conclusions  A vision-assisted millimeter-wave A2G channel prediction method for UAV communications is proposed. By constructing a strictly aligned communication-and-sensing dataset and extracting low-dimensional spatial features with clear physical meaning, the method establishes an effective mapping from environmental geometry to channel parameters. The proposed Transformer-MLP framework achieves accurate prediction of path loss, received power, and RMS delay spread, while offering better interpretability, robustness, and efficiency than the benchmark model.
Transfer Learning Aided CNN for Efficient Data Detection in ReRAM with Sneak-Path Interference
DAI Bin, WU Anni
Available online  , doi: 10.11999/JEIT260354
Abstract:
  Objective  Sneak path interference (SPI) in resistive random-access memory (ReRAM) introduces unpredictable inter-cell correlations, significantly increasing the complexity of signal detection. Traditional detection methods typically rely on assumptions about known channel noise states, resulting in limited generalization capability in practical applications. To address this issue, three data detection methods based on convolutional neural networks (CNNs) are proposed, which can effectively model and mitigate interference without relying on prior channel information: first, a method combining constrained coding with a multi-layer CNN, which uses constrained coding to determine the sneak path interference state and recover data; second, a dual-CNN framework that first employs a lightweight CNN for sneak path interference identification, followed by a multi-layer CNN for refined detection; third, an approach incorporating transfer learning, which maintains detection accuracy while reducing the required training sample size to one-thousandth of that of traditional methods. Simulation results demonstrate that the proposed method achieves superior bit error rate (BER) performance under unknown channel conditions, with a BER reduction of at least half relative to existing algorithms, approaching the theoretical performance limit. Moreover, the integration of transfer learning reduces the required training samples from \begin{document}$ {10}^{6} $\end{document} to \begin{document}$ 1000 $\end{document}, corresponding to a reduction of three orders of magnitude.  Methods  To address distinct challenges in sneak path interference detection, this paper proposes three methods sequentially:1. The integrated constrained coding aided convolutional neural network (CC-CNN) detection framework effectively addresses the complex inter-cell correlations introduced by sneak path interference. This approach first employs constrained coding to detect the presence of interference and subsequently utilizes a CNN to learn and capture the random correlations under the influence of interference, thereby achieving accurate signal recovery.2. The dual-CNN-based detection method resolves the code rate loss associated with traditional constrained coding. By directly leveraging a CNN to learn and identify sneak path interference patterns from raw data, this method eliminates the need for redundant coding or additional overhead. It ensures high-precision interference detection while preserving the overall code rate performance of the system.3. The transfer learning-based CNN (TL-CNN) detection method overcomes the dependence of high-performance CNNs on large-scale training datasets. By reusing knowledge from pre-trained models, this method enables rapid adaptation to ReRAM signal detection tasks. It significantly reduces the required number of training samples while maintaining high detection accuracy and resource efficiency, thereby enhancing the feasibility of the solution in practical scenarios.  Results and Discussions  Simulation results demonstrate that the performance of the three proposed methods consistently approaches the theoretical lower bound (Fig.6), outperforming baseline methods such as the Belief Propagation (BP) detector, Deep Neural Network (DNN) detector, and Elementary Signal Estimator (ESE) detector. The two-step network achieves performance comparable to that of the single-step network while successfully avoiding code rate loss. Notably, the transfer learning-aided CNN attains near-optimal BER with only 1000 target domain samples, and its performance stabilizes when the sample size exceeds 1000 (Fig.7), fully validating its data efficiency. The integration of SK modules enables the models to effectively capture SPI-induced spatial correlations, while the transfer learning strategy ensures the models’ robust performance under different noise conditions.  Conclusions  The crossbar array architecture of ReRAM is susceptible to sneak-path interference during storage operations, leading to reduced data reliability. To address this issue, this paper proposes three deep learning-based detection methods. Type-I integrates constrained coding with a CNN to achieve efficient and fast interference detection. Type-II adopts a two-stage processing approach: it first classifies interference patterns in the memory array and then performs detection specifically on affected units, thereby ensuring high detection accuracy while minimizing coding rate loss. Type-III introduces a transfer learning framework that leverages a pre-trained model from the source domain, significantly reducing the number of training samples required in the target domain and effectively lowering training overhead. Experimental results show that under different noise conditions, all three proposed methods achieve performance close to the theoretical lower bound, providing an effective solution for enhancing the reliability of ReRAM storage systems.
Physical-layer Security in Visible Light Communications: Fundamental Theories, Key Techniques, and Future Challenges
WANG Jinyuan, YAN Xinrun, LIN Zihan, LI Yuanyuan, LI Zheng, ZHANG Xin
Available online  , doi: 10.11999/JEIT260338
Abstract:
  Significance   Due to the broadcast nature of optical signals, information security represents a critical research direction in visible light communication (VLC). Conventional encryption techniques address network security issues at the upper layers of the protocol stack through access control, cryptographic protection, and end-to-end encryption. However, their security relies on the assumption that eavesdroppers possess limited computational capabilities, an assumption that currently faces significant challenges. In recent years, physical layer security (PLS) has emerged as a novel information security paradigm and has attracted considerable attention from researchers worldwide. PLS exploits the randomness, heterogeneity, and distinctiveness between the main channel and the eavesdropping channel to achieve secure information transmission at the physical layer. To date, extensive research achievements have been made regarding PLS techniques in conventional radio frequency wireless communications (RFWC). Nevertheless, due to substantial differences in frequency bands, transmitted signals, power representations, and channel characteristics, PLS research results from RFWC systems cannot be directly applied to VLC. Although scholars worldwide have conducted research on VLC PLS technology, the foundational theories, key techniques, and future challenges involved in VLC PLS still lack a systematic review. To bridge this gap, this paper presents a comprehensive survey of VLC PLS technology.  Progress   To evaluate and enhance system performance, a classic VLC PLS system model—comprising the received signal model, the input constraint model, and the channel gain model—is initially established. A comprehensive theoretical framework for performance evaluation is then developed, encompassing instantaneous performance metrics, statistical performance metrics, and asymptotic performance metrics. Specifically, to characterize instantaneous performance, existing works on instantaneous secrecy capacity and instantaneous secrecy rate across different scenarios are summarized. As statistical performance metrics, average secrecy capacity, average secrecy rate, secrecy outage probability, probability of strictly positive secrecy capacity, and interception probability are analyzed. To demonstrate asymptotic performance, secrecy diversity order and secrecy degrees of freedom are derived. Furthermore, to enhance the PLS performance, advanced technologies, including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity, are summarized.  Prospects   Despite existing research achievements, numerous challenges remain in VLC PLS. This paper identifies four critical challenges: (i) Accurate PLS performance limit: Deriving exact expression of secrecy capacity under VLC's unique physical constraints remains challenging. (ii) Incomplete evaluation framework: Some key metrics widely used in RFWC have not been investigated in VLC, and the construction of a comprehensive VLC PLS performance evaluation framework remains unresolved. (iii) Limitations of existing methods: Conventional PLS performance enhancement methods typically adopt a “modeling-optimization-verification” separated research paradigm, often falling into a vicious cycle of “inaccurate modeling-suboptimal solutions-limited performance gains”. Therefore, it is imperative to integrate novel technologies (such as deep learning, reinforcement learning, and digital twins) to construct a data-model dual-driven framework for VLC PLS performance enhancement. (iv) Hardware platform gap: The absence of dedicated hardware platforms featuring adversarial topologies and real-time processing capabilities significantly impedes the practical deployment of VLC PLS technologies. Therefore, addressing these challenges is essential for transitioning VLC PLS from theoretical advances to commercial applications.  Conclusions  The broadcast nature of optical signals renders VLC systems vulnerable to eavesdropping attacks. This paper presents a comprehensive survey of PLS in VLC, covering system models, performance metrics (instantaneous, statistical, and asymptotic), and key performance enhancement technologies including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity. Despite significant progress, challenges remain in establishing accurate performance bounds, complete evaluation frameworks, novel enhancement techniques, and practical hardware implementations. By exploiting channel disparities at the physical layer without relying on complex encryption, PLS represents a paradigm shift in security assurance, paving the way for next-generation secure and reliable VLC networks.
Performance Optimization and Gate Oxide Electric Field Analysis of 1200V Trench SiC MOSFET Based on PCL-CSL Collaborative Design
FANG Shaoming, LI Hongda, GAO Yuan
Available online  , doi: 10.11999/JEIT260164
Abstract:
  Objective  1 200 V Silicon Carbide (SiC) trench Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) are key devices in medium- and high-voltage power conversion systems. They feature high switching performance, low conduction loss, and high-temperature stability. However, conventional trench structures suffer from electric-field concentration at the trench corner and bottom gate oxide. This effect can cause the peak gate oxide electric field to exceed the industrial reliability criterion of 3 MV/cm, reducing long-term reliability. In addition, strong trade-offs exist among breakdown voltage, specific on-resistance, threshold voltage, and peak gate oxide electric field. These trade-offs make it difficult to achieve high efficiency and high reliability at the same time. To address these issues, this work studies a synergistic structure that combines deep P-type Column (PCL), Carrier Storage Layer (CSL), and locally thickened gate oxide. The aim is to regulate the electric-field distribution, suppress electric-field concentration, improve carrier transport, and achieve balanced device performance. This study provides a systematic design method for high-reliability and high-performance 1 200 V Trench SiC MOSFETs for industrial applications.  Methods  Numerical device simulations were performed using a Technology Computer-Aided Design (TCAD) platform to analyze and optimize the electrical performance of 1 200 V Trench SiC MOSFETs. To ensure reliable simulations, physical models were used for bandgap narrowing, Shockley-Read-Hall (SRH) recombination, Auger recombination, avalanche breakdown, incomplete dopant ionization, doping- and temperature-dependent mobility, and high-field mobility saturation. A device structure with deep PCL, CSL, and locally thickened bottom gate oxide is constructed to reduce the peak gate oxide electric field and improve device reliability. Key structural and process parameters were swept and quantitatively analyzed. These parameters included epitaxial layer thickness (TEpi), epitaxial layer doping concentration (NEpi), trench width, trench depth, P-Well (PW) implantation dose, PCL spacing, and CSL implantation dose. Static electrical characteristics, including threshold voltage (Vth), specific on-resistance (Ron,sp), Breakdown Voltage (BV), and peak gate oxide electric field (Eox,max) are extracted and evaluated. The final parameter combination is finally determined through a trade-off analysis between conduction performance and long-term device reliability.  Results and Discussions  The simulation results show that the deep PCL structure redirects electric-field lines away from the trench bottom gate oxide and reduces electric-field concentration. When this structure is combined with the locally thickened bottom gate oxide, Eox-max is reduced below 3 MV/cm, meeting the industrial reliability criterion. The CSL broadens the vertical conduction path, reduces current crowding, and decreases Ron,sp. Parameter optimization shows that TEpi, NEpi, trench dimensions, PW implantation dose, and CSL implantation dose determine the trade-off between BV and conduction performance (Fig. 5, Fig. 6, Fig. 9, Fig. 10, and Fig. 19). PCL spacing has a strong effect on electric-field shielding and gate oxide protection (Fig. 16 and Fig. 17). After multi-parameter optimization, the device achieves VTH=4.7 V, BV=1 708 V, Ron,sp=1.57 mΩ·cm2, and Eox-max=2.5 MV/cm (Table 2). These results indicate balanced performance for high-voltage power applications.  Conclusions  A synergistic PCL-CSL structural design for 1 200 V Trench SiC MOSFETs is studied and validated through TCAD simulation. The design addresses key limitations of conventional Trench SiC MOSFETs, including high peak gate oxide electric field, limited breakdown capability, and the trade-off between conduction performance and reliability. The effects of TEpi, NEpi, trench dimensions, PW implantation dose, PCL spacing, and CSL implantation dose on device performance and gate oxide reliability are clarified through parameter sweeping and comparative analysis. With coordinated structural optimization, the optimized device achieves low Ron,sp, high BV, suitable VTH, and suppressed electric-field concentration near the trench bottom oxide. Eox-max is controlled below the 3 MV/cm industrial reliability criterion, which reduces the risk of oxide degradation under high-bias operation. The proposed structural strategy and optimization method provide guidance for the design, simulation, and process development of high-voltage, high-reliability SiC power devices.
Research on Energy Efficiency Optimization of Rotatable Hybrid Intelligent Reflecting Surface Communication
ZHANG Guangchi, GUO Xuan, WANG Luyao, CUI Miao, FU Hao
Available online  , doi: 10.11999/JEIT260119
Abstract:
  Objective  With the evolution of 6G communication networks, reconfigurable intelligent surfaces (RIS) have emerged as a pivotal technology for reshaping wireless environments and enhancing spectral efficiency. However, conventional fixed RIS architectures face two critical challenges in practical deployment: the “angle mismatch” loss, where the effective aperture significantly diminishes when users are located at large angles from the RIS normal, and the “energy consumption bottleneck,” caused by the high cumulative power consumption of radio frequency (RF) circuits and static control elements in large-scale arrays. Existing research often treats mechanical rotation and element switching in isolation, lacking a unified framework to balance the trade-off between mechanical/circuit energy consumption and communication gain. To address these limitations, this paper investigates a rotatable and switchable hybrid RIS (H-RIS) assisted downlink communication system. The primary objective is to maximize the system’s energy efficiency (EE) by jointly optimizing the base station transmit power, subarray activation states, physical rotation angles, and electronic phase shifts. This approach aims to introduce mechanical rotation degrees of freedom to compensate for path loss and employ dynamic switching mechanisms to reduce redundant power consumption, thereby achieving sustainable green communication.  Methods  A joint optimization framework is established for the H-RIS aided single-user multiple-input single-output (MISO) system. The system model explicitly accounts for the dynamic power consumption induced by mechanical rotation and the static power consumption of active subarrays. The resulting optimization problem is formulated as a non-convex mixed-Integer non-linear programming (MINLP) problem, involving coupled binary variables (activation status) and continuous variables (power, angles, phases). To solve this challenging problem, a block coordinate descent (BCD)-based alternating optimization (AO) algorithm is proposed to decouple the variables into three sub-problems.Firstly, to tackle the exponential complexity caused by binary switching variables, a channel contribution-based ranking strategy is developed. By performing eigenvalue decomposition on the cascaded channel correlation matrix, the priority of each subarray is quantified, reducing the search space from exponential to linear.Secondly, for the power allocation sub-problem, the non-convex fractional objective function is transformed into a parametric subtractive form using the Dinkelbach algorithm, which is then solved via the interior-point method.Thirdly, for the physical rotation and electronic phase optimization, the problem is decomposed into single-variable sub-problems. A Golden Section Search algorithm is employed to iteratively find the optimal rotation angle and phase shift for each subarray within bounded constraints, ensuring the monotonic convergence of the objective function.  Results and Discussions  Extensive simulations are conducted to evaluate the performance of the proposed H-RIS scheme compared with benchmark schemes, including “Only-Rotation” (always on), “Only-Switching” (fixed angle), and “Conventional” (fixed and always on).The simulation results regarding the maximum transmit power Pmax(Fig. 2 and Fig. 3) demonstrate that the proposed method achieves the highest energy efficiency across the entire power range. Specifically, in the low power regime, the proposed algorithm intelligently turns off redundant subarrays where the rate gain cannot offset the circuit power cost, thereby significantly outperforming the “Only-Rotation” scheme which suffers from high static power consumption.The impact of user distance is also analyzed (Fig. 4 and Fig. 5). Results indicate that the proposed scheme maintains high spectral efficiency comparable to the “Only-Rotation” scheme by dynamically adjusting the rotation angles to align with the Line-of-Sight (LoS) path, effectively compensating for the angle mismatch loss observed in the “Only-Switching” and “Conventional” schemes.Furthermore, the activation pattern of the subarray varies in a “U” shape with distance (Table 1), which allows for flexible adjustment of array size and orientation according to user-RIS geometry.  Conclusions  This paper proposes an energy-efficient transmission scheme for H-RIS aided communication systems by integrating mechanical rotation and dynamic switching capabilities. A low-complexity BCD-based algorithm is developed to jointly optimize the transceiver design. The results confirm that introducing mechanical rotation significantly mitigates the angle mismatch loss, while the proposed channel contribution-based switching strategy effectively eliminates redundant energy consumption. The proposed H-RIS architecture offers a superior trade-off between spectral efficiency and energy efficiency compared to traditional fixed RIS architectures, providing a viable solution for future green 6G networks.
CRLB Optimization for O-RIS-Assisted VLP Systems
ZHANG Zengjie, WU Qi, ZHANG Jian, DUAN Ruijie, FENG Yunhan
Available online  , doi: 10.11999/JEIT260120
Abstract:
  Objective  With the rapid development of indoor location-based services, Visible Light Positioning (VLP) has emerged as a promising high-accuracy positioning technology. The integration of Optical Reconfigurable Intelligent Surfaces (O-RIS) into VLP systems can effectively enhance signal coverage and improve positioning performance. However, optimizing the positioning accuracy and fairness across different user areas in RIS-assisted VLP systems remains a challenging issue. This study focuses on optimizing the Cramer-Rao Lower Bound (CRLB) of the system under both near-field and far-field channel models, aiming to enhance overall positioning precision and fairness through RIS configuration.  Methods  Under the far-field channel model assumption, the RIS orientation optimization problem is formulated as a received power maximization problem. A positioning algorithm combining Particle Swarm Optimization (PSO) and N-step iteration is proposed to dynamically adjust the RIS orientation optimally without prior knowledge of the receiver’s position. Under the near-field channel model assumption, the allocation problem between RIS elements and LEDs is constructed as a Markov Decision Process (MDP). A reinforcement learning method based on experience replay and knowledge utilization is designed to solve this problem, aiming to minimize the CRLB while ensuring positioning fairness for users in different regions.  Results and Discussions  Simulation results demonstrate that the proposed algorithms effectively enhance system positioning performance under both models. In the far-field model, the PSO-based iterative algorithm achieves dynamic optimization of RIS orientation, significantly improving positioning accuracy (Fig. 3). Under the near-field model, the reinforcement learning approach not only minimizes the CRLB but also considerably improves positioning fairness across the entire area, with a noticeable reduction in performance disparity among users in different zones (Fig. 5, Fig. 6). Comparative experiments show that the proposed methods outperform conventional RIS configuration strategies in terms of both average positioning error and fairness index (Table 1).  Conclusions  This paper investigates CRLB optimization methods for O-RIS-assisted VLP systems under near-field and far-field channel models. In the far-field scenario, a PSO-based iterative algorithm is proposed to optimize RIS orientation, enhancing positioning accuracy without requiring prior receiver location information. In the near-field scenario, a reinforcement learning-based approach is designed to optimize RIS element–LED allocation, which effectively minimizes the CRLB and improves positioning fairness across the whole area. Simulation results validate the effectiveness of the proposed algorithms in both models. Future work may consider more practical channel impairments and multi-user scenarios to further improve the robustness and scalability of the system.
A Radio Frequency Fingerprint Open-set Identification MethodCombining Multi-scale Wavelet Front-end and Hyperspherical Metric Learning
TIAN Xinyu, LI Zirui, ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260214
Abstract:
  Objective  Open-set Radio Frequency Fingerprint (RFF) identification under low Signal-to-Noise Ratio (SNR) conditions is challenging because fingerprint features are easily masked by noise, multipath effects induce nonlinear distortions, and existing methods struggle with feature extraction and unknown device detection. This study proposes a deep learning framework that integrates a multi-scale wavelet front-end with hyperspherical metric learning to achieve robust open-set RFF identification.  Methods  The proposed method, MS-RANet, comprises three key components. First, a multi-scale wavelet front-end based on one-dimensional stationary wavelet transform performs full-resolution, multi-scale decomposition of I/Q signals, preserving discriminative fingerprint information while suppressing noise. Second, a multi-scale residual attention network incorporates deep residual learning, global self-attention, and Bidirectional LSTM (BiLSTM) to enhance sensitivity to subtle fingerprint features and capture long-range temporal dependencies. Third, hyperspherical metric learning constrains the feature space onto a unit hypersphere, optimizing angular margins to produce compact intra-class and separable inter-class feature distributions. Unknown devices are subsequently detected using cosine similarity.  Results and Discussions  Experiments on a high-fidelity IEEE 802.11 simulation dataset demonstrate the effectiveness of MS-RANet. The method achieves an average classification accuracy of 65.34% across SNR levels from –5 dB to 20 dB, and an Area Under the Curve (AUC) of 0.81 at –5 dB SNR, outperforming DNN, GRU, CNN-LSTM, ResNet50, and DRSN-CA. Confusion matrices and Receiver Operating Characteristic (ROC) curves confirm robustness under extreme channel conditions. t-SNE visualization shows well-separated, compact clusters for known devices, while unknown samples are effectively isolated from known class regions. Ablation studies verify the contributions of the multi-scale wavelet front-end, global attention, BiLSTM, and hyperspherical metric learning modules.  Conclusions  This study presents a robust open-set RFF identification method combining a multi-scale wavelet front-end with hyperspherical metric learning. The framework exhibits strong noise resilience, enhanced feature discrimination, and reliable detection of unknown devices under low-SNR and multipath fading conditions. Future work will focus on reducing computational complexity, improving inference speed, evaluating generalization across diverse scenarios and protocols, and integrating the method with complementary physical-layer security mechanisms for collaborative authentication.
Semantic Relation-enhanced Adaptive Graph Representation Learning for Next POI Recommendation
WANG Zhuolu, XU Shenghua, WANG Yong, JIANG Shunshun
Available online  , doi: 10.11999/JEIT251357
Abstract:
  Objective  In recent years, next Point Of Interest (POI) recommendation has played an increasingly important role in Location-Based Social Networks (LBSNs). However, existing Graph Representation Learning (GRL)-based recommendation methods have struggled to balance node distributions across different domains (i.e., node types) effectively and have often overlooked feature differences among heterogeneous relations. Thus, complex semantic dependencies in contextual information cannot be fully captured when users’ temporal preference patterns are modeled.  Methods  To address these issues, a next POI recommendation method based on Semantic Relation-enhanced adaptive Graph Representation Learning (SR-GRL) is proposed. A heterogeneous transition graph is constructed to integrate three entity types, namely POIs, POI categories, and regions, and their complex interrelationships. An adaptive balanced random walk sampling strategy is designed to balance node distributions across different domains dynamically and to reduce information redundancy. A type-aware attention mechanism is then used to learn semantic associations among nodes through relation-specific transformation matrices, so that feature differences across node types can be identified effectively. The obtained disentangled POI representations are then used for spatiotemporal encoding of user check-in sequences, and a self-attention mechanism is applied to aggregate users, temporal preference features. Finally, next POI recommendation is generated through a Softmax function.  Results and Discussions  Experiments on the Foursquare datasets from Tokyo and New York and the Sina Weibo dataset from Shanghai show that, compared with state-of-the-art baselines, the SR-GRL method achieves Recall@10 improvements of 2.22%\begin{document}$ \sim $\end{document}24.16%, F1@10 improvements of 1.16%\begin{document}$ \sim $\end{document}10.48%, and NDCG@10 improvements of 3.01%\begin{document}$ \sim $\end{document}17.37%, indicating better recommendation performance.  Conclusions  Overall, the SR-GRL approach can balance the distributions of different node types dynamically and strengthen the modeling of complex semantic dependencies in heterogeneous contextual information.
Communication, Computation, and Caching Resource Collaboration for Heterogeneous Artificial Intelligence Generated Content Service Provisioning
WU Mengru, GAO Yu, ZHAO Bo, XU Bo, SUN Hao, GUO Lei
Available online  , doi: 10.11999/JEIT251300
Abstract:
  Objective  In the Artificial Intelligence of Things (AIoT), Edge Servers (ESs) provide intelligent content generation services to AIoT devices by utilizing cached Artificial Intelligence Generated Content (AIGC) models. However, the limited computing resources and caching capacity of ESs make it difficult to support the large-scale caching demands of heterogeneous AIGC services. To address this issue, a communication, computation, and caching resource collaboration scheme is proposed based on a combined cloud-edge and edge-edge collaborative framework. The scheme considers three representative AIGC services: lightweight AIGC services, computation-intensive AIGC services, and preprocessing-based AIGC services. The objective is to minimize the total AIGC service latency through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions.  Methods  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. First, an AIGC service-oriented AIoT system model is established to incorporate both cloud-edge and edge-edge collaboration. An optimization problem is then formulated to minimize the total latency of AIGC services through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions. Because the formulated problem is non-convex, an Alternating Optimization (AO) algorithm is proposed. The original problem is decomposed into three subproblems. These subproblems are solved using the Successive Convex Approximation (SCA) method, Karush-Kuhn-Tucker (KKT) conditions, and an improved Harris Hawks Optimization (HHO) algorithm.  Results and Discussions  Simulation experiments compare the proposed joint optimization scheme with three baseline methods: Particle Swarm Optimization (PSO), fixed resource allocation, and random offloading and caching. First, the convergence of the proposed AO algorithm is verified (Fig. 2). The results show that the algorithm converges rapidly within a limited number of iterations across different subproblems. Second, increasing transmission bandwidth significantly reduces the total AIGC service latency (Fig. 3). This occurs because each device obtains more bandwidth resources for task transmission, and the ES can allocate more bandwidth to deliver generated content in the downlink. Furthermore, the total AIGC service latency decreases as the ES storage capacity increases for all schemes (Fig. 4). Greater storage capacity enables the ES to store more AIGC models, which reduces the transmission delay between the ES and the cloud server. Moreover, when the required floating-point operations per bit increase, the total AIGC service latency rises significantly across all schemes (Fig. 5). Finally, the total AIGC service latency decreases as the maximum transmit power of the Base Station (BS) increases (Fig. 6). This occurs because higher BS transmit power improves the downlink signal-to-noise ratio, which increases the downlink transmission rate and reduces overall service latency. The proposed scheme demonstrates better performance than the baseline schemes, particularly under high computational demand.  Conclusions  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. The objective is to minimize total AIGC service latency through joint optimization of the transmit power of AIoT devices and BSs, computing resource allocation, AIGC model deployment, and service offloading decisions under computation and caching resource constraints. Because the formulated problem is a mixed-integer nonlinear programming problem, an efficient AO algorithm is developed. The original optimization problem is decomposed into three subproblems, which are solved using the SCA algorithm, KKT conditions, and the HHO algorithm, respectively. Simulation results show that the proposed algorithm reduces the total AIGC service latency compared with the baseline schemes.
Adversarial Attacks on 3D Target Recognition Driven by Gradient Adaptive Adjustment
LIU Weiquan, SHEN Xiaoying, LIU Dunqiang, SUN Yanwen, CAI Guorong, ZANG Yu, SHEN Siqi, WANG Cheng
Available online  , doi: 10.11999/JEIT251264
Abstract:
  Objective   Robust environmental perception is essential for intelligent driving systems. Light Detection And Ranging (LiDAR) provides high-resolution 3D point cloud data and serves as a core information source for object detection and recognition. However, deep learning models for 3D point cloud recognition show notable vulnerability to adversarial attacks. Small, imperceptible perturbations can cause severe classification errors and threaten system safety. Existing attack methods have improved the Attack Success Rate (ASR), but the perturbations they generate often lack concealment, create outliers, and show poor imperceptibility because they do not adequately preserve the geometric structure of point clouds. This reduces their suitability for realistic security evaluation of optoelectronic perception systems. Developing an attack method that maintains a high success rate while preserving geometric consistency and imperceptibility is therefore critical. This study addresses this need by proposing a framework that incorporates point cloud geometry into perturbation generation.  Methods   A Gradient Adaptive Adjustment (GAA) adversarial attack method for 3D point cloud recognition is proposed. The framework (Fig. 2) includes three coordinated modules. The 3D Point Cloud Salient Region Extraction module evaluates decision-level vulnerability using Shapley value analysis to identify and rank point subsets with the strongest influence on classifier output. Perturbations are then concentrated in these sensitive regions. A curvature-weighted gradient mechanism integrates local geometric priors. For each point in the salient region, a local covariance matrix is computed from its k-nearest neighbors. Principal component analysis generates eigenvalues and eigenvectors, which are used to compute a curvature measure. A Gaussian kernel function produces curvature-dependent weights that are applied to backpropagated gradients. This suppresses perturbations in high-curvature areas and encourages them in low-curvature regions to preserve local shape morphology. A principal curvature direction constrained 0ptimization module further refines the perturbation direction. The weighted gradient is projected onto the principal curvature directions, and the projection components are fused using coefficients derived from the corresponding eigenvalues. This aligns the perturbation with natural geometric trends and avoids unnatural deformation. An adaptive optimization algorithm then minimizes a multi-objective loss balancing attack success, geometric similarity (via chamfer distance and hausdorff distance), and perturbation sparsity. The adversarial point cloud is iteratively updated based on the saliency map, curvature-weighted gradients, and principal direction constraints.  Results and Discussions   Experiments on ModelNet40, ShapeNetPart, and KITTI were conducted using PointNet, DGCNN, and PointConv. The GAA method showed strong performance. On ModelNet40 with PointNet, it achieved a 97.69% ASR with an average of 28 perturbed points, outperforming ten baselines such as AL-Adv (92.92% ASR, 40 points) and Kim et al. (89.38% ASR, 36 points) (Table 1). It also produced lower geometric distortion, as indicated by smaller Chamfer Distance and Hausdorff Distance values. Visual results (Fig. 4) show that GAA produces fewer outliers and more natural adversarial point clouds compared with methods such as AL-Adv. The method generalized well across architectures, reaching 99.78% ASR on DGCNN and 96.91% on PointConv (Table 2), with similar performance on ShapeNetPart (Table 3). Ablation experiments on the number of salient regions (K) showed consistent improvements in ASR and reduced geometric distortion as K increased from 1 to 6 (Table 4, Fig. 5), confirming the advantage of targeting multiple critical regions. Tests on the KITTI dataset demonstrated strong performance in real-world, noisy environments. The method maintained high ASRs, such as 99.33% on PointNet, with limited perturbations (Table 5). An ablation study on K indicated that K=4 offers an effective balance between success rate and perturbation cost for PointNet (Table 6).  Conclusions   This study presents a GAA method for adversarial attacks on 3D point cloud recognition. By combining a Shapley value-based saliency analyzer, a curvature-weighted gradient mechanism, and a principal curvature direction constraint, the method generates adversarial examples that achieve high attack success while preserving geometric consistency. Experiments show that GAA minimizes perceptual distortion and perturbs fewer points across datasets and models. The method provides a practical tool for vulnerability analysis and supports the development of more robust and secure optoelectronic perception systems for intelligent driving. Future work will examine robustness under adverse conditions and assess physical-world implications.
Performance Analysis and Rapid Prediction of Long-range Underwater Acoustic Communications in Uncertain Deep-sea Environments
CHEN Xiangmei, TAI Yupeng, WANG Haibin, HU Chenghao, WANG Jun, WANG Diya
Available online  , doi: 10.11999/JEIT251244
Abstract:
  Objective  In complex and dynamically changing deep-sea environments, the performance of underwater acoustic communications shows substantial variability. Feedback-based channel estimation and parameter adaptation are impractical in long-range scenarios because platform constraints prevent reliable feedback channels and the slow propagation of sound introduces significant delay. In typical long-range systems, environmental dynamics are often ignored and communication parameters are selected heuristically, which frequently leads to mismatches with actual channel conditions and causes communication failures or reduced efficiency. Predictive methods able to assess performance in advance and support feed-forward parameter adjustment are therefore required. This study proposes a deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications under uncertain environmental conditions to enable efficient and reliable parameter–channel matching without feedback.  Methods  A feed-forward method for underwater acoustic communication performance analysis and rapid prediction is developed using deep-learning-based sound-field uncertainty estimation. A neural network is first used to estimate probability distributions of Transmission Loss (TL PDFs) at the receiver under dynamic environments. TL PDFs are then mapped to probability distributions of the Signal-to-Noise Ratio (SNR PDFs), enabling communication performance evaluation without real-time feedback. Statistical channel capacity and outage capacity are analyzed to characterize the theoretical upper limits of achievable rates in dynamic conditions. Finally, by integrating the SNR distribution with the bit-error-rate characteristics of a representative deep-sea single-carrier communication system under the corresponding channel, a rate–reliability prediction model is constructed. This model estimates the probability of reliable communication at different data rates and serves as a practical tool for forecasting link performance in highly dynamic and feedback-limited underwater acoustic environments.  Results and Discussions  The method is validated using simulation data and sea trial data. The TL PDFs predicted by the deep learning model show strong consistency with the traditional Monte Carlo (MC) method across multiple receiver locations (Fig. 6). Under identical computational settings, deep-learning-based TL PDF prediction reduces computation time by 2\begin{document}$ \sim $\end{document}3 orders of magnitude compared with the MC method. The chained mapping from TL PDFs to SNR PDFs and then to channel capacity metrics accurately represents the probabilistic features of communication performance under uncertain conditions (Fig. 7 and Fig. 8). The rate–reliability curves derived from the deep-learning-based TL PDFs are highly consistent with MC-based results. In the high sound-intensity region, prediction errors for reliable communication probabilities across data rates range from 0.1% to 3%, and in the low sound-intensity region errors are approximately 0.3% to 5% (Fig. 12). Sea trial results further indicate that predicted rate–reliability performance agrees well with measured data. In the convergence zone, deviations between predicted and measured reliability probabilities at each rate range from 0.9% to 4%, and in the shadow zone from 1% to 9% (Fig. 18). Under a 90% reliability requirement, the maximum achievable rates predicted by the method match the measurements in both the convergence and shadow zones, demonstrating accuracy and practical applicability in complex channel environments.  Conclusions  A deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications in uncertain deep-sea environments is developed and validated. The framework builds a chained mapping from environmental parameters to TL PDFs, SNR PDFs, and communication performance metrics, enabling quantitative capacity assessment under dynamic ocean conditions. Predictive “rate–reliability’’ profiles are obtained by integrating probabilistic propagation characteristics with the performance of a representative deep-sea single-carrier system under the corresponding channel, providing guidance for parameter selection without feedback. Sea trial results confirm strong agreement between predicted and measured performance. The proposed approach offers a technical pathway for feed-forward performance analysis and dynamic adaptation in long-range deep-sea communication systems, and can be extended to other communication scenarios in dynamic ocean environments.
Inverse Design of a Silicon-Based Compact Polarization Splitter-Rotator
HUI Zhanqiang, ZHANG Xinglong, HAN Dongdong, LI Tiantian, GONG Jiamin
Available online  , doi: 10.11999/JEIT250858
Abstract:
  Objective  The Polarization Splitter-Rotator (PSR) is a key device used to control the polarization state of light in Photonic Integrated Circuits (PICs). Device size has become a major constraint on integration density in PICs. Traditional design methods are time-consuming and tend to yield larger device footprints. Inverse design, by contrast, determines structural parameters through optimization algorithms according to target performance and enables compact devices to be obtained while maintaining functionality. This strategy is now applied to wavelength and mode division multiplexers, all-optical logic gates, power splitters, and other integrated photonic components. The objective of this work is to use inverse design to address size limitations in silicon-based PSRs by combining the Momentum Optimization algorithm with the Adjoint Method. This combined approach improves the integration level of PICs and provides a feasible pathway for the miniaturization of other photonic devices.  Methods  The design region is defined on a 220 nm Silicon-on-Insulator (SOI) wafer and is discretized into 25×50 cylindrical elements. Each element has a 50 nm radius, a 150 nm height, and an initial relative permittivity of 6.55. The adjoint method is used to obtain gradient information across the design region, and this gradient is processed with the Momentum Optimization algorithm. The relative permittivity of each element is then updated according to the processed gradient. During optimization, the momentum factor is dynamically adjusted with the iteration number to accelerate convergence, and a linear bias is applied to guide the permittivity toward the values of silicon and air as the iterations progress. After optimization, the elements are binarized based on their final permittivity: values below 6.55 are assigned to air, whereas values above 6.55 are assigned to silicon. This results in a structure containing irregularly distributed air holes. To compensate for performance loss introduced during binarization, the etching depth of air holes with pre-binarization permittivity between 3 and 6.55 is optimized. Adjacent air holes are merged to reduce fabrication errors. The final device consists of air holes with five radii, among which three larger-radius types are selected for further refinement. Their etching radii and depths are optimized to recover remaining performance loss. Device performance is evaluated through numerical analysis. Calculated parameters include Insertion Loss (IL), Crosstalk (CT), Polarization Extinction Ratio (PER), and bandwidth. Tolerance analysis is also conducted to assess robustness under fabrication variations.  Results and Discussions   A compact PSR is designed on a 220 nm SOI wafer with dimensions of 5 μm in length and 2.5 μm in width. During optimization, the momentum factor in the Momentum Optimization algorithm is dynamically adjusted. A larger momentum factor is applied in the early stage to accelerate escape from local maxima or plateau regions, whereas a smaller momentum factor is used in later iterations to increase the weight of the current gradient. Compared with other optimization strategies, this algorithm requires only 20%~33% of the iteration count needed by alternative methods to reach a Figure of Merit (FOM) of 1.7, which improves optimization efficiency. Numerical analysis shows that the device achieves stable performance across the 1 520~1 575 nm wavelength range. The IL remains low (TM0 < 1 dB, TE0 < 0.68 dB), and the CT is effectively suppressed (TM0 < –23 dB, TE0 < –25.2 dB). The PER is high (TM0 > 17 dB, TE0 > 28.5 dB). Tolerance analysis indicates strong robustness to fabrication variations. Within the 1 520~1 540 nm range, performance remains stable under etching depth offsets of ±9 nm and etching radius offsets of ±5 nm, demonstrating reliable manufacturability.  Conclusions   Numerical analysis demonstrates that combining the adjoint method with the Momentum Optimization algorithm is a feasible strategy for designing an integrated PSR. The design principle relies on controlling light propagation through adjustments to the relative permittivity, which determine the distribution and placement of air holes to achieve polarization splitting and rotation. Compared with traditional design approaches, inverse design uses the design region more efficiently and enables a more compact device structure. The proposed PSR is markedly smaller and shows enhanced fabrication tolerance. It is suitable for future large-scale PICs and provides useful guidance for the miniaturization of other photonic devices.
Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning in Wireless Networks
LIU Jingyuan, MA Ke, XU Runchen, CHANG Zheng
Available online  , doi: 10.11999/JEIT251221
Abstract:
  Objective  Multimodal Federated Learning (MFL) uses complementary information from multiple modalities, yet in wireless edge networks it is restricted by limited energy and frequent missing modalities because many clients store only images or only reports. This study presents Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning (CREEMFL), which applies selective completion and joint communication–computation optimization to reduce training energy under latency and wireless constraints.  Methods  CREEMFL completes part of the incomplete samples by querying a public multimodal subset, and processes the remaining samples through zero padding. Each selected user downloads the global model, performs image-to-text or text-to-image retrieval, conducts local multimodal training, and uploads model updates for aggregation. An energy–delay model couples local computation and wireless communication and treats the required number of global rounds as a function of retrieval ratios. Based on this model, an energy minimization problem is formulated and solved using a two-layer algorithm with an outer search over retrieval ratios and an inner optimization of transmission time, Central Processing Unit (CPU) frequency, and transmit power.  Results and Discussions  Simulations on a single-cell wireless MFL system show that increasing the ratio of completing text from images improves test accuracy and reduces total energy. In contrast, a large ratio of completing images from text provides limited accuracy gain but increases energy consumption (Fig. 3, Fig. 4). Compared with four representative baselines, CREEMFL achieves shorter completion time and lower total energy across a wide range of maximum average transmit powers (Fig. 5, Fig. 6). For CREEMFL, increased system bandwidth further reduces completion time and energy consumption (Fig. 7, Fig. 8). Under different user modality compositions, CREEMFL also attains higher test accuracy than local training, zero padding, and cross-modal retrieval without energy optimization (Fig. 9).  Conclusions  CREEMFL integrates selective cross-modal retrieval and joint communication–computation optimization for energy-efficient MFL. By treating retrieval ratios as variables and modeling their effect on global convergence rounds, it captures the coupling between per-round costs and global training progress. Simulations verify that CREEMFL reduces training completion time and total energy while preserving classification accuracy in resource-constrained wireless edge networks.
Breakthrough in Solving NP-Complete Problems Using Electronic Probe Computers
XU Jin, YU Le, YANG Huihui, JI Siyuan, ZHANG Yu, YANG Anqi, LI Quanyou, LI Haisheng, ZHU Enqiang, SHI Xiaolong, WU Pu, SHAO Zehui, LENG Huang, LIU Xiaoqing
Available online  , doi: 10.11999/JEIT250352
Abstract:
This study presents a breakthrough in addressing NP-complete problems using a newly developed Electronic Probe Computer (EPC60). The system employs a hybrid serial–parallel computational model and performs large-scale parallel operations through seven probe operators. In benchmark tests on 3-coloring problems in graphs with 2,000 vertices, EPC60 achieves 100% accuracy, outperforming the mainstream solver Gurobi, which succeeds in only 6% of cases. Computation time is reduced from 15 days to 54 seconds. The system demonstrates high scalability and offers a general-purpose solution for complex optimization problems in areas such as supply chain management, finance, and telecommunications.  Objective   NP-complete problems pose a fundamental challenge in computer science. As problem size increases, the required computational effort grows exponentially, making it infeasible for traditional electronic computers to provide timely solutions. Alternative computational models have been proposed, with biological approaches—particularly DNA computing—demonstrating notable theoretical advances. However, DNA computing systems continue to face major limitations in practical implementation.  Methods  Computational Model: EPC is based on a non-Turing computational model in which data are multidimensional and processed in parallel. Its database comprises four types of graphs, and the probe library includes seven operators, each designed for specific graph operations. By executing parallel probe operations, EPC efficiently addresses NP-complete problems.Structural Features:EPC consists of four subsystems: a conversion system, input system, computation system, and output system. The conversion system transforms the target problem into a graph coloring problem; the input system allocates tasks to the computation system; the computation system performs parallel operations via probe computation cards; and the output system maps the solution back to the original problem format.EPC60 features a three-tier hierarchical hardware architecture comprising a control layer, optical routing layer, and probe computation layer. The control layer manages data conversion, format transformation, and task scheduling. The optical routing layer supports high-throughput data transmission, while the probe computation layer conducts large-scale parallel operations using probe computation cards.  Results and Discussions  EPC60 successfully solved 100 instances of the 3-coloring problem for graphs with 2,000 vertices, achieving a 100% success rate. In comparison, the mainstream solver Gurobi succeeded in only 6% of cases. Additionally, EPC60 rapidly solved two 3-coloring problems for graphs with 1,500 and 2,000 vertices, which Gurobi failed to resolve after 15 days of continuous computation on a high-performance workstation.Using an open-source dataset, we identified 1,000 3-colorable graphs with 1,000 vertices and 100 3-colorable graphs with 2,000 vertices. These correspond to theoretical complexities of O(1.3289n) for both cases. The test results are summarized in Table 1.Currently, EPC60 can directly solve 3-coloring problems for graphs with up to n vertices, with theoretical complexity of at least O(1.3289n).On April 15, 2023, a scientific and technological achievement appraisal meeting organized by the Chinese Institute of Electronics was held at Beijing Technology and Business University. A panel of ten senior experts conducted a comprehensive technical evaluation and Q&A session. The committee reached the following unanimous conclusions:1. The probe computer represents an original breakthrough in computational models.2. The system architecture design demonstrates significant innovation.3. The technical complexity reaches internationally leading levels.4. It provides a novel approach to solving NP-complete problems.Experts at the appraisal meeting stated, “This is a major breakthrough in computational science achieved by our country, with not only theoretical value but also broad application prospects.” In cybersecurity, EPC60 has also demonstrated remarkable potential. Supported by the National Key R&D Program of China (2019YFA0706400), Professor Xu Jin’s team developed an automated binary vulnerability mining system based on a function call graph model. Evaluation of the system using the Modbus Slave software showed over 95% vulnerability coverage, far exceeding the 75 vulnerabilities detected by conventional depth-first search algorithms. The system also discovered a previously unknown flaw, the “Unauthorized Access Vulnerability in Changyuan Shenrui PRS-7910 Data Gateway” (CNVD-2020-31406), highlighting EPC60’s efficacy in cybersecurity applications.The high efficiency of EPC60 derives from its unique computational model and hardware architecture. Given that all NP-complete problems can be polynomially reduced to one another, EPC60 provides a general-purpose solution framework. It is therefore expected to be applicable in a wide range of domains, including supply chain management, financial services, telecommunications, energy, and manufacturing.  Conclusions   The successful development of EPC offers a novel approach to solving NP-complete problems. As technological capabilities continue to evolve, EPC is expected to demonstrate strong computational performance across a broader range of application domains. Its distinctive computational model and hardware architecture also provide important insights for the design of next-generation computing systems.
Personalized Federated Learning Method Based on Collation Game and Knowledge Distillation
SUN Yanhua, SHI Yahui, LI Meng, YANG Ruizhe, SI Pengbo
Available online  , doi: 10.11999/JEIT221203
Abstract:
To overcome the limitation of the Federated Learning (FL) when the data and model of each client are all heterogenous and improve the accuracy, a personalized Federated learning algorithm with Collation game and Knowledge distillation (pFedCK) is proposed. Firstly, each client uploads its soft-predict on public dataset and download the most correlative of the k soft-predict. Then, this method apply the shapley value from collation game to measure the multi-wise influences among clients and quantify their marginal contribution to others on personalized learning performance. Lastly, each client identify it’s optimal coalition and then distill the knowledge to local model and train on private dataset. The results show that compared with the state-of-the-art algorithm, this approach can achieve superior personalized accuracy and can improve by about 10%.
The Range-angle Estimation of Target Based on Time-invariant and Spot Beam Optimization
Wei CHU, Yunqing LIU, Wenyug LIU, Xiaolong LI
Available online  , doi: 10.11999/JEIT210265
Abstract:
The application of Frequency Diverse Array and Multiple Input Multiple Output (FDA-MIMO) radar to achieve range-angle estimation of target has attracted more and more attention. The FDA can simultaneously obtain the degree of freedom of transmitting beam pattern in angle and range. However, its performance is degraded due to the periodicity and time-varying of the beam pattern. Therefore, an improved Estimating Signal Parameter via Rotational Invariance Techniques (ESPRIT) algorithm to estimate the target’s parameters based on a new waveform synthesis model of the Time Modulation and Range Compensation FDA-MIMO (TMRC-FDA-MIMO) radar is proposed. Finally, the proposed method is compared with identical frequency increment FDA-MIMO radar system, logarithmically increased frequency offset FDA-MIMO radar system and MUltiple SIgnal Classification (MUSIC) algorithm through the Cramer Rao lower bound and root mean square error of range and angle estimation, and the excellent performance of the proposed method is verified.
Radar, Sonar,Navigation and Array Signal Processing
Indoor Visible Light Positioning Based on CNN-MLP Multi-Feature Fusion under Random Receiver Tilt Conditions
JIA Kejun, WANG Jian, MAO Lifei, YOU Wei, HUANG Ziyang, PENG Duo
Available online  , doi: 10.11999/JEIT251021
Abstract:
  Objective  Traditional Visible Light Positioning (VLP) methods based on Received Signal Strength (RSS) are unstable when the receiver undergoes orientation perturbations. Such perturbations disrupt the correspondence between optical power and spatial position, which makes reliable three-dimensional (3D) positioning difficult. Existing approaches usually rely on Inertial Measurement Units (IMUs) to obtain orientation information. However, sensor fusion increases system complexity and hardware cost and also introduces cumulative errors. To address these issues, this paper proposes a positioning method that fuses incidence-angle cosine estimation derived from a Photodiode (PD) array with RSS information, which enables high-accuracy 3D indoor positioning under receiver orientation perturbations.  Methods  In the proposed fusion-based positioning method, a multi-PD array structure is first adopted, and a Local Coordinate System (LCS) is established at the array center. Constraint equations are then constructed from differences in the optical power received by the PDs in the array. A Gauss-Newton iterative algorithm is used to estimate the incident light direction vector. By exploiting the orthogonal rotation invariance between the LCS and the Global Coordinate System (GCS), the incident-angle cosine is estimated without orientation sensors. A serial CNN-MLP fusion network is then constructed, in which the estimated incident-angle cosine is introduced as an additional positioning feature beyond RSS-based localization. The network jointly models the RSS and incident-angle cosine information received by the PD array and maps them to 3D spatial coordinates. Finally, training samples are generated by Latin Hypercube Sampling (LHS) to uniformly sample spatial positions and orientation dimensions, thereby improving the representativeness of the training dataset.  Results and Discussions  Simulation experiments are conducted in a 4 m × 4 m × 2.5 m indoor environment. First, the effects of different numbers of PDs and different tilt angles on the accuracy of incident-angle cosine estimation and spatial coverage are evaluated (Fig. 6), and the Cumulative Distribution Functions (CDFs) of positioning errors under different array configurations are compared (Fig. 7). The results show that a 3-PD array with a tilt angle of 40° achieves the best balance of cost, coverage, and positioning accuracy. Next, positioning performance under different receiver tilt angles is analyzed. When the tilt angle is small, more than 70% of positioning errors are below 5 cm. Even when the receiver is tilted by up to 55°, the average error remains within 11.7 cm (Fig. 8). Comparisons of error components show that the error along the Z-axis is significantly smaller than those along the X- and Y-axes (Fig. 9). Further tests are conducted at a height of 0.0 m, which is covered by the training data, and at an unseen height of 0.6 m, which is not included in the training set (Fig. 10). The results show that the proposed model does not strongly depend on a specific height plane and maintains stable 3D positioning performance at unseen heights. Finally, the proposed method is compared with related positioning schemes. It outperforms existing methods in terms of CDF convergence speed, RMSE, and standard deviation (Fig. 11), with an average error reduction of about 2.5 cm and an RMSE reduction of 31.58% compared with Ref. [13].  Conclusions  This paper estimates the incident-angle cosine at the receiver by exploiting differences in the optical power received by different PDs in an array, and introduces this cosine value as a joint positioning feature into conventional RSS-based localization. This design alleviates the instability of position mapping caused by relying only on RSS under random receiver perturbations. By combining the spatial feature extraction capability of CNNs with the nonlinear modeling strength of MLPs, the proposed method effectively maps positioning features to 3D spatial coordinates. The approach reduces reliance on orientation sensors such as IMUs, while overcoming the sensitivity of traditional geometric positioning methods to noise and high-dimensional nonlinear features. Under varying heights and receiver orientations, the proposed algorithm shows clear advantages in both positioning accuracy and stability.
Multi-projection Plane InISAR 3D Reconstruction Method for Complex Moving Ship Targets
LI Ning, NIU Jinfa, WANG Weibin, HU Xingwang, WU Lin
Available online  , doi: 10.11999/JEIT251268
Abstract:
  Objective  Interferometric Inverse Synthetic Aperture Radar (InISAR) is a three-dimensional (3D) reconstruction technique for non-cooperative targets. However, the complex 3D rotational motion of ship targets causes unstable Doppler frequency variation. Inverse Synthetic Aperture Radar (ISAR) imaging also inevitably suffers from target overlap and occlusion. These factors make high-precision and complete 3D reconstruction under a single projection plane difficult. Therefore, a multi-projection plane InISAR 3D reconstruction method for complex moving ship targets based on point cloud fusion is proposed. The method supplements target 3D information through efficient and high-precision point cloud registration and fusion, thereby significantly improving 3D reconstruction quality.  Methods  This method fully exploits the advantage of multi-plane observation enabled by the severe motion of ship targets. The ship centerline is extracted, and the vertical rotation vector is estimated by Principal Component Analysis (PCA) to select the optimal imaging times corresponding to different Imaging Projection Planes (IPPs). ISAR imaging and InISAR 3D reconstruction are then completed. In addition, a point cloud fusion algorithm that combines Weighted Random Sample Consensus (RANSAC) and hierarchical Iterative Closest Point (ICP) is proposed. The random sampling process is optimized through a feature stability weighting strategy, which enables efficient extraction and matching of corresponding feature points in InISAR images and achieves high-precision point cloud fusion under multiple IPPs.  Results and Discussions  Experimental results show that the proposed method significantly improves reconstruction accuracy and target completeness. For simulated ship point-target data, Fig. 7 shows excellent results, with a significant reduction in reconstruction error. Signal-to-Noise Ratio (SNR) analysis shows that the quality of 3D fusion imaging improves steadily as the SNR increases from –10 dB to 10 dB, and robust fusion performance is maintained even under low-SNR conditions. For simulated destroyer Radar Cross Section (RCS) data, the method achieves strong registration performance. The detail recovery and structural integrity of the fused image are also significantly improved, effectively addressing the incomplete reconstruction of 3D information caused by scattering-point overlap and occlusion.  Conclusions  To address the low reconstruction accuracy and information loss caused by target rotation, overlap, and occlusion in traditional InISAR methods for 3D reconstruction of complex moving ship targets, a multi-IPP InISAR 3D reconstruction method based on point cloud fusion is proposed. The method uses a PCA-based optimal imaging time selection strategy. Weighted RANSAC and hierarchical ICP algorithms are then applied to achieve efficient and high-precision registration and fusion of InISAR point clouds under multiple IPPs, thereby producing high-quality 3D reconstruction results. Multi-scenario experiments are conducted by constructing both a ship model with ideal scattering points and an electromagnetic simulation RCS model with occlusion effects. The results verify the accuracy of the proposed method under ideal conditions and demonstrate its applicability in complex real-world scenarios.
Wireless Communication and Internet of Things
A Joint Source-Channel Coding Modulation Scheme for the Transmission of Gaussian Sources
LÜ Yaping, MA Xiao
Available online  , doi: 10.11999/JEIT251224
Abstract:
  Objective  The Separated Source-Channel Coding (SSCC) scheme has been proven to incur no performance loss when the source block length tends to infinity. However, SSCC usually requires a large buffer and causes long delay. It may also lead to error propagation when a single symbol error occurs in the communication channel. To alleviate these issues, Joint Source-Channel Coding (JSCC) schemes have been studied for Gaussian source transmission. In this paper, a Joint Source-Channel Coding Modulation (JSCCM) scheme is proposed for Gaussian sources. A Gaussian source reconstruction scheme and its reconstruction expression are also provided.  Methods  The Gaussian source sequence is quantized into an M-ary symbol sequence by a Lloyd-Max quantizer. For the M-ary quantized symbol sequence, a matching M-ary Fourier Transform Pair (FTP) code is constructed. The corresponding M-ary Pulse Amplitude Modulation (M-PAM) scheme is adopted for modulation. The modulated M-ary symbol sequence is transmitted using Block Markov Superposition Transmission (BMST), forming a BMST-FTP code. In addition, a Geometric Shaping (GS) scheme is proposed to obtain shaping gain. In the proposed source reconstruction scheme, the system output is the weighted average of the representative elements of the Lloyd-Max quantizer, rather than a single representative element.  Results and Discussions  Simulations are conducted over Additive White Gaussian Noise (AWGN) channels with M-PAM modulation and BMST-FTP codes over Galois Field (GF) orders 3 and 5, denoted GF(3) and GF(5). For FTP codes with random mapping, the Word Error Rate (WER) approaches the Union Bound (UB) at high Signal-to-Noise Ratio (SNR). Similarly, FTP codes with m repeated transmissions show WER performance close to the corresponding UBs. The WER performance of BMST-FTP codes with memory m also approaches the UBs in the high SNR region (Fig. 6). In terms of Symbol Error Rate (SER), the GF(3) BMST-FTP code outperforms the GF(5) BMST-FTP code (Fig. 7(a)). For the GF(5) BMST-FTP code, GS provides an SER performance gain of approximately 0.3 dB (Fig. 8(a)). In terms of distortion performance, the GF(3) BMST-FTP code performs better in the low SNR region, whereas the GF(5) BMST-FTP code performs better in the high SNR region (Fig. 7(b)). Compared with other work, the GF(3) BMST-FTP code with m = 1 achieves similar performance, whereas the GF(5) BMST-FTP code with m = 1 achieves better performance (Fig. 7(b)).  Conclusions  This work proposes a JSCCM scheme for Gaussian source transmission. In the proposed scheme, two types of BMST-FTP codes are constructed. Each code is matched with a corresponding Lloyd-Max quantizer and M-PAM modulator. A Gaussian source reconstruction scheme and its reconstruction expression are also provided. Simulation results show that an appropriate transmission scheme can be selected according to the target performance. The proposed GS scheme provides an SER gain of approximately 0.3 dB and improves distortion performance in the waterfall region.
A Two-layer Closed-loop Cooperative Resource Allocation Framework for Improving QoS of MEC Network Slicing
XU Juntao, FAN Xinggang, XU Changfu, SHEN Minyang, LIANG Yuzhu, WANG Tian
Available online  , doi: 10.11999/JEIT260156
Abstract:
  Objective  Driven by 5G/6G networks, Multi-access Edge Computing (MEC) environments face major challenges in ensuring Quality of Service (QoS) for heterogeneous network slices while improving resource utilization. Existing resource allocation methods often lack dynamic adaptability in heterogeneous settings. They also fail to jointly optimize caching, bandwidth, and computing resources, which reduces resource utilization and service success rates. This paper addresses these limitations by proposing a robust framework for joint multi-dimensional resource optimization under strict QoS constraints. The framework provides a tailored solution for heterogeneous MEC environments.  Methods  The network slicing resource allocation problem is first formally defined. Its NP-hardness is proved by reducing the NP-hard multidimensional 0-1 knapsack problem to this problem. To solve this complex optimization problem, a cache-aware two-layer closed-loop cooperative framework, termed QCache, is proposed. The framework uses a synergistic “generate-evaluate-feedback” loop. In the upper Global Exploration Layer, a hybrid heuristic algorithm is designed by combining a population evolution strategy, including selection, crossover, and mutation, with a particle update mechanism guided by historical individual and global best positions. This layer broadly explores the solution space and generates high-quality candidate resource allocation schemes under complex constraints. Adaptive parameter adjustment and elite retention strategies are used to avoid local optima. In the lower Multi-Dimensional Weight Evaluation Layer, a quantitative assessment model is constructed. This model converts low-latency and high-bandwidth service demands into explicit QoS constraints by normalizing key performance indicators, including delay and rate, and by dynamically assigning weights through the entropy weight method. The weighted score reflects different slice priorities. The evaluated score (SCORE) from this layer is fed back to the upper layer as the fitness value, guiding the iterative evolution of candidate solutions until convergence.  Results and Discussions  Extensive simulations are conducted to validate the effectiveness of QCache against several baseline methods, including Genetic Algorithm (GA), PSO-Leader, GraphSAGE, and the No-Consideration-of-Service-Quality (NCSQ) scheme. Under identical resource and user demand scenarios, the overall comparison (Fig. 3) shows that QCache achieves the highest resource utilization rate of 80.30% and the best average service score of 1.109. Compared with the baseline methods, QCache improves resource utilization by 2.29% to 24.50% and increases user service scores by 4.13% to 59.34%. Experiments with varying total cache resources (Fig. 4) show that QCache maintains superior performance across different cache states. It improves resource utilization by 2.43% to 27.53% and service scores by 10.81% to 119.56%, confirming its cache-aware adaptability. Tests with increasing user numbers (Fig. 5) show that QCache scales effectively, achieving up to 85.83% resource utilization and a service score of 3.06. These results demonstrate its ability to handle dense access scenarios. Experiments with time-varying user demands (Fig. 6) further confirm the dynamic robustness of the framework. In these tests, QCache achieves average improvements of 9.20% in resource utilization and 23.45% in service score over the baselines.  Conclusions  This paper studies the NP-hard resource allocation problem in dynamic MEC environments with heterogeneous network slices. The proposed cache-aware two-layer closed-loop cooperative framework, QCache, jointly optimizes caching, bandwidth, and computing resources under explicit QoS constraints. The upper-layer hybrid heuristic provides strong global search capability. The lower-layer multi-dimensional weight model supports accurate QoS quantification and dynamic feedback. Comprehensive experimental results show that QCache outperforms existing methods in both resource utilization efficiency and user QoS satisfaction. Future work will explore reinforcement learning and traffic prediction mechanisms to further improve the response of the framework to bursty traffic and anomalous demands. This may support more intelligent and autonomous MEC network slice resource management.
Image and Intelligent Information Processing
A Cross-Precision Motion Compensation Technique for Security Surveillance Video Coding
JIANG Wei, MA Wei, LU Jinghui, ZHANG Yue, ZHANG Yundong
Available online  , doi: 10.11999/JEIT251301
Abstract:
  Objective  High-altitude dome cameras are widely used in modern security surveillance. They are often deployed at critical locations, such as bridges and tower tops, where they are vulnerable to external interference. Such interference can cause jitter, blur, and distortion in captured videos, creating major challenges for video coding. In video compression, high-precision motion compensation is essential for improving coding efficiency. However, the existing Ultimate Motion Vector Expression (UMVE) technique has limited motion-vector precision and insufficient flexibility in adaptive adjustment. High-precision motion compensation tools, such as Registration Coding Mode (RCM) and Affine Motion Compensation Prediction (AFFINE), can improve compensation accuracy, but they require high computational complexity and hardware cost. These limitations make it difficult to meet the requirements for coding efficiency, power consumption, and real-time processing in high-altitude surveillance scenarios. Therefore, this study aims to design an optimized UMVE scheme that integrates high-precision motion compensation, low computational complexity, and scene adaptability to improve coding efficiency while balancing resource consumption.  Methods  This study proposes UMVE_CPMC, an Ultimate Motion Vector Expression technique supporting Cross-Precision Motion Compensation. The proposed method improves motion compensation accuracy by constructing an extended Up-Precision Motion Vector (UPMV), expressed as UPMV = BaseMV + MMV(p, angle). Here, Base Motion Vector (BaseMV) denotes the base vector obtained by the existing UMVE method, and Micro-Motion Vector (MMV) denotes the fine-adjustment vector defined by a specific precision p and angle. Incremental candidates are provided only at the 1/8 precision level to balance computational complexity and compression efficiency. For step-size adaptive adjustment, a six-mode improved scheme is proposed. It covers enhanced UMVE, conventional UMVE, and four precision-improved modes, allowing the encoder to switch flexibly according to scene characteristics. The average image gradient is used as an objective evaluation index. Test scenes are divided into Class A, representing high-clarity motion scenes, and Class B, representing low-clarity scenes. Different coding configurations, sequences, and parameters are used to compare coding gains and computational efficiency under different modes.  Results and Discussions  Experiments show that UMVE_CPMC improves performance under different scenes and modes. In Class A high-clarity motion scenes, with both the adaptive strategy and RCM disabled, the average gains of the Y, U, and V components in Fusion Mode 1 are –2.912%, –1.656%, and –1.654%, respectively. The average coding time is reduced to 94.55% of the baseline. In Independent Mode 1, the average Y-component gain reaches –2.925%, and the coding time is reduced to 91.91% of the baseline. Compared with conventional UMVE, when CPMC Independent Mode 1 is enabled with RCM and other tools working together, the gain improves from –0.276% to –1.310%, indicating higher cost effectiveness. In Class B low-clarity scenes, adaptive adjustment significantly reduces the losses of coding gain in Fusion Mode 1 and Fusion Mode 0. The average losses of coding gain are limited to 0.071% and 0.108%, respectively, which maintains the original coding gain. In multi-scene tests with RCM and AFFINE disabled, 9 of 10 test sequences in adaptive Fusion Mode 1 show positive gains. The Y-component gain reaches –10.691% for the yuxuedaolu sequence and –11.400% for the BQTerrace sequence. When all existing coding tools are enabled, the Y-component gains of the dianjing, yuxuedaolu, and BQTerrace sequences reach –1.29%, –2.05%, and –1.21%, respectively. The coding time is reduced to 94%~96% of the baseline. In addition, correlation analysis shows a clear positive relationship between the average image gradient and the coding gain. Images with a high average gradient, corresponding to high clarity, gain more from UMVE_CPMC, whereas images with a low average gradient, corresponding to low clarity, benefit little. Principle analysis shows that pixel changes in low-clarity images are smooth, so high-precision interpolation cannot generate effective new pixel values. The compensation effect is therefore limited. The performance differences among modes are consistent with their computational complexity. The fusion mode balances gain and stability, whereas the independent mode further reduces computation. The six step-size adaptive modes can meet the real-time and precision requirements of different scenes.  Conclusions  The proposed UMVE_CPMC technique integrates Cross-Precision Motion Compensation with the UMVE algorithm. It addresses the limited precision of conventional UMVE and the high computational complexity of high-precision motion compensation tools. It also achieves a favorable balance among coding efficiency, computational complexity, and scene adaptability. In Class A high-clarity motion scenes, UMVE_CPMC achieves notable coding gains. The gain exceeds 10% for some sequences when other high-precision motion compensation tools are disabled and reaches 1%~2% when used with other tools. In Class B low-clarity scenes, the original coding gain is maintained through a frame-level adaptive adjustment interface. In addition, the fusion mode does not increase hardware complexity, whereas the independent mode significantly reduces coding time. These features make the proposed method suitable for encoder designs with limited resources or simplified requirements. UMVE_CPMC provides an effective approach for improving the coding efficiency of high-altitude dome camera videos affected by jitter and blur. It also enriches the video coding toolset and provides practical guidance for optimizing video coding technologies in security surveillance. Future work will further optimize the adaptive strategy, explore integration with other advanced coding tools, develop scenario-specific coding schemes, and improve performance in complex scenes.
Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement
ZHANG Chunxiang, ZHANG Huibin, GAO Xueyao
Available online  , doi: 10.11999/JEIT260061
Abstract:
  Objective  Biomedical Word Sense Disambiguation (WSD) is important for biomedical text mining and clinical data analysis. However, biomedical terms are susceptible to semantic noise, fine-grained semantic classes are difficult to distinguish, and model generalization is limited in small-sample settings. A three-branch parallel WSD framework with contrastive learning is proposed to address these problems. Electra, mDeBERTa, and Flan-T5 (FT5) are integrated to obtain complementary semantic features. Chi-square attention, a Focal Loss and Margin Loss hybrid loss, and hard sample mining are further incorporated to improve feature discrimination and model robustness.  Methods  The proposed framework uses Electra, mDeBERTa, and FT5 to extract complementary semantic features from biomedical terms. A chi-square attention module is designed to identify representative biomedical terms and guide token-level attention allocation. Focal Loss and Margin Loss are combined to improve the learning of hard samples and class boundaries under class imbalance. A two-stage hard sample mining strategy is developed by jointly considering training loss and prediction uncertainty. In addition, a constrained contrastive learning mechanism based on core disambiguation tokens is introduced. Random cropping is used to generate semantically equivalent augmented views, and the NT-Xent loss is applied to optimize the representation space.  Results and Discussions  Experiments are conducted on the MSH WSD dataset, which contains 203 ambiguous biomedical terms. The proposed model achieves an accuracy of 95.27%, a precision of 90.33%, a recall of 85.49%, and an F1 score of 87.84%. It outperforms the Neural Concept Embeddings model, which achieves an accuracy of 94.34%, by 0.93 percentage points. Ablation experiments show that the successive addition of Word2Vec, FT5, hard sample mining, chi-square attention, the hybrid loss, and contrastive learning improves model performance. Among the evaluated contrastive learning settings, NT-Xent with a temperature of 0.10 and a contrastive weight of 1.00 achieves the best performance. The proposed hard sample mining strategy, which combines training loss and prediction uncertainty, also outperforms alternative sample selection methods.  Conclusions  A three-branch parallel contrastive learning framework is proposed for biomedical WSD. Electra, mDeBERTa, and FT5 are integrated to capture complementary semantic features. Chi-square attention is used to strengthen representative feature extraction, whereas the Focal Loss and Margin Loss hybrid loss improves learning of hard samples and class boundaries. Two-stage hard sample mining and constrained contrastive learning further enhance the discrimination of fine-grained semantic classes. Experimental results demonstrate that the proposed method improves biomedical WSD performance and reduces confusion among semantically similar biomedical concepts. The framework is evaluated on English biomedical texts from the MSH WSD dataset and can be further extended to multilingual biomedical corpora and domain-specific knowledge graphs.
Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism
ZHANG Xindi, CHEN Hui, ZHANG Hongyun, LIAN Feng, ZHANG Guanghua, YIN Zhipeng
Available online  , doi: 10.11999/JEIT260045
Abstract:
  Objective  Multi-target tracking requires sensor management to adapt the observation process to clutter, missed detections, target-number variations, and target-motion changes. Conventional methods typically search over a finite set of sensor actions, which increases computational cost and limits control resolution. Furthermore, reward functions constructed from multiple single-target metrics may not adequately characterize the joint multi-target posterior. To address these limitations, a continuous-action sensor management method that integrates Twin-Delayed Deep Deterministic Policy Gradient (TD3) with the Labeled Multi-Bernoulli (LMB) filter is proposed to optimize the mobile-sensor heading angle according to the multi-target belief state.  Methods  The LMB posterior, including target existence probabilities and state densities, is used to construct the belief state. At each filtering step, the mobile sensor selects a continuous heading angle that determines the sensor-target geometry, detection probability, and LMB update. Predicted target states and candidate heading actions are used to generate pseudo measurements and obtain pseudo-updated LMB densities. The Cauchy-Schwarz (CS) divergence between the predicted and pseudo-updated LMB densities is adopted to construct an information-gain reward. TD3 employs twin critics, target policy smoothing, and delayed policy updates to reduce value-estimation bias. Random control, Policy Gradient (PG)-based sensor management, CS divergence-based sensor management, and Deep Deterministic Policy Gradient (DDPG)-based sensor management are used for comparison.  Results and Discussions  DDPG-LMB and TD3-LMB produce smoother sensor trajectories than the discrete-action methods (Fig. 2). TD3-LMB achieves the highest or near-highest detection probabilities for most targets (Fig. 3) and yields larger CS divergence values during most time steps, while random control consistently produces lower values (Fig. 4). TD3-LMB also achieves the lowest overall Optimal Subpattern Assignment (OSPA) distance in the evaluated scenario, while DDPG-LMB generally outperforms the discrete baseline methods (Fig. 5). These results demonstrate that continuous heading-angle control improves observation quality and enhances the overall tracking performance of the LMB filter.  Conclusions  A TD3-based continuous-action sensor management framework for the LMB filter is presented. Candidate heading actions are evaluated using pseudo-updated LMB densities and CS divergence, directly associating action selection with the joint multi-target posterior. Simulation results demonstrate smoother sensor trajectories, higher detection probabilities for most targets, greater information gain, and lower OSPA distances in the evaluated scenario. Future work will consider higher-dimensional action spaces and cooperative multi-sensor management.
Resilience-Aware Cooperative Mission Planning Algorithm for Multiple UAV Systems in Complex Dynamic Environments
ZHAO Xuejian, XIE Lulu, WANG Enliang
Available online  , doi: 10.11999/JEIT260138
Abstract:
  Objective  This paper addresses the strongly coupled problem of task allocation and route planning in cooperative task and route planning for multiple UAV systems operating in complex dynamic environments, where dynamic task arrivals, UAV failures, no-fly-zone constraints, and link quality degradation occur simultaneously.  Methods  A Resilience-Aware Hybrid Swarm Optimization (RAHSO) algorithm is proposed. First, an integrated task-route planning model is established by jointly considering task value, route cost, energy consumption, interference penalties, time-window constraints, platform capability constraints, conflict resolution, and link quality within a unified optimization framework. High-quality initial solutions are generated through clustering-based and genetic initialization. A hybrid optimization framework that integrates the Dung Beetle Optimizer (DBO), Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Variable Neighborhood Search (VNS) is then employed to perform global exploration and local refinement. In addition, Tarjan-based deadlock detection and repair are incorporated to guarantee feasible task assignments. Finally, an event-driven Proximal Policy Optimization (PPO) online replanning module is designed to rapidly update affected task subsets in response to emergent tasks, UAV failures, and network topology changes.  Results and Discussions  Comparative and ablation experiments are conducted under static, large-scale, dynamic-event, and interruption scenarios. The results demonstrate that the proposed method consistently outperforms representative baseline algorithms in task completion rate, accumulated task value, average energy consumption, recovery time, and resilience index while maintaining satisfactory online replanning latency.  Conclusions  The proposed method provides an effective solution for resilient cooperative task and route planning for multiple UAV systems operating in complex dynamic environments.
Infrared Small Target Detection Enhanced by Multi-Dimensional Fusion Attention
LI Weixing, WANG Shuai, CHEN Huaiyu, SHENG Weidong
Available online  , doi: 10.11999/JEIT260040
Abstract:
  Objective  Infrared imaging offers advantages including long operating range, wide coverage, high concealment, and all-weather operation, making it suitable for aerospace surveillance, maritime emergency rescue, forest fire monitoring, and remote sensing. Infrared sensors mounted on satellites and aircraft typically acquire long-range images in which targets occupy only a few pixels and lack discriminative texture and shape features. Moreover, weak thermal radiation signals are easily overwhelmed by background clutter. Existing infrared small target detection networks face two major challenges. First, repeated downsampling used to enlarge the receptive field causes small-target features to disappear in deep networks. Second, small-target features become increasingly diffused during deep feature extraction. Therefore, achieving accurate and reliable infrared small target detection under long-range imaging and complex background conditions remains an active research topic.  Methods  A Multi-Dimensional Fusion Attention Module (MFAM) based on channel-spatial attention is proposed to address feature diffusion caused by the small size of infrared targets. The proposed module captures feature dependencies across the channel, height, and width dimensions. Feature fusion and cross-dimensional interaction are jointly exploited to suppress the diffusion of small-target features in deep networks and strengthen the representation of weak infrared targets. Channel attention and spatial attention are applied in parallel directly to the input feature map, enabling feature extraction and information interaction in the channel and spatial domains, respectively. In the channel domain, compressed features are processed using a Multi-Layer Perceptron (MLP) to model inter-channel dependencies. In the spatial domain, Global Average Pooling (GAP) and Global Maximum Pooling (GMP) encode spatial information along the height and width dimensions. The outputs of the two branches are fused through a Sigmoid activation function to generate the final attention map. Unlike conventional hybrid attention mechanisms that connect channel attention and spatial attention sequentially, the proposed parallel architecture performs refined encoding of the channel, height, and width dimensions directly from the original input feature map. This design preserves both shallow spatial details and deep contextual semantics while reducing information loss during feature propagation. Owing to its plug-and-play design, MFAM can be seamlessly integrated into backbone networks such as ResNet and DNA-Net without introducing complex additional structures, demonstrating excellent compatibility.  Results and Discussions  The proposed method is evaluated on the publicly available NUDT-SIRST dataset. After MFAM is integrated into the baseline DNA-Net, the Intersection over Union (IoU), detection rate (Pd), and false alarm rate (Fa) reach 87.34%, 98.72%, and 3.22×10–6, respectively. Compared with CBAM, BAM, GAM, CA, and TA, MFAM improves IoU by 0.40%, 1.68%, 2.46%, 2.11%, and 1.61%, respectively (Table 1). CBAM applies spatial attention after channel attention, which increases the risk of information loss during feature propagation. CA and BAM rely solely on GAP for feature encoding and therefore fail to exploit the complementary information provided by GMP. TA models cross-dimensional dependencies through rotation operations but cannot achieve simultaneous interaction among the channel, height, and width dimensions. By jointly integrating channel and spatial attention, MFAM enables effective cross-domain information interaction. The combined use of GAP and GMP further strengthens contextual and local feature representation, resulting in more accurate infrared small target detection. MFAM is also incorporated into ALCNet and AMFU-Net, improving IoU by 0.43% and 0.28%, respectively, compared with CBAM (Table 2). Ablation experiments further demonstrate the effectiveness of the proposed design. Compared with the serial attention architecture, the parallel fusion strategy improves IoU and Pd by 0.37% and 0.39%, respectively, while reducing Fa by 1.41×10–6. These improvements result from applying channel attention and spatial attention directly to the original input feature map, thereby alleviating the diffusion of small-target features in deep networks. To verify inference performance on edge devices, the proposed method is deployed on an FPGA-GPU heterogeneous platform based on an FPGA and an NVIDIA Jetson AGX Xavier. Experimental results demonstrate successful edge deployment, with an average inference latency of 46.7 ms per 256×256 image (Fig. 8), satisfying real-time processing requirements.  Conclusions  A MFAM is proposed for infrared small target detection. The module effectively aggregates channel and spatial salient features while enabling adaptive cross-dimensional interaction, thereby improving the preservation of small-target features in deep networks and enhancing detection robustness. Owing to its lightweight plug-and-play design, MFAM can be readily integrated into existing detection networks. Experimental results demonstrate superior performance over existing methods in terms of IoU, Pd, and Fa. Furthermore, a lightweight intelligent processing unit based on an FPGA-GPU heterogeneous platform is developed to enable real-time deployment of the proposed algorithm on edge devices, demonstrating its practicality for engineering applications.
An Anomalous Traffic Detection Method Integrating Flow Data Compression and Self-supervised Graph Learning
XIA Jiqiang, ZHAO Jianjin, WANG Zihao, TIAN Le, HU Yuxiang, LI Menglong
Available online  , doi: 10.11999/JEIT260118
Abstract:
  Objective  As network traffic volumes continue to grow and attack methods become increasingly sophisticated, efficient and intelligent anomalous traffic detection is essential for protecting critical information infrastructure. However, existing detection methods still face substantial challenges in large-scale network environments. On one hand, analyzing raw packet sequences and using deep learning-based end-to-end models incur considerable computational and storage overhead, making them difficult to deploy in line-rate processing scenarios. On the other hand, flow records are usually treated as independent samples, while the topological structure and contextual information of inter-host communications are often ignored. This limitation makes it difficult to detect distributed and correlated threats from a global perspective. In addition, supervised learning methods rely heavily on large amounts of labeled data, which are difficult to obtain in practical deployments and limit generalization to unknown threats. Therefore, an anomalous traffic detection method that supports efficient flow feature extraction under limited resources and enables accurate detection without labeled data is needed.  Methods  During training, the original communication graph and augmented negative samples are simultaneously input into a graph encoder to learn edge embeddings. The resulting embeddings are then fed into a discriminator. Mutual information scores are estimated by contrasting the edge embeddings of positive and negative samples with a global graph summary. The training objective is to maximize the scores of positive samples and minimize those of negative samples. Through this self-supervised optimization, the encoder parameters are refined to improve the discriminative capability of the edge embeddings. The process is iterated using gradient descent until convergence. After training, the encoder parameters are fixed, and the resulting edge embeddings are used for downstream anomalous traffic detection. During inference, traffic is converted into a communication graph by the feature extractor and then fed into the trained graph encoder to generate the corresponding edge embeddings. A lightweight classifier takes these embeddings as input to perform end-to-end anomalous traffic detection and output the final classification results.  Results and Discussions  Comprehensive experiments are conducted on four public datasets, namely CAIDA, CIC-IDS2018, UNSW-NB15, and TON-IoT. For feature extraction, under identical memory configurations, the Average Relative Error (ARE) and per-flow Weighted Mean Relative Error (WMRE) of counter features measured by MFSketch-OP are reduced by 31.5% and 31.0%, respectively, compared with the baseline MFSketch with a fixed structure. For bitmap features, the corresponding reductions are 36.1% and 34.9%, respectively (Fig. 4). High throughput is maintained across datasets with different degrees of traffic skewness, with an average throughput of approximately 12 Mpps achieved on CAIDA (Fig. 4(c)). For detection accuracy, when combined with Principal Component Analysis (PCA), Histogram-Based Outlier Score (HBOS), or Isolation Forest (IF), SketchGNN consistently achieves an accuracy of at least 95.2%, a macro-F1 score of at least 90.1%, and a weighted-F1 score of at least 96.7% on CIC-IDS2018 and UNSW-NB15. These results generally outperform the baseline methods and show more stable performance across datasets (Figs. 5 and 6). For detection efficiency, HBOS provides high and stable throughput among the three classifiers (Fig. 7(a)). The end-to-end packet-level equivalent throughput of SketchGNN reaches 640 kpps, approximately 17 times that of Kitsune (37 kpps), and is comparable in magnitude to that of Whisper accelerated by the Data Plane Development Kit (DPDK) (1.3 Mpps) (Fig. 7(b)). In addition, the performance variation across different datasets remains within 3%, indicating robust generalization to normal traffic fluctuations and diverse flow-level anomalous behaviors.  Conclusions  To address the high overhead of feature extraction, insufficient use of traffic context, and strong dependence on labeled data in existing anomalous traffic detection methods, SketchGNN, an anomalous traffic detection framework integrating flow data compression with self-supervised graph learning, is proposed. A dynamically configurable sketch, MFSketch, is used to efficiently extract and accurately measure diverse flow features under limited resource constraints. A self-supervised graph neural network is then used to model host communication graphs and learn traffic representations, enabling efficient anomalous traffic detection without labeled data. Experimental results show that MFSketch dynamically optimizes its data structure according to traffic distribution and provides high-throughput and high-precision feature inputs for downstream detection. The edge embeddings generated through self-supervised graph learning achieve higher detection accuracy than the baseline methods when combined with different unsupervised classifiers. In future work, hybrid detection mechanisms that combine Deep Packet Inspection (DPI) with programmable data planes will be explored to further improve the detection of application-layer anomalous traffic.
Multi-Agent Deep Reinforcement Learning Strategy for Multi-Spacecraft Long-Distance Orbital Pursuit-Evasion Games
DI Peng, YIN Zengshan, LIN Zheng, YAO Ye
Available online  , doi: 10.11999/JEIT251384
Abstract:
This paper presents a novel research scenario for the multi-spacecraft Orbital Pursuit-Evasion Game (OPEG), which has not yet been systematically studied. To improve spacecraft decision-making and enable robust policies in complex multi-agent games, a Multi-Agent Deep Reinforcement Learning (MADRL) algorithm based on a Progressive Adversarial Training Framework (PATF) is proposed to solve the game policies of each spacecraft. Two numerical cases with different orbital characteristics and four simulation setups are designed for verification. Behavioral deviation analysis is also conducted to evaluate policy robustness. The effects of different orbital characteristics, simulation setups, and behavioral deviations on spacecraft game policies are analyzed. Simulation results show that the proposed method enables each spacecraft to develop effective game policies that satisfy all prescribed constraints and show good robustness.  Objective  As the space environment becomes increasingly complex, space security has become a major research topic. A large amount of space debris and failed spacecraft pose serious threats to high-value spacecraft in orbit. Therefore, the OPEG for non-cooperative target spacecraft has attracted considerable attention. Existing studies mainly focus on two-spacecraft OPEGs, whereas multi-spacecraft OPEGs remain less explored. When more than two participants are included, a zero-sum game formulation is no longer feasible, and the problem becomes difficult to solve using traditional methods. Moreover, existing studies often ignore engineering dynamic constraints and simplify the dynamics or define the problem in a two-dimensional scene, which may introduce considerable errors. To address these limitations, this paper proposes a novel multi-spacecraft OPEG scenario. The aim is to investigate the application of MADRL to solving the approximate steady-state policies of each spacecraft in a long-distance multi-spacecraft OPEG. This study highlights the advantages of MADRL in multi-spacecraft OPEGs and provides a feasible approach for future autonomous multi-spacecraft game decision-making.  Methods  A Multi-Agent Proximal Policy Optimization (MAPPO) algorithm based on PATF is used to solve the approximate steady-state policy for each spacecraft in the multi-spacecraft OPEG. First, a multi-constrained multi-spacecraft OPEG model is established based on practical engineering constraints, and the problem is formulated as a Partially Observable Stochastic Game (POSG). Second, to improve the decision-making ability of agents in complex multi-agent game environments and develop more robust game policies, a novel PATF is proposed. Different reward functions are designed for the specific missions of each spacecraft. Finally, two numerical cases with different orbital characteristics are designed. Four simulation setups are then used for simulation and behavioral deviation analysis.  Results and Discussions  The proposed PATF-based MAPPO algorithm is compared with the original MAPPO algorithm (Fig. 3). The results show that the proposed method learns effective policies more rapidly, reduces ineffective exploration, and achieves a higher final convergence reward with smaller fluctuations in the reward curve. These results also show that PATF can significantly improve the decision-making ability of agents and help them develop robust policies more effectively. Simulation verification is conducted using two numerical cases under four different setups (Figs. 4, 5, 6, and 7). The simulation results (Tables 3 and 4) show that the proposed method performs well in both cases. The results further show that the pursuer is more likely to be intercepted when the pursuer and interceptor are in the same orbital plane. When the interceptor and the target are not in the same orbital plane, the interception mission becomes relatively easier. This paper also analyzes behavioral deviations on both sides of the game by adding control noise. The simulation results (Tables 5 and 6) show that both sides adopt relatively conservative policies to counter control noise. The game policy obtained by the proposed method is an approximate steady-state policy. Behavioral deviations reduce the deviating side’s payoff and increase the opponent’s payoff, while the game policy maintains good robustness.  Conclusions  The proposed method can be effectively applied to long-distance OPEGs with multiple spacecraft in non-coplanar elliptical orbits, enabling each spacecraft to develop effective game policies. PATF improves spacecraft decision-making in complex multi-spacecraft dynamic systems, and robust control policies are developed by both the pursuer and the interceptors. The results also demonstrate the accuracy and effectiveness of the reward function design. Based on two numerical cases and simulation results under different setups, the effects of different orbital characteristics on the policies of both sides are analyzed. When the interceptors have different maximum thrusts, the decision-making of each spacecraft changes accordingly. Behavioral deviation analysis shows that the game policies of each spacecraft have good robustness. When one side’s behavior deviates, the approximate steady-state policy balance changes, which reduces its own payoff and increases the opponent’s payoff. The research scenario proposed in this paper expands the scope of existing studies on multi-spacecraft game problems.
A Hierarchical Cross-layer Closed-loop Learning Framework andCoordination Mechanism for Complex Multi-agent Systems
ZHANG Long, HUANG wenbo, LEI Zhen, FENG Xuanming, WANG Ying
Available online  , doi: 10.11999/JEIT260143
Abstract:
Complex Multi-Agent Systems (MAS) in dynamic and uncertain environments face challenges in unified modeling, adaptive coordination, and interpretable effectiveness evaluation. Existing methods usually address individual decision-making, inter-agent coordination, and high-level policy evolution separately. This separation leads to fragmented decision chains and weak cross-layer coupling. It also makes it difficult to explain how local learning gains are transformed into global effectiveness improvements under mission variation, observation disturbance, and structural damage. To address this issue, a Hierarchical Cross-layer Closed-loop Learning (HCCL) framework is proposed. The framework couples individual autonomy, system-level coordination, and system-of-systems learning to build a computable path from local policy optimization to overall effectiveness enhancement.   Methods   HCCL adopts a unified three-layer architecture. At the individual autonomy layer, each agent is modeled as a Partially Observable Markov Decision Process (POMDP) to describe decision-making under partial observability. At the system-level coordination layer, multi-agent coordination is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and represented by a dynamic directed weighted coordination graph. A Graph Neural Network (GNN) is used to encode interaction dependencies, structural coupling, and joint value information. At the system-of-systems learning layer, a Meta-Decentralized Partially Observable Markov Decision Process (Meta-Dec-POMDP) is established to describe task-context adaptation and rule evolution. A cross-layer closed-loop mechanism is further designed. In the bottom-up behavior induction pathway, local state and capability features are aggregated into graph-level structural representations and supplied to the upper rule-learning process. In the top-down rule-shaping pathway, learned high-level rules are converted into control parameters and fed back to lower layers to regulate local policies and coordination relationships. Simulations are conducted under baseline, mission-variation, observation-disturbance, and structural-damage scenarios. The full HCCL model is compared with a non-closed-loop model and an upward-induction-only model. Interface ablation studies are also performed to analyze the contributions of cross-layer feature reporting, structural induction, and rule shaping.   Results and Discussions   The full HCCL model consistently outperforms the comparison models and ablated variants. In the baseline scenario, it achieves a task success rate of 88.6% and a comprehensive system effectiveness of 0.842. Under mission variation, it reduces the adaptation process to 16±2 rounds. Under structural damage, it achieves a recovery rate of 81.4% and restores coordination-structure stability to 0.742 within 20 steps. These results indicate that HCCL improves task performance, adaptation speed, and structural recovery. Ablation results show that removing any cross-layer interface reduces performance, while removing the top-down rule-shaping pathway causes the largest loss. This result indicates that upward structural perception alone is insufficient for sustained system-level improvement. The effectiveness gain mainly arises from closed-loop coupling between bottom-up behavior induction and top-down rule shaping, rather than from simple hierarchical stacking.   Conclusions   The HCCL framework is proposed for complex MAS by integrating POMDP-based individual autonomy modeling, Dec-POMDP- and graph-based coordination modeling, and Meta-Dec-POMDP-based rule evolution. Through bottom-up behavior induction and top-down rule shaping, HCCL provides a computable and interpretable path from local learning to overall effectiveness enhancement. Experimental results verify its advantages in task completion, adaptation, recovery, and coordination stability under multiple disturbances. Future work will focus on larger-scale heterogeneous systems, communication-constrained networking, online continual adaptation, and data-driven evaluation in realistic environments.
Special Topic on Advanced Technologies of Optoelectronic Information
Finite-time Adaptive Sliding Mode Control of Servo Motors Considering Frictional Nonlinearity and Unknown Loads
ZHANG Tianyu, GUO Qinxia, YANG Tingkai, GUO Xiangji, MING Ming
Available online  , doi: 10.11999/JEIT250521
Abstract:
  Objective  Ultra-fast laser processing with an infinite field of view requires servo motor systems with superior tracking accuracy and robustness. However, such systems are highly nonlinear and affected by coupled unknown load disturbances and complex friction, which constrain the performance of conventional controllers. Although Sliding Mode Control (SMC) exhibits inherent robustness, traditional SMC and observer designs cannot achieve accurate finite-time disturbance compensation under strong nonlinearities, thus limiting high-speed and high-precision trajectory tracking. To address this limitation, a novel finite-time adaptive SMC approach is proposed to ensure rapid and precise angular position tracking within a finite time, satisfying the stringent synchronization requirements of advanced laser processing systems.  Methods  A novel control strategy is developed by integrating an adaptive disturbance observer fused with a Radial Basis Function Neural Network (RBFNN) and finite-time SMC. First, the unknown load disturbance and complex frictional nonlinear dynamics are combined into a unified "lumped disturbance" term, improving model generality and the ability to represent real operating conditions. Second, a finite-time adaptive disturbance observer is constructed to estimate this lumped disturbance. The observer utilizes the universal approximation capability of the RBFNN to learn and approximate the dynamic characteristics of unknown disturbances online. Simultaneously, a finite-time adaptive law based on the error norm is introduced to update the neural network weights in real time, ensuring rapid and accurate finite-time estimation of the lumped disturbance while reducing dependence on precise model parameters. Based on this design, a finite-time SMC is developed. The controller uses the observer’s disturbance estimation as a feedforward compensation term, incorporates a carefully formulated finite-time sliding surface and equivalent control law, and introduces a saturation function to suppress control input chattering. A suitable Lyapunov function is then constructed, and the finite-time stability theory is rigorously applied to prove the practical finite-time convergence of both the adaptive observer and the closed-loop control system, guaranteeing that the system tracking error converges to a bounded neighborhood near the origin within finite time.  Results and Discussions  To verify the effectiveness and superiority of the proposed control strategy, a typical Permanent Magnet Synchronous Motor (PMSM) servo system model is constructed in the MATLAB environment, and a simulation scenario with desired trajectories of varying frequencies is established. The proposed method is comprehensively compared with the widely used Proportional–Integral (PI) control and the advanced method reported in reference[7]. Simulation results demonstrate the following: 1. Tracking performance: Under various reference trajectories, the proposed controller enables the system to accurately follow the target trajectory with a tracking error substantially smaller than that of the PI controller. Compared with the method in reference[7], it achieves smoother responses and smaller residual errors, effectively eliminating the chattering observed in some operating conditions of the latter. 2 Disturbance rejection and robustness: The adaptive disturbance observer based on the RBFNN rapidly and effectively learns and compensates for the lumped disturbance composed of unknown load variations and frictional nonlinearities. Even in the presence of these disturbances, the proposed controller maintains high-precision trajectory tracking, demonstrating strong disturbance rejection and robustness to system parameter variations. 3. Control input characteristics: Compared with the reference methods, the control signal of the proposed approach quickly stabilizes after the initial transient phase, effectively suppressing chattering caused by high-frequency switching. The amplitude range of the control input remains reasonable, facilitating practical actuator implementation. 4. Comprehensive evaluation: Based on multiple error performance indices, including Integral Squared Error (ISE), Integral Absolute Error (IAE), Time-weighted Integral Absolute Error (ITAE), and Time-weighted Integral Squared Error (ITSE), the proposed controller consistently outperforms both PI control and the method in reference[7]. It demonstrates comprehensive advantages in suppressing transient errors rapidly and reducing overall error accumulation. The method also improves steady-state accuracy and achieves a balanced response speed with effective noise attenuation. 5. Observer performance: The RBFNN weight norm estimation converges rapidly and stabilizes at a low level after initial adaptation, confirming the effectiveness of the proposed adaptive law and the learning efficiency of the observer.  Conclusions  A finite-time sliding mode control strategy with an adaptive disturbance observer is proposed for servo systems used in ultra-fast laser processing. The method models unknown load disturbances and frictional nonlinearities as a lumped disturbance term. An adaptive observer, integrating an RBF neural network with a finite-time mechanism, accurately estimates this disturbance for real-time compensation. Based on the observer, a finite-time SMC law is formulated, and the practical finite-time stability of the closed-loop system is theoretically proven. Simulations conducted on a permanent magnet synchronous motor platform confirm that the proposed approach achieves superior tracking accuracy, robustness, and control smoothness compared with conventional PI and existing advanced methods. This work offers an effective solution for achieving high-precision control in nonlinear systems subject to strong disturbances.
Research Status and Prospects of Mid-Wavelength Infrared Superlattice Detector Technology
LIU Ming, ZHAO Yaqi, GUAN Xiaoning, ZHANG Fan, LU Pengfei
Available online  , doi: 10.11999/JEIT260083
Abstract:
  Significance   Mid-Wavelength Infrared (MWIR) detectors are widely used in civilian and military applications because of their high sensitivity and excellent temperature discrimination. Type-II SuperLattice (T2SL) materials, especially the InAs/GaSb and InAs/InAsSb systems, have become promising candidates for third-generation infrared photodetectors. This review systematically analyzes the research status and future trends of MWIR T2SL detector technology. It focuses on key photoelectric parameters, including Quantum Efficiency (QE), dark current density, and Specific Detectivity (D*). This work provides a reference for material selection and performance optimization in this rapidly developing field.  Progress   Considerable progress has been made in dark current suppression and photoresponse enhancement for MWIR T2SL detectors. For dark current suppression, advanced barrier structures, such as nBn, XBn, and M-structures, are designed through band-structure engineering. These structures effectively block majority-carrier transport while allowing efficient collection of photogenerated carriers. For instance, an nBn device with an AlAsSb/InAsSb superlattice barrier shows a dark current density of 2.01×10–5 A/cm2 at 150 K (Fig. 2(c)). Strain compensation and optimized epitaxial growth further reduce bulk dark current. One device achieves a dark current density of 4.5×10–7 A/cm2 at 140 K (Fig. 4(f)). Device process optimization, including two-step etching and Zn-diffusion-based planar junction formation, also reduces surface leakage current (Fig. 5, Fig. 6). For photoresponse enhancement, the main strategies include micro/nano-optical structure integration, epitaxial growth optimization, and device process improvement. Monolithically integrated metalenses increase th,e peak responsivity to 9.01 A/W at 300 K (Fig. 7(d)). Guided-mode resonance architectures enable a room-temperature External Quantum Efficiency (EQE) of approximately 60% (Fig. 8(c)). Epitaxial optimization, including stepped absorption layers and interfacial graded doping, increases the QE to 59.4% at 150 K (Fig. 10(c)). Device process optimization, such as substrate removal and Anti-Reflection (AR) coating deposition, also improves QE. An average QE of 63.7% is reported in the 3.7~4.8 μm range (Fig.13(c)). Comparative analysis shows that InAs/GaSb detectors are mainly reported at 77~150 K, whereas InAs/InAsSb detectors show stronger potential for higher-temperature operation, especially near 150 K (Fig. 15, Fig. 16). Overall, at 150K, dark current densities are generally suppressed below 10–4 A/cm2, and peak QEs approach 70%.  Conclusions  T2SL materials, with tunable band structures and low Auger recombination rates, have become a core material platform for high-performance MWIR detection. Current studies have addressed key challenges in dark current suppression and photoresponse enhancement. Through advanced barrier design and device process optimization, dark current densities have been suppressed to the 10–6 A/cm2 level at approximately 150 K. Through optical and epitaxial engineering, QEs have been increased to approximately 60% or higher. The InAs/InAsSb material system is particularly promising for High-Operating-Temperature (HOT) applications.  Prospects  Future development will focus on four main directions. First, the HOT limit should be further increased, with the goal of maintaining diffusion-limited performance at 180 K or higher. Second, large-format Focal Plane Arrays (FPAs) should be developed based on highly uniform material growth through mature Molecular Beam Epitaxy (MBE), aiming for pixel operability higher than 99%. Third, multicolor and multispectral detection should be expanded by precisely tuning superlattice periods, enabling integrated dual-band or multiband MWIR detection with reduced crosstalk. Fourth, new device architectures and coupled physical mechanisms should be explored to extend detector performance and application boundaries.
Excellence Action Plan Leading Column
Efficient and Verifiable Ciphertext Retrieval Scheme Based on Trusted Execution Environment
WU Axin, FENG Dengguo, ZHANG Min, CHI Jialin, YI Yuling
Available online  , doi: 10.11999/JEIT251358
Abstract:
The ciphertext retrieval mechanism enables retrieval over encrypted data. Symmetric Searchable Encryption (SSE) is a critical branch of ciphertext retrieval. However, to save computing resources, cloud servers may return incorrect or incomplete results. Moreover, attackers may exploit information leaked from search and access patterns to reconstruct keyword details. Therefore, protecting the privacy of search and access patterns, while ensuring result verifiability, is necessary and meaningful. Nevertheless, existing verifiable SSE schemes that support search and access pattern privacy usually rely on keyword traversal mechanisms, and their verification mechanisms are inefficient. This results in high computational and communication overhead for users. To address these performance bottlenecks, an efficient and verifiable ciphertext retrieval scheme based on Trusted Execution Environment (TEE) is proposed. To improve ciphertext retrieval efficiency, the scheme uses the collaborative implementation of hardware-level security isolation and oblivious data rearrangement, so that the size of the keyword trapdoor is independent of the size of the keyword dictionary. Meanwhile, the correctness of the returned results is verified by embedding random numbers and blinding the constant terms of polynomials. Owing to these designs, significant efficiency improvements are achieved. Specifically, the scheme ensures that the size of keyword trapdoors depends only on the number of query keywords rather than on the global dictionary size, which effectively reduces communication and computational costs. The scheme also requires only two random numbers to achieve verifiability, which substantially reduces the user's local storage overhead. In addition, techniques such as single-server, single-round result retrieval and symmetric homomorphic encryption are adopted to further improve operational efficiency. Moreover, confidential computing within TEE weakens the security assumptions and trust requirements imposed on TEE. After the security of the proposed scheme is formally proved by simulation-based methods, a comprehensive performance evaluation is conducted. The results confirm that the proposed scheme is substantially more efficient than other schemes with the same functionality.
Satellite Navigation
Research on GRI Combination Design of eLORAN System
LIU Shiyao, ZHANG Shougang, HUA Yu
Available online  , doi: 10.11999/JEIT201066
Abstract:
To solve the problem of Group Repetition Interval (GRI) selection in the construction of the enhanced LORAN (eLORAN) system supplementary transmission station, a screening algorithm based on cross interference rate is proposed mainly from the mathematical point of view. Firstly, this method considers the requirement of second information, and on this basis, conducts a first screening by comparing the mutual Cross Rate Interference (CRI) with the adjacent Loran-C stations in the neighboring countries. Secondly, a second screening is conducted through permutation and pairwise comparison. Finally, the optimal GRI combination scheme is given by considering the requirements of data rate and system specification. Then, in view of the high-precision timing requirements for the new eLORAN system, an optimized selection is made in multiple optimal combinations. The analysis results show that the average interference rate of the optimal combination scheme obtained by this algorithm is comparable to that between the current navigation chains and can take into account the timing requirements, which can provide referential suggestions and theoretical basis for the construction of high-precision ground-based timing system.