Latest Articles

Articles in press have been peer-reviewed and accepted, which are not yet assigned to volumes/issues, but are citable by Digital Object Identifier (DOI).
Display Method:
Resource Allocation Optimization in Dual-RIS Cooperative Rate-Splitting Multiple Access Networks
CHEN Yuang, WU Chang, PENG Mingyu, LU Hancheng
Available online  , doi: 10.11999/JEIT260171
Abstract:
  Objective  In Rate-Splitting Multiple Access (RSMA) systems, the achievable common-stream rate is limited by the user with the weakest channel quality. This constraint reduces scalability, robustness, and user fairness in dense 6G networks. Existing cooperative RSMA architectures partly reduce this bottleneck, but they remain constrained by fixed channel conditions and limited interference management. To address these issues, this paper proposes a dual Reconfigurable Intelligent Surface (RIS) cooperative RSMA system. Two cooperatively deployed RISs create additional controllable propagation paths through cascaded double reflection. The objective is to maximize the system sum rate by jointly optimizing Base Station (BS) Beamforming (BF), Rate Splitting (RS) strategies, and dual-RIS phase configurations, thereby improving spectral efficiency, robustness, and user fairness under users’ Quality of Service (QoS) constraints.  Methods  A tractable system model is developed for the dual-RIS cooperative RSMA system. The model captures cascaded multi-link channels, multi-node channel structures, and interference coupling. Based on this model, a joint optimization problem is formulated to maximize the system sum rate by optimizing BS BF, RS strategies, and the discrete phase shifts of both RISs. Because of strong variable coupling and non-convexity, a low-complexity Alternating Optimization (AO) algorithm is designed. The original problem is decomposed into three subproblems: BS-side RIS phase optimization, user-side RIS phase optimization, and BS BF optimization. Semidefinite Relaxation (SDR) and Successive Convex Approximation (SCA) are used to transform these subproblems into tractable convex forms, which are then solved iteratively with fast convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed dual-RIS cooperative RSMA system. The proposed AO algorithm converges within six iterations under different numbers of RIS reflecting elements, and it reaches 97.8% of the steady-state sum rate within three iterations when M = 170 (Fig. 3). Compared with SOPS and RPS, the proposed phase-configuration scheme obtains 10.6% and 31.8% sum-rate gains when M = 190, respectively (Fig. 4). The proposed RSMA scheme also outperforms NOMA and SDMA by 10.0% and 14.6%, respectively (Fig. 5). Under M = 160 and b = 4, dual-RIS cooperation provides an 11.9% sum-rate gain over the single-RIS scheme, and its performance is close to the CPS upper bound (Fig. 6). Balanced allocation of reflecting elements between the two RISs further improves the sum rate (Fig. 7). The proposed BF strategy also outperforms ZF and RBF, achieving 33.2% and 336.5% gains at a transmit power of 30 dBm, respectively (Fig. 8). Under different RIS cooperation modes, the proposed joint optimization scheme achieves the best overall performance (Fig. 9). These results show that dual-RIS cooperative RSMA improves common-stream decoding, interference suppression, robustness, and user fairness.  Conclusions  This paper investigates a dual-RIS cooperative RSMA communication system. The proposed architecture improves common-stream decoding while mitigating complex interference. To maximize the system sum rate, BS BF vectors, RS vectors, and the discrete phase matrices of two RISs are jointly optimized. A low-complexity AO algorithm based on SDR and SCA is developed to solve the strongly coupled non-convex problem. Simulation results show that the proposed dual-RIS cooperative RSMA scheme achieves clear sum-rate gains over advanced benchmark schemes. Compared with the single-RIS mode and SDMA, it obtains 11.9% and 14.6% rate gains, respectively, while improving system robustness and user fairness.
Intelligent Resource Allocation Algorithm Based on Outdated CSI for Multi-Node URLLC
ZHAO Yizhen, GAO Wei, HU Yulin, ZHU Yao
Available online  , doi: 10.11999/JEIT260216
Abstract:
  Objective  Ultra-Reliable and Low-Latency Communications (URLLC) is widely used in Industrial Internet of Things (IIoT) systems. However, in mobile industrial scenarios such as transportation and inspection, instantaneous Channel State Information (CSI) is difficult to obtain because of feedback overhead. Resource allocation decisions therefore need to be made using outdated CSI. This mismatch restricts system energy efficiency. Traditional convex optimization methods have difficulty addressing this problem. Classical Deep Reinforcement Learning (DRL) algorithms also have limited convergence stability and policy performance under the stringent latency and reliability constraints of URLLC. To address these challenges, this paper considers a multi-node URLLC system under outdated CSI in dynamic scenarios. An energy-efficiency maximization problem is formulated under the Finite BlockLength (FBL) regime, with communication latency and reliability constraints. An efficient and stable algorithm is then designed for joint power and blocklength allocation.  Methods  A Successive Convex Approximation (SCA)-assisted DRL framework is proposed to maximize energy efficiency under outdated CSI. First, an SCA-based algorithm is developed to obtain a pre-allocation solution for transmit power and blocklength. This solution is feasible and physically interpretable, but relatively conservative. Based on this baseline, a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is used for incremental refinement through interaction with the dynamic environment. This process reduces the conservatism of SCA. The SCA solution is used as prior knowledge in the state representation. Node location information is also incorporated into the state space. These designs narrow the policy search space and enable the DRL agent to better capture large-scale channel characteristics and system dynamics under outdated CSI. Learning efficiency and stability are therefore improved.  Results and Discussions  The proposed algorithm is evaluated through simulations and compared with three benchmark algorithms: an SCA-based optimization algorithm, a TD3 algorithm without SCA guidance, and a TD3 algorithm without node location information. The results show that the proposed method outperforms all benchmarks in convergence stability and system energy efficiency. In the training phase (Fig. 3), the average reward of the proposed algorithm increases steadily and converges stably. By contrast, removing node location information leads to lower rewards and stronger fluctuations. Removing SCA guidance causes the algorithm to converge to a much lower reward level. These results confirm the roles of SCA-based prior guidance and location-aware state representation in improving training stability. In the actual operation stage (Fig. 4), the proposed algorithm achieves high and stable energy efficiency and outperforms all comparison algorithms. Under outdated CSI, DRL-based methods can obtain higher energy efficiency than conservative optimization methods when transmission succeeds. However, removing node location information reduces energy efficiency, and removing SCA guidance increases transmission failures. These results verify the effectiveness of both designs in improving energy efficiency and maintaining policy feasibility. The effects of key system parameters are also examined. For basic resource parameters, a moderate increase in the blocklength budget (Fig. 5) or power budget (Fig. 6) improves system energy efficiency. For reliability constraints (Fig. 7), the reliability requirement should be set according to service requirements to avoid resource waste. Finally, the average energy efficiency under different numbers of nodes and different numbers of neurons in the TD3 network is analyzed (Fig. 8). The results provide guidance for algorithm configuration and network-scale design.  Conclusions  This paper addresses energy-efficient resource allocation for multi-node URLLC systems with outdated CSI by integrating SCA and DRL. In the proposed framework, a TD3-based DRL algorithm is guided by an SCA reference solution, and node location information is incorporated into the state representation. This optimization-learning dual-driven framework combines the interpretability and feasibility of model-based optimization with the adaptivity of data-driven learning. Simulation results show that the proposed method achieves higher energy efficiency than SCA-based optimization and conventional TD3 while satisfying URLLC latency and reliability constraints. The SCA reference solution improves policy stability and effectiveness under outdated CSI. Node location information further supports efficient decision-making. This work focuses on a single-cell multi-node scenario under Time Division Multiple Access (TDMA). Practical issues such as multi-cell interference, cooperative scheduling among multiple base stations, and more complex mobility patterns are not considered. Future work will extend the proposed framework to multi-cell and multi-agent scenarios and test its applicability under more severe CSI imperfections.
Load Optimization of Inverter Air-Conditioning Clusters Driven by Constraint Surface Projection and Spatial-Fitness Synergy
ZHENG Bowen, PAN Mingming, WANG Lei, LIU Chang, ZHENG Qingrong, TANG Zhuofan, ZHAO Jianli
Available online  , doi: 10.11999/JEIT260149
Abstract:
  Objective  Supply-demand imbalances in modern distribution networks are intensified by the increasing penetration of distributed renewable energy and frequent extreme high-temperature events. Large-scale Inverter Air-Conditioning (IAC) clusters can be aggregated as virtual energy storage resources for Demand Response (DR), providing an effective way to improve grid flexibility. However, existing dispatch strategies are often limited by the curse of dimensionality. Conventional penalty-function-based soft constraints also fail to strictly satisfy aggregate power equality constraints and may introduce steady-state errors. This paper develops an optimization framework in which grid-side power commands are accurately tracked while user thermal discomfort is reduced and fairness among heterogeneous users is maintained.  Methods  A multi-objective optimization framework based on an Equivalent Thermal Parameter (ETP) model is established to describe the thermodynamic states of heterogeneous buildings. To balance collective comfort and individual fairness, a composite fitness function is designed by integrating a weighted mean-squared error term, a fairness variance term, and a maximum violation suppression term. To eliminate the steady-state errors of traditional penalty-based methods, a Spatial-Fitness Adaptive Particle Swarm Optimization (SFA-PSO) algorithm is proposed. A geometric constraint surface projection mechanism maps particles strictly onto the power-conservation hyperplane, thereby satisfying the aggregate power equality constraint. In addition, the learning factors are dynamically adjusted through a Spatial-Fitness Adaptive (SFA) strategy. This strategy measures the mismatch between a particle’s fitness rank and spatial distance rank, which helps prevent premature convergence in high-dimensional search spaces.  Results and Discussions  Extensive continuous scheduling simulations are conducted in a complex dynamic environment. The environment includes multi-source thermal disturbances, a bidirectional communication packet loss rate of 1%, and Part Load Ratio (PLR) values of 20%, 50%, and 80%. First, ablation experiments confirm that constraint surface projection guarantees power tracking accuracy. Traditional penalty-based methods, such as Penalty Particle Swarm Optimization (Penalty-PSO), produce steady-state power deviations of approximately 10^-1 kW. By contrast, SFA-PSO limits aggregate power tracking errors to within 10^-9 kW (Fig. 3). The SFA strategy also prevents the premature convergence observed in Physical Particle Swarm Optimization (Phy-PSO). It enables continuous fitness reduction, especially in low-load scenarios with narrow feasible regions (Fig. 4). This improvement is attributed to the dynamic evolution of the learning factors. The cognitive factor remains high at the early stage to promote global exploration. It then decreases as the social factor increases, which strengthens local exploitation and improves convergence precision (Fig. 5). Second, continuous dynamic scheduling performance is evaluated through a 6-hour simulation during the peak load period from 12:00 to 18:00. The dispatch interval is 5 min, yielding 72 decision steps. Under tight peak-load constraints, Genetic Algorithm (GA) and Whale Optimization Algorithm (WOA) show severe power-limit violations because their population update rules do not cooperate well with the projection mechanism. By contrast, SFA-PSO maintains strict constraint satisfaction (Fig. 7). SFA-PSO remains at the lowest fitness level throughout the real-time evolution curves, indicating strong robustness against environmental thermal noise and uplink and downlink communication packet loss (Fig. 8). Quantitatively, compared with eight baseline algorithms, including Social Learning Particle Swarm Optimization (SLPSO), Competitive Swarm Optimizer (CSO), and Dynamic State Cluster-Based Particle Swarm Optimization (DSCPSO), SFA-PSO achieves the best overall performance. It obtains an average fitness of 904, a minimum fitness of 243, and the lowest standard deviation of 551 (Table 2). Finally, scalability analyses across cluster sizes from 100 to 1,000 nodes further validate the high-dimensional optimization capability of SFA-PSO. In all scale scenarios, SFA-PSO shows the strongest optimization capacity. It achieves rapid initial descent within the first 20 iterations and maintains continuous exploration in later stages (Fig. 9). Although the projection and SFA mechanisms increase computational time by 30% to 50% compared with basic Particle Swarm Optimization (PSO) (Fig. 6), the absolute optimization time remains stable at approximately 1.5 seconds even for a 1,000-node cluster (Fig. 9). This computational overhead is acceptable for minute-level control cycles and meets the real-time dispatch requirements of modern smart grids.  Conclusions  The proposed SFA-PSO algorithm effectively addresses the steady-state error of traditional soft-constraint methods in aggregate power control. By ensuring accurate tracking of dispatch commands and mitigating high-dimensional search traps, it provides a robust and scalable solution for flexible scheduling of large-scale IAC loads in smart grids. It also maintains a practical balance between grid-side regulation and user-side comfort. The method still has limitations. The constraint projection mechanism depends on the host algorithm, which restricts cross-algorithm generalization. High-precision tracking also increases computational cost. Future work will focus on adaptive constraint handling and lightweight algorithm design. Coordinated scheduling for heterogeneous loads, such as electric vehicles and energy storage, will also be investigated.
Recent Advances in Remote Sensing Image-Text Retrieval Driven by Vision-Language Foundation Models
WU Hui, ZHAO Yan, ZHANG Peirong, HOU Yingyan, QI Xiyu, WANG Lei
Available online  , doi: 10.11999/JEIT260189
Abstract:
  Significance  Remote Sensing Image-Text Retrieval (RS-TIR) connects large-scale Earth observation imagery with natural-language queries and has become an important interface for geospatial intelligence systems. Compared with conventional content-based retrieval, RS-TIR allows users to search for scenes, objects, spatial layouts, and functional regions through semantic descriptions rather than handcrafted visual cues. This capability is increasingly needed in natural resource monitoring, urban governance, disaster response, environmental assessment, and on-demand retrieval from rapidly growing satellite archives. However, RS-TIR remains challenging. Remote sensing imagery is captured from nadir or near-nadir perspectives, shows strong rotation invariance, and contains extreme scale variation, ranging from tiny vehicles to large airports. It also requires domain-specific semantic descriptions, such as land-use attributes, spatial distributions, and geoscientific relations. Meanwhile, high-quality image-text annotations remain limited relative to the scale of remote sensing data. These properties widen the cross-modal semantic gap between images and language and limit the generalization ability of traditional cross-modal retrieval methods. Against this background, this review examines how Vision-Language Foundation Models (VLMs) reshape RS-ITR through large-scale contrastive pre-training, stronger transferable representations, and more flexible multimodal interaction mechanisms. It also explains why remote sensing adaptation is needed and why a focused synthesis of architectures, datasets, alignment mechanisms, and future directions is timely for this field.  Progress   The technical development of RS-ITR is reviewed from three complementary perspectives. First, this review summarizes the domain-specific challenges that shape the task, including visually isotropic topology with extreme scale variation, professional and fine-grained textual semantics, and the compounded cross-modal semantic gap between overhead imagery and natural-language descriptions (Fig. 3). The overall survey structure is then presented to show the logical progression from task formulation to future challenges (Fig. 1). From a methodological perspective, RS-ITR has evolved from handcrafted visual descriptors and shallow semantic mapping to deep representation learning, and then to VLM-driven paradigms with stronger generalization and zero-shot transfer capability (Fig. 4, Table 2). Early methods rely on color, texture, shape, and hash-based retrieval. However, they struggle to model high-level geospatial semantics and complex scene composition. Deep learning methods improve retrieval by learning joint embedding spaces, adopting dual-encoder or interaction-based architectures, and using multi-scale feature fusion and region-aware matching. These methods improve semantic consistency, but they still depend heavily on labeled data and often show limited robustness in open or cross-sensor scenarios. Second, this review summarizes the benchmark ecosystem used to evaluate these methods. Representative datasets range from small-scale test sets, such as Sydney-Caption and UCM-Caption, to mainstream benchmarks, such as RSICD and RSITMD, and recent large-scale training resources, such as RS5M and SkyScript (Table 1). These datasets show a clear transition from small manually annotated corpora to web-scale or automatically generated image-text pairs. This transition supports domain pre-training and large model adaptation. Third, this review analyzes the core VLM techniques that now drive progress in RS-ITR. The model spectrum and representative architecture families are systematically summarized, including contrastive dual-encoder models, multimodal interaction models, and remote sensing foundation models integrated with large language models (Fig. 5, Fig. 6, Table 3). Domain adaptation routes are further grouped into continued remote sensing pre-training, parameter-efficient transfer learning, adapter-based tuning, prompt learning, and instruction tuning. At the semantic alignment level, this review focuses on contrastive joint embedding, fine-grained multi-scale alignment, and the use of remote sensing priors, such as spatial topology and geolocation. Performance comparisons on RSICD and RSITMD show that remote sensing VLMs, especially RemoteCLIP, GeoRSCLIP, iEBAKER, and LRSCLIP, yield consistent gains in mean Recall (mR) and overall retrieval robustness (Table 4). In parallel, this review tracks the extension of retrieval capability into unified multi-task remote sensing models, in which retrieval, grounding, segmentation, and reasoning begin to share a common multimodal representation space.  Conclusions  Several conclusions are drawn from the comparative analysis. First, VLMs establish a dominant paradigm for RS-ITR because they narrow the cross-modal semantic gap and improve transferability across datasets and scenes. Second, no single architecture is universally optimal. Dual-encoder models remain attractive for large-scale retrieval because of their efficiency, whereas interaction-based or instruction-enhanced models provide finer semantic alignment at a higher computational cost. Third, domain adaptation is indispensable. Continued pre-training on remote sensing image-text corpora, parameter-efficient tuning, and prompt-based adaptation consistently outperform direct reuse of internet-trained VLMs. This finding indicates that remote sensing imagery differs too strongly from natural-image distributions for generic pre-training alone to be sufficient. Fourth, the most effective recent methods do not improve performance through scale alone. They also exploit remote sensing-specific information, including multi-scale structures, foreground objects, explicit keyword reasoning, and spatial priors. Finally, this review shows that the field is shifting from isolated retrieval models toward more general geospatial multimodal systems. Retrieval is no longer treated only as a matching task. It is also becoming a key capability that supports question answering, instruction following, knowledge augmentation, and coordinated reasoning in remote sensing applications.  Prospects   Future research is expected to advance in four closely related directions. The first direction is the unified representation of multi-source heterogeneous data, especially the integration of optical imagery with Synthetic Aperture Radar (SAR), hyperspectral data, thermal infrared observations, and multi-temporal acquisitions. The second direction is knowledge-enhanced retrieval, in which geospatial priors, land-use rules, remote sensing terminology, and external knowledge bases are incorporated into multimodal alignment and retrieval-augmented reasoning. The third direction is lifelong and open-world learning. Real deployment requires models to remain reliable under seasonal variation, sensor updates, regional domain shifts, cloud contamination, and newly emerging categories, while avoiding catastrophic forgetting. The fourth direction is efficiency and deployability. Practical remote sensing systems often operate under tight computational budgets. Therefore, lightweight tuning, sparse computation, token reduction, model compression, and on-orbit and edge inference will become increasingly important. Interactive and explainable retrieval is also likely to gain importance. It allows analysts to refine queries through dialogue and inspect the image regions or semantic cues that support retrieval decisions. Overall, continued progress in data construction, domain adaptation, semantic alignment, and efficient multimodal modeling is expected to make RS-ITR a more robust infrastructure capability for Earth observation applications.
A Cryptographic Side-Channel Security Modeling and Formal Verification Method
WANG Xingxin, HU Wei, HUANG Xuan, LIN Chenyu, ZHOU Yi
Available online  , doi: 10.11999/JEIT260631
Abstract:
  Objective  Compared with post-silicon side-channel security analysis, pre-silicon side-channel security verification during the design phase enables the earlier identification of potential side-channel security vulnerabilities in cryptographic core designs, thereby effectively reducing the cost and time of post-silicon remediation. However, most existing pre-silicon side-channel security assessment approaches rely on data-driven statistical analysis or artificial intelligence techniques and require complex calculations on large amounts of data to mitigate the impact of insufficient coverage on the assessment results. In addition, existing methods typically adopt independent modeling strategies for different types of side channels, lacking a unified side-channel security modeling approach. A cryptographic side-channel security modeling and formal verification method is proposed, supporting unified and automated modeling of different types of side channels by constructing a side-channel security model. The method can identify potential timing side-channel, power side-channel and fault injection vulnerabilities in cryptographic core designs, and analyze the effectiveness of side-channel countermeasures based on side-channel security property checking.  Methods  The proposed cryptographic side-channel security modeling and formal verification method includes side-channel security model construction, side-channel security property extraction, and side-channel security verification. The side-channel security model uses information flow analysis to characterize timing side-channel leakage, power side-channel leakage, and fault propagation behavior in cryptographic core designs, providing an effective mathematical model for side-channel security verification. Specifically, the side-channel security model utilizes changes in signal labels to analyze information flows during the encryption process by assigning a label to a signal bit and defining label propagation rules. Side-channel security properties formally describe the behavioral characteristics of side-channel leakage, including timing properties, power properties, and fault properties, providing theoretical support for side-channel security verification. Side-channel security verification uses the extracted security properties as verification constraints and employs formal verification tools to identify potential timing side-channel, power side-channel, and fault injection vulnerabilities in cryptographic core designs. Furthermore, the method can analyze the effectiveness of masking and fault injection countermeasures against side-channel vulnerabilities.  Results and Discussions  The proposed side-channel security verification method utilizes formal verification techniques to accurately identify potential side-channel security vulnerabilities in various block cipher core designs, and evaluate the effectiveness of side-channel countermeasures based on side-channel security property constraints. The timing side-channel verification results demonstrate that the proposed method can accurately identify timing side-channel security vulnerabilities in AES, SM4, LED, PRESENT, and IDEA core designs within 20s (Table 2, Fig. 7). No timing side-channel vulnerabilities are identified in the other cryptographic core designs, except for IDEA, which exhibits timing side-channel vulnerabilities caused by modular multiplication operations. The power side-channel security verification results show that formal checks based on controllability property, key–power distinguishability coupling property and key–power nonlinear coupling property can accurately identify target modules with potential power side-channel security vulnerabilities in AES, SM4, LED and PRESENT within 1 minute (Table 3). The key expansion module in cryptographic core designs does not cause key leakage through key-dependent power consumption, as it fails to satisfy the controllability property. In addition, the experimental results indicate that the masking protection in the RSM core design can prevent the correct key from being distinguished through random masking (Fig. 8). The fault injection security verification results for three AES core designs with infective countermeasures demonstrate that the proposed method can analyze the effectiveness of fault infection countermeasures. The results show that the infection countermeasure requires not only altering the fault propagation path but also disrupting the algebraic relationships among faults (Table 4, Fig. 10).  Conclusions  This paper proposes a cryptographic side-channel security modeling and formal verification method to address the lack of formal mathematical models and the limited completeness of existing data-driven side-channel security assessment methods. The proposed method first achieves unified modeling of timing leakage, power leakage, and fault propagation behaviors from the perspective of information flow analysis. Based on the constructed side-channel security model, side-channel security properties are extracted to formally characterize the behavioral features of side-channel information leakage and propagation during the encryption process. Potential side-channel security vulnerabilities in cryptographic core designs are then identified through formal checking using the extracted security properties as constraints. The proposed method provides an effective solution for the unified modeling and formal security verification of different types of side channels. Experimental results obtained from the side-channel security verification of various block cryptographic core designs demonstrate that: (1) the proposed method can uniformly model timing side-channel leakage, power side-channel leakage, and fault propagation behaviors in cryptographic core designs; (2) the proposed method can accurately identify timing side-channel vulnerabilities, power side-channel vulnerabilities, and fault injection vulnerabilities in cryptographic core designs, including AES, SM4, IDEA, LED and PRESENT; (3) the proposed method can analyze the effectiveness of masking and fault infection countermeasures. However, this study only qualitatively identifies side-channel security vulnerabilities in cryptographic core designs; pre-silicon quantitative assessment of side-channel leakage should be investigated in future work.
A Lightweight Dual-Stream Convolutional Network Feature Fusion Method for UAV RF Recognition
DONG Pengyu, XIANG Xin, LV Siting, LIANG Yuan, WANG Rui, MAO Hu
Available online  , doi: 10.11999/JEIT260464
Abstract:
  Objective  With the rapid proliferation of Unmanned Aerial Vehicles (UAVs) and the escalating demand for airspace security, radio frequency (RF) fingerprint recognition has emerged as a pivotal technology for identifying non-cooperative UAVs. However, existing methods grapple with significant challenges, including poor robustness in low signal-to-noise ratio (SNR) environments and prohibitive computational complexity, which severely hinder their deployment on resource-constrained tactical edge devices. To address these critical limitations, this paper proposes a novel lightweight dual-stream convolutional network tailored for UAV RF recognition. This network is designed to extract static spectral texture features and dynamic temporal gradient features in parallel, complemented by a meticulously crafted lightweight feature fusion strategy.  Methods  The proposed network architecture is ingeniously designed to process RF signals. The input signal undergoes a Short-Time Fourier Transform (STFT) to generate a two-dimensional spectrogram, which serves as the primary input. The network is bifurcated into two parallel streams: a static stream and a dynamic stream. The static stream is engineered to capture the inherent static spectral patterns and energy distributions within the STFT spectrogram. It comprises a series of stacked convolutional blocks, each integrating convolutional layers, batch normalization, and ReLU activation functions, followed by max-pooling layers to progressively downsample the feature maps and increase the channel depth. Conversely, the dynamic stream is dedicated to enhancing feature discriminability, particularly in low-SNR scenarios. It begins by computing the temporal gradient of the input spectrogram, effectively suppressing static background noise and accentuating dynamic signal variations. This gradient map is then processed by a symmetric set of convolutional blocks, mirroring the structure of the static stream. To maintain model efficiency, an element-wise addition fusion strategy is employed to integrate the features from both streams, ensuring a balance between feature complementarity and computational overhead. The fused features are subsequently fed into a classification head, consisting of an adaptive average pooling layer, dropout layers for regularization, and fully connected layers to produce the final classification output. Extensive experiments are conducted on the publicly available DroneRF dataset, encompassing ablation studies to dissect the contribution of each component, comparative analyses of various fusion strategies, and rigorous evaluations of the model’s lightweight characteristics.  Results and Discussions  The experimental results unequivocally demonstrate the efficacy of the proposed method. The dual-stream network achieves a remarkable 95.65% accuracy on the test set, representing a substantial 5 percentage point improvement over the best-performing single-stream network. A critical analysis reveals that the temporal gradient operation contributes significantly to this enhancement by improving the average SNR by 1 dB, thereby bolstering feature discriminability in challenging low-SNR environments. Furthermore, the model’s lightweight design is a standout feature, with a mere 0.58 million parameters, making it eminently suitable for deployment on tactical edge devices. Ablation studies and feature visualization analyses provide compelling evidence for the complementary nature of static and dynamic features. The static stream adeptly captures broad spectral contours, while the dynamic stream focuses on fine-grained temporal variations. The element-wise addition fusion strategy proves superior, outperforming other approaches like feature concatenation and attention-based fusion in terms of both performance and computational efficiency, thereby validating the rationale behind the lightweight design.  Conclusions  This paper presents a comprehensive solution to the challenges of UAV RF recognition in complex environments by proposing a lightweight dual-stream convolutional network. The method effectively enhances recognition accuracy and robustness through the synergistic combination of dual-stream feature extraction and the SNR-enhancing properties of temporal gradient features, all while maintaining a lightweight architecture suitable for edge deployment. The proposed approach offers a significant advancement in the field, providing a robust and efficient solution for UAV identification. Future research endeavors will focus on further enhancing the model’s adaptability to complex electromagnetic environments, incorporating the effects of sensor noise, and extending the framework to multi-UAV cooperative scenarios.
Complex-domain Joint Spectrum Sensing Method for UAV Swarms in Complex Electromagnetic Environments
QIAN Hui, CHEN Li, YIN Huarui, WANG Weidong
Available online  , doi: 10.11999/JEIT260499
Abstract:
  Objective  The rapid development of the low-altitude economy is increasing the use of unmanned aerial vehicle (UAV) swarms in emergency communication, urban logistics, reconnaissance, and low-altitude network coverage. These applications require reliable spectrum awareness to support cooperative communication, dynamic spectrum access, and interference avoidance. However, low-altitude electromagnetic environments often contain low-SNR signals, multipath propagation, non-cooperative interference, and multiple coexisting transmissions. Conventional energy, cyclostationary-feature, and matched-filter detectors are sensitive to noise uncertainty, computational cost, or prior waveform knowledge. Learning-based methods can improve robustness, but many rely on power spectral density (PSD) or short-time Fourier transform (STFT) representations. These representations may weaken phase information or require costly two-dimensional time-frequency preprocessing. Existing methods also focus mainly on spectrum occupancy and provide limited information about overlapping transmissions. This study therefore develops a low-latency joint sensing method that preserves magnitude and phase information while estimating spectrum occupancy and interference overlap for each frequency bin. The method is intended for local spectrum sensing at UAV nodes under resource and latency constraints.  Methods  The proposed RadioSEUnet pipeline contains two stages: magnitude-phase feature construction and joint spectrum-state estimation. First, each complex baseband in-phase/quadrature (I/Q) sequence is multiplied by a Hann window and transformed using a one-dimensional fast Fourier transform (FFT). The resulting complex spectrum is decomposed into logarithmic magnitude and phase components. The two components are normalized separately and stacked as a two-channel feature tensor. Binary labels indicate spectrum occupancy and multi-signal overlap at each frequency bin. RadioSEUnet adopts a U-shaped encoder-bottleneck-decoder architecture with four encoder stages containing 64, 128, 256, and 512 channels. Each RadioSEBlock combines a complex-parameterized convolution, squeeze-and-excitation channel attention, and a residual connection. The convolution couples the magnitude and phase feature streams through constrained cross-channel operations. Two prediction heads convert the shared representation into a spectrum-occupancy probability mask and an interference-overlap probability mask. The model is optimized using an equally weighted sum of two binary cross-entropy losses. Training uses AdamW, cosine-annealing learning-rate scheduling, early stopping, a batch size of 64, and at most 200 epochs. The complete data collection contains 72,000 complex I/Q records, including 48,000 simulated records and 24,000 measured records. The simulated subset covers Wi-Fi, BLE, ZigBee, LoRa, QPSK/16QAM, FM, and AM signals. Signal-to-noise ratios range from –15 dB to 10 dB under additive white Gaussian noise and Rayleigh fading. The measured subset was collected using a USRP N310 in the 2.4–2.5 GHz ISM band at 100 MS/s over a 1 ms observation interval. The controlled quantitative evaluation uses an 8:1:1 split of the simulated subset. A separate simulated-to-measured protocol is defined in the main text to examine cross-domain generalization. RadioSEUnet is compared with six PSD- or STFT-based baselines under matched data splits and hardware conditions. Performance is measured using intersection over union (IoU), precision, recall, preprocessing time, inference time, and total sensing latency.  Results and Discussions  The SNR-dependent quantitative results reported here are obtained using the controlled simulated-data protocol. At -15 dB, RadioSEUnet achieves an IoU of 0.768 and a recall of 0.846 for spectrum occupancy detection. Compared with the second-best STFT-RADN baseline, these values correspond to absolute improvements of 0.186 and 0.166, respectively. For interference-overlap detection, RadioSEUnet achieves an IoU of 0.456 and a precision of 0.768 at –15 dB. The corresponding improvements over STFT-RADN are 0.246 and 0.275. The lower IoU for interference-overlap detection indicates that weak overlap boundaries remain difficult to separate from strong-signal sidelobes and background noise. The latency evaluation is conducted on the workstation specified in the main text. Magnitude-phase preprocessing requires 21.04 ms, and network inference requires 4.77 ms, producing a total sensing latency of approximately 25.81 ms. STFT-YOLOv3 requires 98.9 ms under the same hardware setting, so the proposed pipeline is approximately 3.8 times faster in this comparison. Ablation experiments show that magnitude-phase preprocessing, complex-parameterized feature coupling, and channel attention each improve low-SNR sensing performance. Removing the magnitude-phase preprocessing produces the largest degradation. These results indicate that preserving complementary magnitude and phase information is useful for weak-signal and interference-overlap detection. They do not, however, establish performance on airborne hardware or across unreported radio environments.  Conclusions  RadioSEUnet combines a magnitude-phase representation, constrained cross-channel feature coupling, channel attention, multiscale feature fusion, and dual-head prediction. It jointly estimates spectrum occupancy and interference-overlap states while avoiding two-dimensional STFT preprocessing. Under the controlled simulated-data protocol, the method provides higher point estimates than the six evaluated baselines at low SNR and reduces total sensing latency on the evaluated workstation. The present evidence is limited to the reported signal types, channel models, hardware configuration, and the 2.4–2.5 GHz measurement band. Quantitative simulated-to-measured results, tests on wider bands, additional interference types, repeated trials, and deployment on airborne edge hardware are still required. Future work will therefore focus on cross-domain validation, lightweight deployment, boundary-aware interference modeling, and integration with spectrum resource management for UAV networks.
Design of a Channel-Adaptive Denoiser for Digital Semantic Communications
WU Yanjun, LIU Zhangyuhang, YANG Wenxin, YAN Mubiao, ZHOU Hao, ZHAO Yajun, XIE Zhuochen, LIANG Xuwen
Available online  , doi: 10.11999/JEIT260523
Abstract:
  Objective  Practical semantic communication should simultaneously satisfy two requirements: compatibility with existing digital communication infrastructures and robustness under varying channel conditions. Semantic-oriented modulation (SOM) provides a feasible way to map continuous semantic features into layered digital constellation symbols, thereby making semantic transmission compatible with conventional digital systems. However, the digitization process also introduces structured quantization distortion, which makes receiver-side recovery more difficult than in continuous semantic transmission. Although diffusion models have shown strong capability in channel-adaptive semantic recovery, directly applying them to SOM-based digital semantic communication is still limited by the structured distortion introduced by SOM. Therefore, this paper focuses on channel-adaptive receiver design for digital semantic communication and investigates how to compensate SOM-induced structured distortion before subsequent recovery.  Methods  An SOM-based digital semantic communication system for image transmission over an additive white Gaussian noise (AWGN) channel is considered. The proposed receiver adopts a two-stage structure composed of a Quantization Noise Predictor (QNP) and a diffusion recovery module. In the first stage, QNP estimates and compensates the structured quantization distortion introduced by SOM from the layer-wise soft received symbols. In the second stage, the compensated semantic representation is further refined by a diffusion denoiser, whose inference step number is adaptively selected according to the estimated signal-to-noise ratio (SNR). The QNP includes a shared feature extraction frontend, a classification branch exploiting discrete SOM constellation priors, and a regression branch performing fine-grained continuous distortion compensation. A Feature-wise Linear Modulation (FiLM) mechanism is used to incorporate SOM parameters and channel-state information, so that the same QNP can adapt to different modulation configurations and channel conditions. In addition, a composite loss with classification loss, regression loss, and distribution regularization is designed to improve the statistical properties of the compensated residual noise.  Results and Discussions  Experiments are conducted on the CLIC dataset using PSNR and MS-SSIM. First, the proposed method is compared with VAE, VAE+Diff, VAE+SOM, VAE+SOM+QNP, VAE+SOM+Diff, and JCM. The results show that direct SOM-based digitization causes noticeable performance degradation, while the proposed method consistently improves reconstruction quality over digital semantic baselines. In particular, QNP alone already provides stable gains over the SOM-only receiver, indicating that its effectiveness does not rely on diffusion recovery itself. Moreover, VAE+SOM+QNP achieves performance close to VAE+SOM+Diff while requiring much lower computational cost, and combining QNP with diffusion yields the best overall performance. Second, two training strategies, namely independent QNP training and diffusion-assisted fine-tuning, are compared. The results show that diffusion-assisted fine-tuning provides only limited additional gains but significantly increases training cost and complexity, so independent training offers a more practical balance. Third, experiments under different SOM configurations and different SNR conditions verify that QNP provides stable gains across different modulation orders and SOM layer settings. Latency analysis further shows that QNP introduces only a small fixed overhead, whereas the diffusion module dominates the total inference time; therefore, the adaptive diffusion-step schedule is selected according to the measured latency-PSNR trade-off. Fourth, Gaussianity analysis based on the Kullback-Leibler divergence and Wasserstein distance shows that QNP compensation significantly improves the Gaussianity of the residual noise, while the version with distribution regularization achieves the best statistical consistency.  Conclusions  This paper proposes a channel-adaptive receiver for digital semantic communication, in which QNP-based front-end compensation is combined with diffusion-based semantic recovery. The main contribution lies in introducing a lightweight and independently effective QNP module to compensate SOM-induced structured quantization distortion before subsequent recovery. Experimental results show that QNP alone can already stably improve digital semantic reconstruction under different SNR conditions and different SOM configurations, while its combination with diffusion recovery yields the best overall performance. Therefore, the proposed method provides an effective way to improve semantic reconstruction quality and channel adaptability while preserving compatibility with existing digital communication infrastructures.
Fusing Global Perspective Rectification and Fine-grained SemanticDecoupling for Language-conditioned Robotic Grasp Detection
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng
Available online  , doi: 10.11999/JEIT260442
Abstract:
  Objective  Accurate grasp detection from language instructions is essential for service robots to achieve natural human-robot interaction. Existing methods primarily rely on large-scale data-driven training or hierarchical feature fusion to align visual perception with textual instructions. However, they generally overlook the strong coupling between target objects and background clutter in low-level visual features, leading to degraded compositional generalization under cross-view and unseen-scene conditions. To address this limitation, a dual-view cross-scene grasp detection and object localization dataset is constructed to systematically evaluate and improve the compositional generalization of existing models. Based on this benchmark, a Simultaneous Grasp detection and object Localization Network (SGL-Net) is proposed to jointly predict object locations and optimal grasp poses. The proposed framework enables service robots to manipulate objects according to natural language instructions in real-world dynamic environments, providing technical support for embodied intelligence.  Methods  The proposed SGL-Net is illustrated in Fig. 1. First, a Cross-modal Global Context Modulation Module (CGCMM) is proposed to exploit semantic priors from language instructions for adaptive viewpoint correction and background suppression during the early stage of visual feature extraction. Second, a Word-Pixel Cross-modal Alignment Module (WPCAM) is designed to achieve fine-grained semantic decoupling through a flattening-based cross-modal attention mechanism, thereby improving semantic understanding in complex dynamic scenes. Finally, a unified decoder jointly predicts object locations and optimal grasp poses from the fused multimodal features.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on the reconstructed dual-view cross-scene dataset containing bottom-view and top-view scenes and on a real-world robotic grasping platform. Comparative results demonstrate that SGL-Net consistently outperforms mainstream CNN-based and CLIP-based methods in both grasp detection and object localization (Tables 2 and 3). Ablation studies further verify the effectiveness of CGCMM and WPCAM in improving fine-grained semantic alignment and semantic decoupling (Tables 4 and 5). Furthermore, qualitative results (Figs. 46) and real-world robotic experiments (Fig. 7) demonstrate that SGL-Net can be reliably deployed in complex physical environments. Overall, the proposed network exhibits strong generalization capability and excellent potential for practical robotic applications.  Conclusions  To improve cross-view and cross-scene generalization in language-conditioned robotic grasp detection, this paper constructs a dedicated validation dataset and proposes SGL-Net, which jointly performs grasp detection and object localization. By integrating CGCMM and WPCAM, the proposed network accurately localizes instruction-specified objects and predicts optimal grasp poses. Experimental results obtained on multiple benchmark scenarios and a real-world robotic platform demonstrate the superior performance and practical applicability of the proposed method. Future work will focus on integrating Large Multimodal Models (LMMs) and adapting the proposed framework through fine-tuning to further improve zero-shot robotic grasp detection.
Decision Learning Correction Network: Fusion Classification of Hyperspectral Images and LiDAR Data
WANG Haoyu, LIU Nuofei, CHENG Yuhu, LIU Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260362
Abstract:
  Objective  HyperSpectral Images (HSI) and Light Detection And Ranging (LiDAR) provide complementary information for land-cover classification. HSI captures rich spectral information for material discrimination, while LiDAR provides elevation and structural information for spatial characterization. However, most existing fusion methods treat multimodal fusion as a one-shot static aggregation process, implicitly assuming that a fixed fusion strategy is applicable to all pixels and regions. This assumption is difficult to satisfy in complex remote sensing scenes, where class-boundary and cross-modal heterogeneous regions exhibit high information density but account for only a small proportion of samples (Fig. 1). To address this limitation, this paper proposes a Decision Learning Correction Network (DLCN) that reformulates static HSI-LiDAR fusion as a context-dependent sequential decision-making process.  Methods  The proposed DLCN consists of feature extraction, fusion decision learning, and classification. First, HSI and LiDAR are processed through two parallel branches to extract spectral and spatial features and elevation and structural features, respectively. The extracted features are then concatenated to form the current state and are fed into an Actor-Critic framework. The Actor network generates fusion actions to adaptively adjust modality contributions, while the Critic network evaluates the long-term value of each action for classification. To improve learning from difficult samples, a key-sample-oriented sampling module assigns higher sampling probabilities to samples with larger modal fidelity loss. Meanwhile, a modal fidelity constraint mechanism evaluates spectral fidelity, feature consistency, structural preservation, and resolution matching, and corrects destructive actions during fusion. Through this closed-loop framework, DLCN performs dynamic generation, evaluation, and correction of fusion actions, thereby producing high-quality fusion features for classification (Fig. 2).  Results and Discussions  Experiments are conducted on the Houston2013, Trento, and MUUFL datasets. DLCN achieves the highest Overall Accuracy (OA) of 97.85%, 99.58%, and 94.38% on the three datasets, respectively, outperforming CHNet, DSymFuser, mPMCL, MEDFN, S3F2Net, and MSAF. The classification maps demonstrate that DLCN effectively reduces misclassification in class-boundary, mixed land-cover, and structurally complex regions, producing results that more closely match the ground-truth maps across all three datasets (Figs. 35). Ablation studies further demonstrate that the value-guided policy optimization mechanism, key-sample-oriented sampling module, and modal fidelity constraint mechanism each improve classification performance. Compared with the baseline models, the complete DLCN consistently increases OA on Houston2013, Trento, and MUUFL, validating the effectiveness of the proposed decision-learning-correction framework. Time-step analysis shows that DLCN progressively improves classification accuracy while maintaining stable spectral-angle variation during sequential decision making (Fig. 6). Furthermore, DLCN achieves inference times of 1.32 s, 0.86 s, and 2.23 s on the three datasets, respectively, ranking first among the compared methods. These results indicate that the additional computation introduced by the Actor-Critic decision framework and modal fidelity constraint mechanism is effectively translated into improved classification performance without imposing excessive computational cost.  Conclusions  This paper proposes a DLCN for HSI and LiDAR fusion classification. Unlike conventional static fusion methods, DLCN formulates multimodal fusion as a sequential decision-making process and adaptively adjusts fusion strategies according to the local context. Its closed-loop framework enables fusion actions to be generated, evaluated, and corrected throughout the decision process, thereby producing high-quality fusion features for classification. Experimental results demonstrate that DLCN produces more accurate classification maps in heterogeneous remote sensing scenes, and the time-step analysis further confirms the stability of the sequential decision-making process. Future work will focus on more fine-grained feature representation and more robust policy optimization to improve model generalization in complex remote sensing scenes.
Radiation-Hardened Ga2O3 MOSFET Design Featuring NiO Heterojunction and Comb-Shaped Gate Modulation
GAO Sheng, ZHANG Lin, WU Yanjun, WANG Qi, JING Liang
Available online  , doi: 10.11999/JEIT260396
Abstract:
  Objective  Gallium Oxide Metal-Oxide-Semiconductor Field-Effect Transistor (Ga2O3 MOSFET) is regarded as a promising power device for high-voltage applications, particularly in aerospace and satellite power systems, because of its ultra-wide bandgap and high critical breakdown field. However, the Conventional MOSFET (C-MOSFET) exhibits limited reliability in space radiation environments. Under off-state conditions, the electric field is highly concentrated near the gate edge. Heavy-ion irradiation generates dense electron-hole pairs along the ion track. Driven by the intense electric field, these carriers undergo avalanche multiplication through impact ionization, causing the drain current to increase sharply without recovery and ultimately leading to irreversible Single-Event Burnout (SEB) at relatively low drain bias. This failure mechanism severely limits the application of Ga2O3 MOSFETs in harsh radiation environments. Furthermore, the lack of reliable and efficient p-type doping restricts the implementation of conventional radiation-hardening techniques, including junction termination extension and junction isolation. Therefore, ionization-induced carriers readily accumulate in sensitive regions, increasing susceptibility to Single-Event Effect (SEE). The extremely low thermal conductivity of Ga2O3 further promotes local heat accumulation following heavy-ion irradiation, producing localized hot spots that increase the likelihood of thermal burnout. Existing hardening approaches, including field-plate optimization and dielectric engineering, provide only limited improvement. Moreover, the application of heterojunction structures for radiation hardening has rarely been investigated, and systematic hardening strategies have not yet been established. To address these limitations, this paper proposes a Comb-Shaped Gate Metal-Oxide-Semiconductor Field-Effect Transistor (CSG-MOSFET) incorporating a NiO heterojunction. The proposed structure redistributes the channel electric field, suppresses electric-field crowding at the conventional gate edge, and significantly improves SEB tolerance, providing an effective solution for Ga2O3 power devices operating in harsh radiation environments.  Methods  Technology Computer-Aided Design (TCAD) simulations are performed to evaluate the electrical characteristics and SEB performance of the proposed CSG-MOSFET in comparison with the C-MOSFET. The simulations incorporate high-field mobility, Shockley-Read-Hall recombination, Auger recombination, impact ionization, and heavy-ion models. Based on the charge-compensation effect of the p-NiO/n-Ga2O3 heterojunction, the proposed structure utilizes the extended depletion region formed at the heterointerface to redistribute the channel electric field. This heterojunction-induced depletion region improves electric-field uniformity and enhances SEB tolerance. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field away from the conventional gate edge, suppress local electric-field crowding, and improve device reliability under high-voltage and radiation conditions.  Results and Discussions  Simulation results demonstrate that the optimized Double Comb-Shaped Gate MOSFET (DCSG-MOSFET) significantly improves radiation hardness compared with the C-MOSFET. The SEB Threshold Voltage (VSEB) increases from 240 V to 2 280 V, while the Breakdown Voltage (BV) increases from 2 000 V to 3 500 V. Meanwhile, the specific on-resistance decreases. Therefore, the Baliga Figure of Merit (BFOM) and the SEB-based figure of merit are substantially improved. The NiO heterojunction and comb-shaped gate columns effectively redistribute the electric field, shifting the peak electric field from the conventional gate edge to the outer gate-column edge and suppressing local electric-field crowding. These improvements substantially enhance the radiation hardness of the device.  Conclusions  A radiation-hardened DCSG-MOSFET incorporating a NiO heterojunction is proposed and evaluated using TCAD simulations. The optimized structure significantly improves SEB tolerance while maintaining excellent electrical performance. Compared with the C-MOSFET, both VSEB and BV are substantially increased, demonstrating enhanced blocking capability. Charge compensation at the p-NiO/n-Ga2O3 heterojunction forms an extended depletion region that effectively redistributes the channel electric field and suppresses electric-field crowding near the conventional gate edge. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field to the outermost gate-column edge, thereby suppressing impact ionization induced by heavy-ion irradiation and effectively mitigating SEB. The reduced specific on-resistance further improves the BFOM and the SEB-based figure of merit. These results demonstrate that the proposed DCSG-MOSFET is a promising candidate for power electronic applications in harsh radiation environments, including aerospace and satellite systems.
Indoor Visible Light Positioning Based on CNN-MLP Multi-Feature Fusion under Random Receiver Tilt Conditions
JIA Kejun, WANG Jian, MAO Lifei, YOU Wei, HUANG Ziyang, PENG Duo
Available online  , doi: 10.11999/JEIT251021
Abstract:
  Objective  Traditional Visible Light Positioning (VLP) methods based on Received Signal Strength (RSS) are unstable when the receiver undergoes orientation perturbations. Such perturbations disrupt the correspondence between optical power and spatial position, which makes reliable three-dimensional (3D) positioning difficult. Existing approaches usually rely on Inertial Measurement Units (IMUs) to obtain orientation information. However, sensor fusion increases system complexity and hardware cost and also introduces cumulative errors. To address these issues, this paper proposes a positioning method that fuses incidence-angle cosine estimation derived from a Photodiode (PD) array with RSS information, which enables high-accuracy 3D indoor positioning under receiver orientation perturbations.  Methods  In the proposed fusion-based positioning method, a multi-PD array structure is first adopted, and a Local Coordinate System (LCS) is established at the array center. Constraint equations are then constructed from differences in the optical power received by the PDs in the array. A Gauss-Newton iterative algorithm is used to estimate the incident light direction vector. By exploiting the orthogonal rotation invariance between the LCS and the Global Coordinate System (GCS), the incident-angle cosine is estimated without orientation sensors. A serial CNN-MLP fusion network is then constructed, in which the estimated incident-angle cosine is introduced as an additional positioning feature beyond RSS-based localization. The network jointly models the RSS and incident-angle cosine information received by the PD array and maps them to 3D spatial coordinates. Finally, training samples are generated by Latin Hypercube Sampling (LHS) to uniformly sample spatial positions and orientation dimensions, thereby improving the representativeness of the training dataset.  Results and Discussions  Simulation experiments are conducted in a 4 m × 4 m × 2.5 m indoor environment. First, the effects of different numbers of PDs and different tilt angles on the accuracy of incident-angle cosine estimation and spatial coverage are evaluated (Fig. 6), and the Cumulative Distribution Functions (CDFs) of positioning errors under different array configurations are compared (Fig. 7). The results show that a 3-PD array with a tilt angle of 40° achieves the best balance of cost, coverage, and positioning accuracy. Next, positioning performance under different receiver tilt angles is analyzed. When the tilt angle is small, more than 70% of positioning errors are below 5 cm. Even when the receiver is tilted by up to 55°, the average error remains within 11.7 cm (Fig. 8). Comparisons of error components show that the error along the Z-axis is significantly smaller than those along the X- and Y-axes (Fig. 9). Further tests are conducted at a height of 0.0 m, which is covered by the training data, and at an unseen height of 0.6 m, which is not included in the training set (Fig. 10). The results show that the proposed model does not strongly depend on a specific height plane and maintains stable 3D positioning performance at unseen heights. Finally, the proposed method is compared with related positioning schemes. It outperforms existing methods in terms of CDF convergence speed, RMSE, and standard deviation (Fig. 11), with an average error reduction of about 2.5 cm and an RMSE reduction of 31.58% compared with Ref. [13].  Conclusions  This paper estimates the incident-angle cosine at the receiver by exploiting differences in the optical power received by different PDs in an array, and introduces this cosine value as a joint positioning feature into conventional RSS-based localization. This design alleviates the instability of position mapping caused by relying only on RSS under random receiver perturbations. By combining the spatial feature extraction capability of CNNs with the nonlinear modeling strength of MLPs, the proposed method effectively maps positioning features to 3D spatial coordinates. The approach reduces reliance on orientation sensors such as IMUs, while overcoming the sensitivity of traditional geometric positioning methods to noise and high-dimensional nonlinear features. Under varying heights and receiver orientations, the proposed algorithm shows clear advantages in both positioning accuracy and stability.
Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement
ZHANG Chunxiang, ZHANG Huibin, GAO Xueyao
Available online  , doi: 10.11999/JEIT260061
Abstract:
  Objective  With the rapid growth of biomedical literature, biomedical word sense disambiguation (WSD) has become essential for medical text mining and clinical data analysis. However, existing methods suffer from semantic noise, fine-grained category discrimination, and limited generalization in low-resource scenarios. This study proposes a three-branch parallel WSD framework with contrastive learning, integrating multi-pretrained models, chi-square attention, Focal+Margin hybrid loss, and hard sample mining. The proposed method improves semantic representation, robustness, and discrimination ability, providing an effective solution for biomedical semantic mining.  Methods  The proposed framework integrates Electra, mDeBERTa, and Flan-T5 to extract complementary contextual features from biomedical terms. A chi-square attention module is designed to select representative features, while a Focal+Margin hybrid loss improves discrimination under class imbalance. In addition, a two-stage hard sample mining strategy and a core-term constrained contrastive learning mechanism are introduced to enhance the learning of difficult samples and semantic boundaries.  Results and Discussions  The proposed framework integrates chi-square attention and contrastive learning into a three-branch parallel architecture for biomedical WSD. Experiments on the MSH dataset show that the proposed model achieves an accuracy of 95.27%, outperforming the state-of-the-art Neural Concept Embeddings by 0.93%. Ablation studies on contrastive parameters further demonstrate its effectiveness in enhancing semantic discrimination and generalization ability. The model also reduces confusion among similar biomedical terms, achieving an F1-score of 92.1% on minority semantic classes.  Conclusions  This study proposes a three-branch parallel contrastive learning framework for biomedical WSD in complex semantic environments. The framework integrates Electra, mDeBERTa, and FT5 to capture complementary semantic features, while combining chi-square attention, contrastive learning, and two-stage hard sample training to enhance feature discrimination and robustness. Experimental results demonstrate that the proposed method effectively improves disambiguation performance and reduces confusion among semantically similar biomedical concepts. However, this study is limited to monolingual English biomedical texts. Future work will explore multilingual biomedical corpora and integrate domain-specific knowledge graphs to further improve semantic representation.
A Multi-Station Emitter TDOA Deinterleaving Method for Severe Pulse-Loss Environments
LIU Yuchen, ZHAO Yaqin, WU Longwen
Available online  , doi: 10.11999/JEIT260401
Abstract:
  Objective  Modern electronic reconnaissance systems must deinterleave dense and overlapping radar pulse streams in non-cooperative environments. As radar emitters increasingly employ agile waveforms, similar pulse descriptor words, and low-intercept-probability strategies, conventional single-station methods based on carrier frequency, pulse width, and Pulse Repetition Interval (PRI) become less reliable. Multi-station deinterleaving based on Time Difference of Arrival (TDOA) provides a more stable geometric observable, but severe pulse loss still causes sparse cross-station pairing, weak true TDOA peaks, ambiguity-induced spurious peaks, isolated pulses, and fragmented trajectories across time slices. These effects increase false alarms and weaken track continuity. To address these issues, a closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments, with Time of Arrival (TOA) sequences used as the core observables.  Methods  A slice-based framework is developed for continuous reconnaissance. Residual unmatched pulses are carried forward by a sliding window to alleviate cross-slice misalignment. First, candidate pulse pairs satisfying geometric TDOA constraints are generated, and pulse descriptor word constraints on carrier frequency and pulse width are used to remove inconsistent pairs. To reduce the sparsity and binning sensitivity of conventional histograms, multiscale Kernel Density Estimation (KDE) is introduced to reconstruct the TDOA density from sparse candidate differences. Gaussian kernels with different bandwidths are fused, and candidate peaks are adaptively extracted using local statistics and peak widths. Second, a dynamic memory matrix is designed to suppress ambiguity-induced spurious peaks in high pulse repetition frequency scenarios. Since dependent spurious peaks collapse after the dominant peak is extracted and removed, a collapse-rate criterion is defined, and the spurious regions are recorded in a memory mask for subsequent iterations. Third, Dynamic Time Warping (DTW) is used to compare incomplete TOA sequences of isolated pulses with extracted pulse sequences, enabling reassignment of unequal-length and incomplete sequences. Finally, a Kalman-filter-based state-space model tracks multi-baseline TDOA trajectories across successive slices. Predicted and observed TDOA residuals are jointly used for association, so intermittent observations can still be linked to the correct track. In this way, the proposed method forms a closed-loop processing chain that links weak-peak reconstruction, spurious-peak suppression, isolated-pulse reassignment, and trajectory association (Fig. 3).  Results and Discussions  Four simulation scenarios are designed: a high pulse repetition frequency scenario dominated by ambiguity-induced spurious peaks, a parameter-overlapping scenario dominated by isolated pulse reassignment, an ablation scenario for evaluating the memory matrix and DTW modules, and a 10-emitter mixed-regime scenario including fixed PRI, staggered, jittered, frequency-agile, pulse-group frequency-agile, frequency-agile jittered-PRI, linear-sliding, and sinusoidal-sliding PRI signals. In the mixed-regime scenario, the total reconnaissance duration is 1 s and the slice duration is 0.1 s. The environmental pulse loss rate is fixed at 10%, and the receiver-specific loss rate increases from 0% to 40%. Both loss rates are calculated with respect to the initial theoretical number of transmitted pulses; therefore, the total loss rate is their sum, ranging from 10% to 50%. The proposed method is compared with an extended TDOA histogram method under constrained criteria, a cloud-model-based multi-station sorting method, and a Dirichlet Process Mixture Model (DPMM)-based method (Table 5). In the high pulse repetition frequency scenario, the proposed method maintains near-zero false alarms by identifying the collapse of dependent spurious peaks and suppressing them through the memory matrix, whereas the comparison methods show severe false alarms (Figs. 4 and 5). In the parameter-overlapping scenario, DTW-based reassignment improves isolated-pulse recovery, while the memory matrix suppresses spurious TDOA peaks. Their combination improves extraction reliability and reduces false alarms (Figs. 6 and 7). The ablation results verify their complementary roles: at a 50% loss rate, the memory matrix reduces the TDOA false alarm rate from 29.92% to 3.32%, DTW increases pulse extraction accuracy from 66.71% to 91.99%, and the complete method achieves a TDOA detection rate of 99.25% with a false alarm rate of 0.88% (Fig. 8). In the 10-emitter mixed-regime scenario, the proposed method achieves a favorable overall trade-off. At an overall pulse loss rate of 50%, its pulse extraction accuracy remains 92.72%, and the TDOA false alarm rate is limited to 7.75%, lower than 35.39%, 35.05%, and 37.58% for the DPMM, cloud-model, and constrained recursive histogram methods, respectively. After cross-slice trajectory association, the mean number of identity switches decreases to 3.83, compared with 10.64, 9.85, and 12.64 for the three comparison methods (Fig. 9 and Table 5).  Conclusions  A closed-loop multi-station emitter TDOA deinterleaving method is proposed for severe pulse-loss environments. By integrating multiscale KDE-based weak peak reconstruction, dynamic memory-matrix-based spurious peak suppression, DTW-based isolated pulse reassignment, and Kalman-filter-based trajectory association, the method addresses the coupled failure mechanisms caused by severe pulse loss. Simulation results demonstrate high extraction accuracy, low TDOA false alarm rates, and strong trajectory continuity in high-loss and mixed-regime scenarios. These results demonstrate the effectiveness of the method under the simulated conditions and indicate its application potential for persistent multi-station passive reconnaissance.
LLM-Aided Secure Routing Method in Industrial IoT Against Flooding Attacks
LI Jieling, XIAO Liang, WANG Chengyao, FANG Mingyang, CHEN Chen, LEI Yan
Available online  , doi: 10.11999/JEIT260400
Abstract:
  Objective  Industrial Internet of Things (IIoT) routing forwards and schedules control commands, equipment status information and sensing data to support critical tasks such as collaborative equipment control, safe system operation and environmental monitoring, but the routing process is prone to congestion and resource exhaustion under flooding attacks. Existing intelligent secure routing methods apply reinforcement learning (RL) to optimize next-hop selection based on network topology, but the heterogeneity in queue capacity and link bandwidth of IIoT terminals is often overlooked, leading to load imbalance and local congestion, and limiting performance under high load or malicious traffic attacks. Therefore, we propose a large language model (LLM)-based global situation-aware assisted secure routing method in IIoT against flooding attacks, which applies RL to optimize multi-path selection and achieve load balancing across the network.  Methods  Based on global security awareness, queue congestion of neighboring nodes, queue capacity, link bandwidth, and service types, the proposed secure routing method applies RL to optimize multi-path selection against flooding attacks. The cloud–edge large model infers global security situational awareness including global load distribution and anomalous traffic distribution based on network topology, node resource occupancy and link state information, and feeds the inference result back to IIoT terminals to construct RL states and evaluate routing policies risks. In addition, a risk-aware function is formulated to quantify the routing disruption potential by integrating end-to-end latency, packet delivery ratio and node vulnerability to attacks. An experience replay buffer that incorporates both reward and risk is constructed, where both factors are considered during routing parameter updates to guide routing policy selection, thereby balancing safe path exploration and optimization efficiency.  Results and Discussions  Simulations are conducted using 30 industrial nodes under varying configurations, including bandwidths of 5 MHz, 10 MHz, 20 MHz, and queue capacities ranging from 100 to 500 packets. The global security situational awareness is inferred by the Qwen3.5-27B-AWQ-4bit, which is deployed on a cloud–edge server equipped with dual 24 GB RTX 4090 GPUs. In each time slot, each terminal sends 5 packets of 2 KB each to the industrial gateway. A flooding attacker injects \begin{document}$ y\in \{10,20,30\} $\end{document} packets into neighboring queues per time slot to excessively consume network resources. Compared with the baseline method EEMR, the proposed secure routing method improves 39.4% packet delivery ratio, reduces 48.2% end-to-end latency and 41.1% routing energy consumption. Compared with the baseline method RLMR, the proposed secure routing method improves packet delivery ratio by a factor of 1.48, reduces end-to-end latency by 53.8% and routing energy consumption by 54.5%. This is because the proposed method leverages an LLM to infer global security situational awareness, integrating load distribution and anomalous traffic patterns to assist in selecting low-load nodes while avoiding high-load nodes, potential attack nodes, and abnormal or faulty nodes.  Conclusions  This paper proposes an LLM-based global situation-aware assisted secure routing method for IIoT against flooding attacks, which applies RL to optimize multi-path selection based on global security situational awareness including load distribution and anomalous traffic distribution. A risk assessment network is constructed based on attack behavior characteristics and service requirements to evaluate the risk level of routing performance degradation, thereby enabling risk-aware rerouting. Simulation results show that the proposed method increases the packet delivery ratio by 39.4%, reduces the end-to-end latency by 48.2% and the routing energy consumption by 41.1%.
DroneRFc-MM: Anti-UAV Multi-modal Detection Measured Dataset
YU Taosong, YANG Qianqian, HU Zhuo, LI Mingkai, WU Jiajun, SU Yifan, PAN Junyu, SHI Zhiguo, CHEN Jiming
Available online  , doi: 10.11999/JEIT260889
Abstract:
Multimodal data fusion can effectively improve the generalization, robustness, and scene adaptability of counter-unmanned aerial vehicle (UAV) detection systems. To address the limitations of existing counter-UAV datasets in sensing modalities, UAV models, and annotation granularity, the paper releases DroneRFc-MM, a multimodal counter-UAV detection dataset. DroneRFc-MM synchronously collects data from six types of sensors, including a pan-tilt-zoom camera, a wide-angle (fisheye) camera, radio-frequency antennas, LiDAR, millimeter-wave radar, and a microphone array. The dataset covers six consumer UAV models and provides fine-grained annotations such as model category, position, attitude, and velocity. It supports multiple tasks, including target detection, UAV model recognition, and motion-direction reasoning, and is accompanied by user-friendly sample extraction tools. Finally, as a demonstration of dataset usage, the paper evaluates the performance of recent Qwen-series models on UAV flight-direction reasoning.Objective: The work aims to build a more comprehensive benchmark for multimodal counter-UAV detection. Existing datasets often cover limited sensing modalities and UAV models, making it difficult to represent realistic low-altitude scenarios and diverse target characteristics. Their annotations are also typically coarse, such as category labels or bounding boxes, and thus cannot fully support downstream tasks requiring precise spatial, motion, and cross-modal information. DroneRFc-MM addresses these gaps by providing synchronized multimodal data, richer UAV coverage, and fine-grained annotations for detection, model recognition, trajectory analysis, motion reasoning, and multimodal fusion evaluation.Methods: The DroneRFc-MM dataset was synchronously captured via six heterogeneous sensors—including a pan-tilt-zoom(PTZ) camera, fisheye camera, radio frequency(RF) antenna, LiDAR, millimeter-wave radar, and microphone array—on an open rooftop of a university in Zhejiang Province. Featuring a representative urban low-altitude scenario, this dataset contains data recordings of six consumer-grade DJI drones. All devices were time-synchronized via network timestamp, and drones flew in rectangular and vertical reciprocating trajectories within 20–60 meters. Fine-grained annotations including drone type, position, attitude and velocity were provided. For flight direction reasoning task, 5-second multimodal clips were generated: videos for cameras and RF spectrograms, audio for microphones, text coordinates for radar point clouds. Zero-shot inference was conducted on Qwen 3.6-Plus and Qwen 3.5-Omni-Plus models with unified prompts, and accuracy and inference time were evaluated by comparing predicted directions with ground truth calculated from drone positioning data.Conclusions: Experiments on Qwen-series models show that general-purpose multimodal large models can capture weak motion-related features from drone-related videos, audio, RF spectrograms and point clouds, but only achieve limited flight-direction reasoning accuracy ranging from 20% to 30%. Meanwhile, the long inference time and unstable latency make them unable to satisfy the real-time and stability demands of practical low-altitude surveillance systems. These results demonstrate that domain-specific pre-training, supervised fine-tuning, knowledge enhancement and lightweight inference optimization are essential for deploying multi-modal LLMs in real anti-UAV detection scenarios. Future work will focus on expanding the dataset scale and enriching application scenarios to support the development of intelligent and efficient low-altitude airspace management systems.
Accelerated Broadband Electromagnetic Scattering Analysis via ACA-Driven Measurement Matrix Interpolation
WANG Zhonggen, WU Chenggang, NIE Wenyan, SUN Yufa
Available online  , doi: 10.11999/JEIT260392
Abstract:
  Objective  Broadband electromagnetic scattering analysis is widely used in radar target recognition, stealth technology, and microwave imaging. Although the Method of Moments (MoM) provides high computational accuracy, it incurs substantial computational and memory costs for electrically large or geometrically complex targets because full impedance matrices must be constructed and solved. Existing acceleration techniques, including the MultiLevel Fast Multipole Method (MLFMM) and Adaptive Cross Approximation (ACA), reduce the computational burden but still require repeated matrix construction and equation solving at every frequency during wideband analysis. Methods such as Asymptotic Waveform Evaluation (AWE), Model-Based Parameter Estimation (MBPE), and impedance matrix interpolation have been proposed to reduce this redundancy. However, AWE is prone to error accumulation over wide frequency bands, MBPE requires expensive initial sampling, and conventional impedance matrix interpolation still requires the computation of full high-dimensional impedance matrices at the sampling frequencies. More recently, Compressive Sensing Method of Moments (CS-MoM) and its extension, CS-HBFM, have improved wideband analysis by employing Hyper-Basis Functions (HBFs). By constructing Characteristic Mode Basis Functions (CMBFs) only once at the highest frequency, CS-HBFM eliminates repeated basis-function generation. Nevertheless, existing CS-HBFM methods rely on nondeterministic random or uniform sampling, require expensive large-scale matrix-vector products, and repeatedly reconstruct and solve impedance equations throughout the frequency sweep.  Methods  A CS-ACA-MMI framework is proposed for broadband electromagnetic scattering analysis by combining dual ACA decomposition with Measurement Matrix Interpolation (MMI). First, CMBFs are constructed at the highest frequency, and dominant HBFs are selected according to the Modal Significance (MS) criterion. ACA is then applied to the full impedance matrix to extract deterministic row indices corresponding to the dominant Rao-Wilton-Glisson (RWG) basis functions. These indices are reused throughout the frequency band, eliminating nondeterministic sampling and repeated index extraction. Second, four sampling frequencies are selected using Chebyshev-Lobatto nodes. Low-dimensional measurement matrices are constructed directly from the extracted row indices, avoiding the generation of full high-dimensional impedance matrices. The measurement impedance elements at the sampling frequencies are corrected according to the geometric distance, interpolated to the target frequency, and then restored to the actual measurement impedance elements, thereby eliminating repeated construction of measurement matrices during frequency sweeping. Third, ACA is applied to the far-field component of the interpolated measurement matrix, converting large-scale matrix-vector products into low-dimensional matrix multiplications. The near-field sensing matrix is obtained directly by multiplying the measurement matrix by the basis functions, enabling rapid construction of the complete sensing matrix. Finally, the dense linear system is transformed into an overdetermined system under the compressive sensing framework, and the least-squares method is used to reconstruct the current coefficients, from which the broadband Radar Cross Section (RCS) is calculated. The Root Mean Square Error (RMSE) is used to evaluate numerical accuracy. Three representative targets, namely a cylinder, a slotted cone, and an almond, are analyzed. Broadband RCS, numerical accuracy, total computation time, and single-frequency measurement-matrix memory consumption are compared with those obtained using MoM and CS-HBFM to validate the proposed framework.  Results and Discussions  Three numerical examples, including a perfect electric conductor cylinder, a slotted cone, and an almond, are used to validate the proposed CS-ACA-MMI framework. The ACA-extracted row indices are concentrated near geometric boundaries and structural junctions, demonstrating the physical validity of the deterministic sampling strategy (Fig. 2). Parametric studies show that appropriate ACA thresholds and four sampling frequencies provide the best balance between computational efficiency and numerical accuracy (Figs. 35). The broadband RCS predicted by the proposed framework agrees closely with the MoM results over the entire frequency band (Figs. 68), and the RMSE remains low, demonstrating high numerical accuracy. Compared with CS-HBFM, the proposed framework reduces the total computation time by 93.4% for the cylinder, 96.7% for the slotted cone, and 80.9% for the almond (Table 2). These improvements result from deterministic index reuse, MMI, and dual ACA acceleration, which substantially reduce the computational cost of broadband frequency-sweeping analysis.  Conclusions  A CS-ACA-MMI framework is proposed by integrating ACA with MMI for efficient broadband electromagnetic scattering analysis. The proposed framework eliminates repeated matrix construction and equation solving during frequency sweeping while overcoming the nondeterministic sampling strategy and the high computational and memory costs of conventional CS-HBFM. Dominant row indices extracted by ACA at the highest frequency provide a deterministic measurement-matrix construction strategy and a stable physical basis for broadband interpolation. By shifting the interpolation target from full impedance matrices to low-dimensional measurement matrices, the computational complexity and redundant matrix construction are substantially reduced. A second ACA decomposition further accelerates sensing-matrix construction by converting large-scale matrix-vector products into low-dimensional matrix multiplications. Numerical results demonstrate that the proposed framework achieves numerical accuracy comparable to that of MoM while reducing total computation time by more than 80% and decreasing single-frequency measurement-matrix memory consumption by up to 65%. Because only the measurement matrices at four sampling frequencies need to be stored, the overall memory requirement is further reduced.
Construction and Performance Analysis of Optimal Low-Hit-Zone Frequency Hopping Sequence Sets
TIAN Xinyu, CHEN Xiaoyu, ZHANG Jitao
Available online  , doi: 10.11999/JEIT260343
Abstract:
  Objective  ElectroMagnetic Interference (EMI) is a critical factor limiting the reliability of synchronization systems. Existing Fifth-Generation (5G) synchronization schemes extensively employ Zadoff-Chu (ZC) sequences to distinguish users through cyclic shifts. However, finite sequence lengths and limited orthogonal resources create substantial capacity bottlenecks in high-density access scenarios. To address these challenges, this paper investigates the problem from two perspectives. At the system level, a synchronization framework is developed by integrating Frequency Hopping (FH) with ZC sequences. By jointly exploiting code, time, and frequency-domain resources, the proposed framework improves concurrent access capability for local clusters while enhancing robustness against complex EMI through frequency diversity. At the sequence-design level, a class of multi-subset Low-Hit-Zone (LHZ) Frequency Hopping Sequence (FHS) sets is constructed to provide an efficient sequence allocation scheme for local-cluster synchronization.  Methods  Based on the theoretical framework proposed by Cai et al., the sequence mapping mechanism is reconstructed, and a disjoint Cyclic Perfect Mendelsohn Difference Family (CPMDF) is introduced to construct FHS sets that are optimal with respect to the Peng-Fan bound. The generating units are further expanded through Cartesian products, and a column-incoherent partitioning strategy is proposed to construct multi-subset LHZ FHS sets. It is proved that every nonempty subset satisfies the Peng-Fan-Lee bound with equality. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure. At the system level, an FH-ZC synchronization architecture is developed by combining predefined FH patterns with the frequency-domain correlation properties of ZC sequences for subband signal detection. A Peak-to-SideLobe Ratio (PSLR) decision metric and an early-termination strategy are adopted to evaluate synchronization preamble detection under interference. Furthermore, a multi-user simulation model is established to evaluate synchronization detection performance under accumulated co-channel collisions and EMI.  Results and Discussions  The proposed construction generates an FHS set that is optimal with respect to the Peng-Fan bound and a class of multi-subset LHZ FHS sets in which every nonempty subset is optimal with respect to the Peng-Fan-Lee bound. Example 2 demonstrates the construction procedure and the intra-subset and inter-subset Hamming correlation properties of the proposed multi-subset LHZ FHS sets. Table 1 shows that, under the same frequency-resource constraints, the proposed construction generates more sequences than existing methods under the compared parameter settings, indicating higher sequence-resource utilization. Table 2 compares the parameters of the proposed sequence sets with representative constructions reported previously and demonstrates that the proposed multi-subset optimal sequence family provides a new parameter combination. To the best of our knowledge, an optimal sequence family with a multi-subset structure has not been reported previously. Figures 2 and 3 demonstrate that the proposed FH-ZC synchronization architecture achieves a higher synchronization detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI. Figure 4 shows that the synchronization detection probability decreases as the number of active users increases because accumulated co-channel collisions degrade synchronization performance. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure characterized by strong intra-cluster competition and weak inter-cluster coupling.  Conclusions  To satisfy the sequence-capacity requirements of massive-access scenarios, this paper proposes a class of multi-subset LHZ FHS sets. By expanding the generating sequence sets through Cartesian products and partitioning subsets using a column-incoherent strategy, the proposed construction achieves both a large family size and optimal LHZ performance. The proposed multi-subset structure is well suited to local-cluster synchronization and substantially improves sequence family size and sequence-resource utilization, thereby providing a richer sequence resource pool for high-density multi-user systems. Simulation results under the considered physical-layer model demonstrate that the proposed LHZ FHS subsets reduce the effect of frequency collisions during multi-user synchronization detection. Furthermore, the FH-ZC synchronization scheme achieves a higher synchronization preamble detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI.
A Phase Transition Obstacle Avoidance Method for UAV Swarms Driven by Multistable Potential Fields
HE Ming, CHEN QiYang, HAN Wei, PAN Fan, MA YiSong
Available online  , doi: 10.11999/JEIT260357
Abstract:
  Objective  Unmanned Aerial Vehicle (UAV) swarms have demonstrated considerable potential for complex missions, such as search, surveillance, and disaster response, because of their distributed coordination and robustness. However, in dynamic environments with dense obstacles and rapidly changing risks, conventional swarm control methods often exhibit discontinuous behavior switching and control chattering, which reduce system stability and coordination efficiency. Existing approaches, including threshold-based switching and Artificial Potential Field (APF) methods with fixed potential weights, rely on abrupt transitions between behavioral modes, leading to oscillatory responses. To address these limitations, a phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. Swarm behavior evolution is modeled as a continuous phase transition process within a unified potential field framework, enabling smooth and adaptive transitions between formation flight and obstacle avoidance.  Methods  An environmental risk assessment model is first established by integrating static obstacle risk, dynamic obstacle risk, and inter-agent proximity risk. A distributed consensus protocol is then employed to establish global risk consensus. Subsequently, a morphology factor is generated through nonlinear mapping of the global risk consensus and is used as an order parameter to characterize the macroscopic swarm state. A unified time-varying potential field, comprising formation, obstacle avoidance, and navigation potentials, is constructed, and the relative weights of these potentials are continuously adjusted by the morphology factor. When the risk level is low, the system exhibits a monostable structure dominated by the formation and navigation potentials. As the risk increases, the potential field continuously evolves into a multistable structure dominated by the obstacle avoidance potential, thereby enabling distributed obstacle avoidance. A distributed consensus control law based on the negative gradient of the unified potential field is further developed. A damping term is incorporated to dissipate system energy and improve stability, while a dynamic compensation term addresses nonlinear dynamics. The control law depends only on local information, ensuring good scalability. The global uniform ultimate boundedness of the closed-loop system is established using Lyapunov theory.  Results and Discussions  Simulation results demonstrate that the proposed method enables the swarm to maintain a compact solid-phase swarm formation in low-risk regions and to transition smoothly to a dispersed liquid-phase swarm configuration when obstacles are encountered, followed by rapid formation recovery after obstacle avoidance. The pitch and roll angles of each UAV vary smoothly without abrupt changes, and both the UAV-to-obstacle distance and the inter-UAV separation remain above the prescribed safety threshold throughout the flight, ensuring collision-free operation. Statistical results obtained from 20 independent simulation runs show that, compared with the threshold-switching method, the proposed method reduces the rate of control input variation by approximately 26% and decreases the peak control input by approximately 18%. Compared with the bio-inspired diversion method, the average formation recovery time after obstacle avoidance is reduced by approximately 16%. Ablation experiments further demonstrate that removing the morphology-driven phase transition mechanism significantly increases trajectory oscillation and control oscillation, confirming the critical role of the multistable continuous phase transition mechanism in maintaining smooth swarm motion. In complex narrow-channel environments, the proposed method effectively avoids the local minimum problem encountered by conventional APF methods and generates smoother flight trajectories with substantially reduced oscillation.  Conclusions  A phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. By introducing a morphology factor and constructing a unified potential field framework, swarm behavior evolution is represented as a continuous phase transition process. The distributed control law enables smooth behavioral transitions while maintaining system stability and scalability. Simulation results demonstrate that the proposed method achieves better safety, smoother control, and higher coordination efficiency than conventional methods.
Energy Efficiency Analysis of Discrete Phase-Shifted Active RIS Enhanced Communication Systems
SHU Feng, LIN Zhiyuan, ZHENG Weihai, WANG Yan, JIANG Hao, WANG Jiangzhou
Available online  , doi: 10.11999/JEIT260462
Abstract:
  Objective  Active Reconfigurable Intelligent Surface (RIS) enhances wireless communication performance by integrating radio frequency amplifiers to mitigate the multiplicative fading inherent to passive RIS. However, amplification noise and additional power consumption are introduced. Furthermore, high-precision digital phase control at the base station incurs considerable communication overhead. Employing low-precision phase shifters is therefore an effective approach for practical RIS deployment. Therefore, characterizing the Energy Efficiency (EE) performance of active RIS-assisted communication systems and quantifying the effect of finite-bit phase quantization errors on EE are essential for system design and practical implementation. To this end, a discrete phase-shifted active RIS-assisted communication system over Rayleigh fading channels is investigated. The EE loss caused by phase quantization errors is analyzed, approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE are derived, and the relationship between RIS EE and user EE is established, providing theoretical guidance for the practical deployment of active RIS.  Methods  Based on the law of large numbers and Taylor series expansion, closed-form expressions for the user EE loss and its approximation are derived. The effects of system parameters on EE are investigated by expressing EE as explicit univariate functions. Ferrari’s method and the Lambert W function are then employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE. Finally, the relationship between RIS EE and user EE is established using the law of large numbers and the Lambert W function.  Results and Discussions  User EE is expressed as a function of six parameters: the number of quantization bit (\begin{document}$ k $\end{document}), power allocation factor (\begin{document}$ \beta $\end{document}), the number of RIS elements (\begin{document}$ N $\end{document}), the total power sum of base station and active RIS (\begin{document}$ {P}_{\text{t}} $\end{document}), the noise at active RIS (\begin{document}$ \sigma _{\text{r}}^{2} $\end{document}), and the noise at user (\begin{document}$ \sigma _{\text{u}}^{2} $\end{document}). First, the EE loss decreases as \begin{document}$ k $\end{document} increases. When \begin{document}$ k $\end{document}=3, the difference between the approximate EE loss and the lossless case is less than 0.026 8 Mbit/J, while the difference between the EE loss and the lossless case is less than 0.026 5 Mbit/J (Fig. 3). Therefore, 3- to 4-bit discrete phase shifters achieve performance close to that of continuous phase shifters. Second, user EE exhibits a unimodal dependence on both \begin{document}$ \beta $\end{document} and \begin{document}$ N $\end{document}. The approximate optimal solution for \begin{document}$ \beta $\end{document} differs from the exact optimal solution obtained by the Dinkelbach algorithm by less than 0.01 (Fig. 4), whereas the approximate and exact optimal solutions for \begin{document}$ N $\end{document} are identical (Fig. 5), demonstrating the high accuracy of the proposed approximations. Third, user EE exhibits a unimodal trend as \begin{document}$ {P}_{\text{t}} $\end{document} increases. Higher phase quantization precision produces a higher EE peak while requiring a lower optimal \begin{document}$ {P}_{\text{t}} $\end{document} to achieve the maximum EE (Fig. 6). In addition, user EE decreases as both \begin{document}$ \sigma _{\text{r}}^{2} $\end{document} and \begin{document}$ \sigma _{\text{u}}^{2} $\end{document} increase. User EE is more sensitive to the amplification noise introduced at the RIS, indicating that reducing the RIS noise power yields a greater EE improvement (Fig. 7). Finally, user EE first increases and then decreases sharply to zero as RIS EE increases. The signal-to-noise ratio at the RIS is identified as the key factor governing the relationship between RIS EE and user EE (Fig. 8).  Conclusions  The EE performance of active RIS-assisted wireless networks employing discrete phase shifters over Rayleigh fading channels is investigated. First, closed-form expressions are derived for the user EE in the lossless case, the lossy case, and the approximate-loss case. Simulation results demonstrate that 3- to 4-bit discrete phase shifters closely approach the performance of continuous phase shifters. Next, explicit functions describing the effects of key system parameters on user EE are established. Ferrari’s method and the Lambert W function are employed to derive approximate optimal solutions for the power allocation factor and the number of RIS elements that maximize EE, and both exhibit negligible errors relative to the exact solutions. Finally, the relationship between RIS EE and user EE is established, demonstrating that user EE initially increases and subsequently decreases to zero as RIS EE increases.
A Spatial-temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming
Available online  , doi: 10.11999/JEIT260512
Abstract:
  Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is essential for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often reduce the stability of target region segmentation, leading to centroid deviation, trajectory jitter, missed detections, and trajectory discontinuity. To address these challenges, a Spatial-Temporal Collaborative Optimization Method is proposed for stable and continuous grab trajectory extraction. While maintaining high inference speed, the proposed method improves both trajectory extraction accuracy and trajectory stability, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates Spatial Representation Enhancement (SRE) and Temporal CONSistency constraint (TCONS). First, CBAM, BiFPN-lite, and shallow feature aggregation are incorporated to improve target-background separability, enhance multi-scale feature representation, and preserve boundary details. TCONS is then imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss to suppress the temporal accumulation of local errors. In addition, a stage-wise training strategy with warm-up epochs and a joint optimization objective is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The proposed method achieves the best segmentation performance on the real portal crane grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also improves performance on DAVIS2016 while maintaining comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, and decreases the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further demonstrate that, under different occlusion ratios, the proposed method maintains good region integrity and continuous target extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementary effects of SRE and TCONS, whereas parameter analysis shows that a TCONS weight of 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  A Spatial-Temporal Collaborative Optimization Method is proposed to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves high segmentation accuracy and stable trajectory extraction on DAVIS2016 and the real portal crane grab dataset, while maintaining comparable segmentation performance on SegTrackV2. It also exhibits strong continuous target extraction capability under occlusion without significantly sacrificing inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further improve the robustness and applicability of the proposed method in real-world environments.
An SO(3)-Manifold-Constrained Registration Method for Twin-Fisheye Panoramic Images
WANG Zhuopeng, LIN Shanling, LIN Jianpu, LÜ Shanhong, LIN Zhixian
Available online  , doi: 10.11999/JEIT260798
Abstract:
  Objective  Twin-fisheye cameras provide near-360° coverage with low hardware complexity and are widely used in immersive imaging, surveillance, and mobile robotics. Their panoramic output depends on registration over a narrow overlapping band, so geometric accuracy and temporal consistency directly affect seam quality and video smoothness. After the two fisheye views are unfolded into the Equirectangular Projection (ERP), three coupled problems arise. First, the near-co-centric lens pair is ideally related by a pure rotation R ∈SO(3), whereas a conventional 8-Degree-of-Freedom (DoF) homography introduces five redundant parameters that may couple with matching noise. Second, ERP sampling is nonuniform with latitude, so identical pixel residuals do not represent identical spherical angular errors. Third, the cyclic ±π longitude boundary splits structures that are continuous on the sphere and weakens correspondences around the seam. Existing planar pipelines and generic learned matchers rarely combine these constraints under a unified rotation-referenced evaluation. This study therefore develops a lightweight registration framework that explicitly exploits spherical rotation geometry while addressing ERP boundary discontinuity, temporal fluctuation, and long-tail residuals.  Methods  The proposed framework contains three modules (Fig. 1). First, an overlapping-band Region-of-Interest (ROI) is cropped around the ERP seam and rearranged with modulo-W wrap-around (Fig. 2). A default longitude half-width of ±15° and an approximately 3° margin on each side preserve cross-boundary feature continuity while restricting the search region. XFeat detects, describes, and matches features under a fixed Top-K budget. Second, the two-dimensional matches are restored to global ERP coordinates, mapped to unit-sphere direction vectors, and processed by rotation-only SO(3)-RANSAC. Spherical angular residuals are used as the inlier criterion with a 0.8° threshold, a maximum of 2000 iterations, confidence 0.999, and at least 12 inliers; the iteration bound is updated adaptively. All inliers are then used for Kabsch/SVD closed-form rotation re-estimation, which reduces the randomness of a minimal sample while preserving the SO(3) constraint. Third, a local increment on the Lie algebra so(3) is optimized under a Huber loss by the Levenberg–Marquardt algorithm. The refinement is triggered only when the inlier-residual P95 exceeds 0.90° and the inlier ratio is below 0.58, thereby concentrating nonlinear optimization on difficult image pairs.  Results and Discussions  Experiments are conducted on PanoraMIS Sequences 3 and 4 under a unified relative inter-frame rotation protocol (Table 1). The proposed method achieves a 97.10% success rate, a 0.549° P95 angular residual, and a 0.313° temporal-stability error. Compared with SuperPoint+LightGlue, the P95 and temporal-stability errors are reduced by 22.8% and 77.3%, respectively. Compared with Efficient LoFTR, peak GPU memory and runtime are reduced by 52.7% and 60.9%, although Efficient LoFTR retains the lowest overall P95. Under a unified SO(3)-RANSAC back-end (Table 2), XFeat provides the largest average inlier count of 625.4 and the lowest temporal-stability error of 0.313° at 38.04 ms. The ablation and sensitivity results (Tables 34) show that the ±15° ROI reduces the P95 from 0.720° for the full ERP to 0.600°. Replacing H-RANSAC with SO(3)-RANSAC reduces temporal instability from 0.931° to 0.313°, a 66.3% reduction, while increasing runtime from 24.91 ms to 35.24 ms. Adaptive refinement operates on approximately one third of the image pairs and improves both P95 and temporal stability with lower overhead than always-on refinement; its five-seed mean and median trigger rate are both 36.23%. A Top-K budget of 2048 reaches the saturated accuracy level, because increasing the budget to 4096 yields no further P95 or stability improvement. Five fixed-seed repetitions produce standard deviations no greater than 0.011° for P95 and stability, indicating that the main conclusions are insensitive to RANSAC randomness. In a 5×5 threshold sweep, the maximum changes in P95 and inter-frame rotation jitter within the central neighborhood are 4.40% and 0.037%, respectively, and changing the robust-error truncation from 3° to 2° or 5° does not alter the relative ranking. On the more difficult Sequence 4, characterized by weak texture and unstable overlap, the proposed method obtains 919.7 average inliers and a P95 of 0.383°, the lowest among the evaluated learning-based matchers, although robust-estimation time increases. With an identical standardized stitching back-end, it produces a lower seam-band gradient than H-RANSAC in the small-rotation example (26.90 versus 28.71) and a lower truth-referenced temporal-stability error (0.313° versus 0.931°; Fig. 6). On outdoor Sequence 7-L2, 324 of 346 correspondences are retained, yielding a 93.6% inlier ratio and a 0.524° P95 residual (Fig. 5). Because sequence-specific calibration is unavailable and the fixed inter-lens baseline may cause depth-dependent parallax, this result serves only as a diagnostic consistency check, not as evidence of absolute pose accuracy.  Conclusions  By restoring feature continuity across the ERP boundary, replacing the redundant planar homography with an explicit SO(3) rotation model, and selectively refining difficult image pairs on so(3), the proposed method balances registration accuracy, temporal consistency, and resource cost. It provides a lightweight front-end for twin-fisheye panorama stitching. The current evaluation is limited to pairwise registration on a small number of sequences; future work will address translation compensation, multi-frame global optimization, end-to-end integration with seam finding, exposure compensation, and blending, as well as generalization across additional platforms, dynamic scenes, and illumination conditions.
Research on Channel Multipath Prediction Based on Environmental Graph
ZHANG Zhaoling, JIN Jing, ZHAO Jingbo, YU Li, CAI Yichen, MA Liang, ZHANG Jianhua
Available online  , doi: 10.11999/JEIT260416
Abstract:
  Objective  Environment-aware channel modeling requires a structured representation that can connect physical objects in a propagation environment with the resulting multipath topology. However, conventional data-driven methods generally treat environmental information as unstructured global features and therefore have difficulty representing the interactions among the transmitter (Tx), receiver (Rx), and surrounding scatterers. This study investigates whether an environmental graph can provide an effective intermediate representation for identifying propagation-relevant scatterers and predicting candidate multipath structures.  Methods  An environmental graph was constructed by representing the Tx, Rx, and scatterers as graph nodes. The spatial distance and visibility between nodes were encoded as edge features to describe their geometrical relationships. Based on this representation, an edge-aware graph isomorphism network, termed ScatterGNN, was developed to extract structural features and identify effective scatterers involved in signal propagation. Candidate single- and multi-bounce paths were subsequently generated from the detected scatterers. A path-ranking network, PathRankingNet, was then designed to estimate the validity scores of candidate paths and rank them using a listwise ranking loss. The proposed framework was evaluated in a controlled indoor Industrial Internet of Things scenario generated using Wireless InSite. The scenario covered an area of 150 m × 63 m × 22 m and contained 3,381 Rx sampling locations.  Results and Discussions  For effective scatterer detection, the proposed method achieved an average precision of 0.9579, while 77.1% of the test samples obtained complete detection of all effective scatterers. For candidate path prediction, the model converged stably after approximately 30 training epochs. Precision@3 ranged from 0.60 to 0.64, Recall@3 ranged from 0.42 to 0.45, and Hit@3 ranged from 0.90 to 0.92. These results indicate that the environmental graph preserves useful structural information related to propagation-path topology. In particular, the model was able to retain at least one reference propagation path among the three highest-ranked candidates for more than 90% of the test samples.  Conclusions  The proposed framework provides a graph-based approach for transforming environmental geometry into structured representations of effective scatterers and candidate propagation paths. Rather than replacing ray tracing or channel measurements, the method is intended to reduce the candidate search space before detailed path-parameter calculation or channel reconstruction. The current results demonstrate its feasibility within a single simulated environment and primarily reflect its ability to approximate the path-topology labels generated by Wireless InSite. Further validation using independent environments, measured channel data, and path-level parameters such as power, delay, and phase is required before its cross-scenario generalization and practical applicability can be established.
Hierarchical Prototype Learning with Shared Subspace Factorization for Generalizable Deepfake Detection
PENG Shufan, LU Tianliang, HE Chunhao, ZHANG Lu, ZHAO Kai
Available online  , doi: 10.11999/JEIT260426
Abstract:
  Objective  Deepfake detectors often exhibit substantial performance degradation when applied to unseen manipulation methods, cross-dataset distribution shifts, diffusion-generated faces, or common image degradations. This limitation is critical in forensic applications because the generation process, data source, and post-processing history of a questioned sample are usually unknown. Existing approaches often represent the entire fake class using a single feature center, despite the heterogeneous patterns produced by different generation methods, source datasets, and processing conditions. As a result, transferable forensic evidence may become entangled with mode-specific artifacts, weakening generalization to unknown domains. To address this issue, a hierarchical prototype learning framework with shared subspace factorization (HPL-SF) is proposed. The framework exploits within-class diversity to estimate a low-rank structure shared across latent fake modes and adaptively uses this structure for individual test samples.  Methods  A pretrained DINOv2 Vision Transformer (ViT-L/14) is used as the backbone. Its original parameters are frozen, and low-rank adaptation modules are inserted into the query and value mappings of the self-attention layers for parameter-efficient training. All features and prototypes are L2-normalized so that cosine similarity is used consistently for prototype assignment, binary classification, and test-time representation adaptation. HPL-SF comprises three successive stages (Fig. 1). First, one real prototype and multiple mode-specific fake prototypes are maintained in the normalized feature space. Each training sample is softly assigned to the fake prototypes according to its cosine similarity to each prototype, and the weighted prototypes are aggregated into a sample-adaptive fake representation. The resulting response distribution provides an observable representation of latent within-class modes without requiring forgery-source labels. A binary classification loss, a sample–prototype contrastive loss, and a prototype-diversity loss are jointly minimized to separate real and fake samples, improve sample–prototype alignment, and prevent the fake prototypes from collapsing into a single direction. Second, the normalized mode-specific fake prototypes are arranged as a prototype matrix. Singular value decomposition is applied to this matrix, and the largest gap between consecutive singular values determines the dimension of the shared fake subspace. Each fake prototype is decomposed into a projection within this subspace and an orthogonal residual. Only the shared projections are aggregated to update the shared fake prototype, whereas the residuals retain mode-dependent information and are excluded from this update. The real and shared fake prototypes are updated using exponential moving averages. Gradients are not propagated through subspace construction, dimension selection, or the updates of these semantic prototypes. Third, the backbone, adaptation modules, prototypes, and subspace basis are fixed during inference. A test feature is compared with the real and shared fake prototypes, and their relative responses determine a sample-specific adaptation weight. The original feature is blended with its projection onto the shared subspace, normalized again, and classified according to its cosine similarities to the two semantic prototypes. Thus, the estimated shared structure can be exploited on a per-sample basis without updating model parameters during testing.  Results and Discussions  The method is evaluated under cross-dataset, cross-forgery-type, diffusion-forgery, repeated-run, image-degradation, ablation, and mechanism-analysis protocols. Across seven unseen datasets, HPL-SF achieves the highest area under the receiver operating characteristic curve (AUC) on every dataset and an average AUC of 91.67%, exceeding the second-highest average by 2.10 percentage points (Table 1). When trained on FaceForensics++ and directly evaluated on the Diffusion Facial Forgery dataset, HPL-SF obtains the highest AUC on the text-to-image, image-to-image, face-swapping, and face-editing subsets, with an average AUC of 84.39% (Table 2). In four cross-forgery-type settings on FaceForensics++, the average accuracy and AUC reach 85.29% and 92.41%, respectively, both ranking first among the compared methods. In the high-quality-compression setting where DeepFakes is held out for testing, HPL-SF trails the best method by only 0.91 percentage points in accuracy and 0.39 percentage points in AUC (Table 3). Repeated experiments with multiple random seeds yield the highest mean values for all four aggregate metrics, with standard deviations ranging from 0.38 to 0.62 percentage points (Table 4). Under five severity levels of compression, blur, and noise on the Deepfake Detection Challenge Preview dataset, HPL-SF achieves the best or joint-best AUC at most severity levels and remains comparatively stable under moderate and severe degradations (Fig. 2). All six ablated variants perform worse than the complete model on the four unseen test sets, indicating that mode-specific prototype learning, sample–prototype contrast, prototype diversity, shared subspace factorization, and test-time representation adaptation make complementary contributions (Fig. 3). Performance generally improves as the number of mode-specific fake prototypes increases and begins to plateau when the number reaches 10. The largest spectral gap occurs between the fifth and sixth singular values; accordingly, the shared dimension is set to 5. With this setting, HPL-SF attains a diffusion-forgery AUC of 84.39% ± 0.62%, compared with 79.68% for direct mean aggregation and 81.56% ± 0.80% without test-time representation adaptation (Figs. 4(a)4(c)). Linear-probe results further show that the shared component provides stronger discrimination in unknown domains while retaining less domain-identifying information than the mode-specific residual, supporting the intended separation of transferable and domain-related cues (Fig. 4(d)). The t-distributed stochastic neighbor embedding visualization reveals clearer real–fake separation and closer same-class distributions across data sources (Fig. 5). Prototype-allocation statistics indicate differentiated yet balanced use of the fake prototypes. The adaptation weights are higher for fake samples, particularly for correctly detected fake samples, whereas misclassified samples lie closer to the balanced-response line (Fig. 6). The remaining false positives are mainly associated with low resolution, compression, filters, or occlusion, whereas false negatives usually contain weak or high-quality forgery traces (Fig. 7).  Conclusions  HPL-SF organizes generalizable deepfake detection as a progression from observing within-class variation to estimating and exploiting shared structure. The experimental evidence indicates that separating shared projections from mode-dependent residuals provides transferable decision information under dataset shifts, unseen forgery types, diffusion-generated manipulations, and common image degradations. The framework requires no parameter updates during testing and adaptively exploits the estimated shared structure based on each sample’s relative responses to the real and shared fake prototypes. Nevertheless, errors remain when degradations in real images resemble forgery artifacts or when high-quality forgeries contain only weak traces. Future work will extend the framework to temporal prototypes for video, multimodal forensic evidence, and adaptive discrimination under broader open-world conditions.
Multi-projection Plane InISAR 3D Reconstruction Method for Complex Moving Ship Targets
LI Ning, NIU Jinfa, WANG Weibin, HU Xingwang, WU Lin
Available online  , doi: 10.11999/JEIT251268
Abstract:
  Objective  Interferometric Inverse Synthetic Aperture Radar (InISAR) is a three-dimensional (3D) reconstruction technique for non-cooperative targets. However, the complex 3D rotational motion of ship targets causes unstable Doppler frequency variation. Inverse Synthetic Aperture Radar (ISAR) imaging also inevitably suffers from target overlap and occlusion. These factors make high-precision and complete 3D reconstruction under a single projection plane difficult. Therefore, a multi-projection plane InISAR 3D reconstruction method for complex moving ship targets based on point cloud fusion is proposed. The method supplements target 3D information through efficient and high-precision point cloud registration and fusion, thereby significantly improving 3D reconstruction quality.  Methods  This method fully exploits the advantage of multi-plane observation enabled by the severe motion of ship targets. The ship centerline is extracted, and the vertical rotation vector is estimated by Principal Component Analysis (PCA) to select the optimal imaging times corresponding to different Imaging Projection Planes (IPPs). ISAR imaging and InISAR 3D reconstruction are then completed. In addition, a point cloud fusion algorithm that combines Weighted Random Sample Consensus (RANSAC) and hierarchical Iterative Closest Point (ICP) is proposed. The random sampling process is optimized through a feature stability weighting strategy, which enables efficient extraction and matching of corresponding feature points in InISAR images and achieves high-precision point cloud fusion under multiple IPPs.  Results and Discussions  Experimental results show that the proposed method significantly improves reconstruction accuracy and target completeness. For simulated ship point-target data, Fig. 7 shows excellent results, with a significant reduction in reconstruction error. Signal-to-Noise Ratio (SNR) analysis shows that the quality of 3D fusion imaging improves steadily as the SNR increases from –10 dB to 10 dB, and robust fusion performance is maintained even under low-SNR conditions. For simulated destroyer Radar Cross Section (RCS) data, the method achieves strong registration performance. The detail recovery and structural integrity of the fused image are also significantly improved, effectively addressing the incomplete reconstruction of 3D information caused by scattering-point overlap and occlusion.  Conclusions  To address the low reconstruction accuracy and information loss caused by target rotation, overlap, and occlusion in traditional InISAR methods for 3D reconstruction of complex moving ship targets, a multi-IPP InISAR 3D reconstruction method based on point cloud fusion is proposed. The method uses a PCA-based optimal imaging time selection strategy. Weighted RANSAC and hierarchical ICP algorithms are then applied to achieve efficient and high-precision registration and fusion of InISAR point clouds under multiple IPPs, thereby producing high-quality 3D reconstruction results. Multi-scenario experiments are conducted by constructing both a ship model with ideal scattering points and an electromagnetic simulation RCS model with occlusion effects. The results verify the accuracy of the proposed method under ideal conditions and demonstrate its applicability in complex real-world scenarios.
Design and Verification of Robust Modulation Recognition Framework Under Blind Adversarial Attacks
ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260019
Abstract:
  Objective  Deep learning-based Automatic Modulation Recognition (AMR) models demonstrate strong performance in non-cooperative communication systems such as cognitive radio and spectrum monitoring. However, deep learning models remain vulnerable to adversarial attacks. In these attacks, imperceptible perturbations lead to severe misclassification and create security risks. Existing defense methods, including adversarial training, often rely on prior knowledge of specific attacks. They also introduce considerable computational overhead and reduce accuracy on clean samples. This study designs and verifies a robust modulation recognition framework that operates effectively under blind adversarial attack scenarios without prior knowledge of attack type or strategy. The goal is to support reliable deployment of intelligent communication systems in adversarial environments.  Methods  The proposed framework integrates a feature-purifying autoencoder module with standard modulation classifiers, including Convolutional Neural Network (CNN) and Transformer architectures. The core component is the autoencoder bottleneck layer, which implements a dynamic purification mechanism. First, an adaptive threshold is calculated from the statistical properties of encoded latent features to detect anomalies. Then a Top-K sparsification operation retains the most significant feature activations. This step suppresses noise and adversarial perturbations and preserves essential signal characteristics. The autoencoder is trained using a three-stage curriculum learning strategy. The stages sequentially optimize reconstruction fidelity, feature sparsity, and semantic consistency between purified signals and original clean signals. This process guides the reconstructed signals toward the true modulation manifold. The module is model-agnostic and can be placed before a trained classifier without retraining.  Results and Discussions  Experiments are conducted on a simulated dataset containing 12 digital modulation types under multipath fading channels. The framework produces clear performance gains. Under targeted white-box attacks, recognition accuracy increases to 82.1% for CNN and 83.2% for Transformer. Under non-targeted black-box attacks, accuracy reaches 87.7% and 89.4%, respectively (Table 1). The Attack Success Rate (ASR) and Attack Effectiveness Index (AEI) remain low, indicating strong defense capability. Figure 4 shows that defense performance improves as the Signal-to-Noise Ratio (SNR) increases. The ablation study in Figure 5 confirms the critical role of the autoencoder. Removing this module reduces accuracy by 4.02% for CNN and 2.36% for Transformer under strong attacks. Further analysis in Figure 6 shows that the framework maintains stable robustness across a wide perturbation range (\begin{document}$ \epsilon \leq 0.1 $\end{document}). Parameter sensitivity experiments in Figures 7 and 8 indicate stable performance when the threshold coefficient \begin{document}$ \xi $\end{document} is within [1.5, 1.9] and the sparsity rate k is around 0.7. These results support practical deployment.  Conclusions  A robust blind defense framework for AMR is presented based on a feature-purifying autoencoder. The framework provides three main advantages. First, it defends against different white-box and black-box attacks without requiring prior knowledge of attack methods. Second, as a preprocessing module, it avoids computationally expensive retraining of the primary classifier and remains compatible with different backbone networks. Third, the multi-stage training strategy balances adversarial robustness with high accuracy on clean samples. Experiments on the simulated dataset confirm the effectiveness of the proposed framework. Future work will explore lightweight architectural designs to reduce inference latency and will further investigate performance limits under extremely low SNR conditions combined with nonlinear channel impairments.
Co-Frequency Interference Analysis and Dynamic Simulation Validation of Satellite-Direct-to-Device Systems Against Terrestrial IMT Networks in Cross-Border Scenarios
LIU Quan, ZHAO Weisong, XIAO Na, SONG Yanjun, ZHOU Meng, ZHANG Zhili, WANG Jinhai, WANG Lichong
Available online  , doi: 10.11999/JEIT260263
Abstract:
  Objective   Satellite-Direct-to-Device (SD2D) systems that reuse terrestrial IMT spectrum may generate harmful downlink interference to incumbent IMT networks in neighboring administrations, particularly in cross-border deployments where SD2D downlinks overlap the receive bands of both IMT user equipment (UE) and IMT Base Stations (BSs). A practical coexistence methodology is therefore required to (i) translate IMT receiver protection criteria into explicit Power Flux Density (PFD) and Equivalent Power Flux Density (EPFD) constraints and (ii) validate these constraints using a dynamic simulation framework so that they can be converted into enforceable geographic coordination measures, such as minimum isolation distances. This study focuses on the dominant interference path, namely SD2D downlink interference to IMT receivers, and establishes a traceable workflow from deterministic protection limits to dynamic simulation validation and the corresponding minimum isolation distances.  Methods  A cross-border scenario is modeled in which Country A deploys an SD2D system and Country B operates a terrestrial IMT network. Two representative downlink frequencies, 1 995 MHz and 2 190 MHz, are evaluated for two representative Starlink configurations, Starlink-1 and Starlink-2. The IMT network is modeled using ITU-R- and 3GPP-compliant parameters, with an I/N protection threshold of –6 dB and a target percentile κ (baseline κ=99.5%) for both IMT UEs and BSs. Satellite transmit antennas follow the ITU-R S.1528 reference pattern, IMT BS receive antennas follow the ITU-R F.1336 sector pattern, and IMT UEs are modeled with omnidirectional antennas. A back-lobe blockage model is incorporated into both satellite and BS antenna patterns to account for rear-side shielding. Signal propagation follows the ITU-R P.619 model, using free-space path loss as the conservative baseline, while an optional clutter-loss term is incorporated through a clutter-occurrence probability. Deterministic protection limits are derived by calculating the maximum permissible aggregate PFD for IMT UE protection and the maximum permissible aggregate EPFD for IMT BS protection. A dynamic simulation framework then validates these limits and searches for the required minimum isolation distances (Fig. 4). Co-channel beam isolation angles are optimized using the C/I Complementary Cumulative Distribution Function (CCDF), and a segmented search algorithm determines the minimum UE- and BS-side isolation distances together with the corresponding κ-percentile PFD/EPFD statistics.  Results and Discussions  The deterministic analysis yields a maximum permissible aggregate PFD of –102.72 dBW/m2/MHz at 2 190 MHz for IMT UEs and a maximum permissible aggregate EPFD of –129.53 dBW/m2/MHz at 1 995 MHz for IMT BSs (Fig. 3). For Starlink-1, the C/I design criterion yields a minimum co-channel beam isolation angle pair of (12°, 12°) (Fig. 5). Dynamic simulation shows that, under the representative baseline configuration with an I/N threshold of –6 dB and κ=99.5%, the minimum isolation distances are 195 km for UE protection and 290 km for BS protection (Fig. 6, Fig. 7, and Table 4). The resulting coordination isolation distance is therefore 290 km, and the simulated κ-percentile PFD and EPFD agree with the deterministic protection limits, with a residual margin below 0.5 dB. For Starlink-2, the optimized co-channel beam isolation angles increase to (15°, 15°), and the corresponding minimum isolation distances increase to 272 km for UEs and 420 km for BSs under the same baseline configuration (Table 5). These baseline distances should be interpreted as representative values for the specified simulation configuration rather than unique, strictly converged results. Stability verification shows that, under different sampling intervals, simulation durations, and random seeds, the UE- and BS-side minimum isolation distances remain within 195~210 km and 290~300 km, respectively, for Starlink-1, and within 266~290 km and 370~420 km, respectively, for Starlink-2 (Table 7). Sensitivity analysis for Starlink-1 further indicates that the required minimum isolation distance is governed by the upper tail of the aggregate I/N distribution (Table 6). Increasing κ from 99.5% to 100% increases the UE- and BS-side minimum isolation distances from 195/290 km to 304/560 km. Clutter attenuation substantially reduces the UE-side minimum isolation distance, decreasing it to 173 km when the clutter-occurrence probability is 0.5, while producing little change in BS protection. Polarization reuse increases the UE- and BS-side minimum isolation distances to 222 km and 360 km, respectively, whereas increasing the number of co-channel beams to 16 increases the BS-side minimum isolation distance to 330 km. The minimum service elevation angle and the link establishment strategy are identified as the dominant operational factors. Changing the minimum service elevation angle from 10° to 35° changes the required UE- and BS-side minimum isolation distances from 340/460 km to 101/150 km, whereas replacing the Sat-MaxElevation strategy with the UE-MaxElevation strategy reduces them to 80/180 km.  Conclusions   The proposed workflow converts IMT receiver protection criteria into deterministic protection limits expressed as PFD and EPFD constraints and validates them using a dynamic simulation framework. Under an I/N threshold of –6 dB and κ=99.5%, the baseline and stability analyses jointly indicate representative UE- and BS-side minimum isolation-distance ranges of 195~210 km and 290~300 km for Starlink-1 and 266~290 km and 370~420 km for Starlink-2, rather than unique, strictly converged values. Sensitivity analysis further shows that κ only changes the statistical criterion used to extract tail events from the sample set, whereas clutter attenuation primarily benefits IMT UEs. In contrast, the minimum service elevation angle, polarization reuse, the number of co-channel beams, and the link establishment strategy reshape the worst-case interference geometry and can produce substantial, and sometimes non-monotonic, changes in the required minimum isolation distances. The proposed framework establishes a traceable link between IMT receiver protection criteria and enforceable border coordination measures.
A Novel TDMOSFET and Its Neural Network Modeling for Ternary Logic Applications
LU Bin, LU Haoran, ZHAO Xiaohong, DI Jiayu, XING Linlin
Available online  , doi: 10.11999/JEIT260413
Abstract:
  Objective  Complementary metal-oxide-semiconductor (CMOS) technology is continuously improving, moving toward smaller size and higher integration. As circuit integration increases, short-channel effects and other phenomena lead to a significant rise in leakage current in MOSFET devices, resulting in higher static power consumption. Against the backdrop of rapid advancements in artificial intelligence, traditional binary logic chips face severe limitations in computing and storing massive amounts of data. To meet the demands for higher efficiency, greater density, and lower power consumption, ternary logic technology has attracted widespread attention from researchers. Compared with traditional binary logic, ternary logic offers advantages such as higher information density and lower system complexity. However, the current design of ternary logic circuits faces several challenges, including the need for a large number of components, the involvement of passive elements, and poor compatibility with conventional CMOS processes.  Methods  To address these issues, this paper proposes a novel Tunneling and Drift-Diffusion Metal-Oxide-Semiconductor Field-Effect Transistor (TDMOSFET) that integrates both quantum tunneling and drift-diffusion mechanisms. This device features a constant off-state circuit characteristic, making it suitable for ternary logic applications. The working principle of the TDMOSFET is analyzed in detail, and an artificial neural network (ANN) is employed to model the device. The established ANN model can accurately simulate the current-voltage (IV) and capacitance-voltage (CV) characteristics of the device. Furthermore, the ANN model is converted into a Verilog-A program and embedded into HSPICE to simulate basic ternary logic circuits, including the Standard Ternary Inverter (STI), Negative Ternary Inverter (NTI), Positive Ternary Inverter (PTI), Ternary NOT-AND gate (T-NAND), and Ternary NOT-OR gate (T-NOR).  Results and Discussions  The well-trained ANN model can accurately predicted the current-voltage and capacitance-voltage performance. The maximum relative errors of the ANN model compared with the TCAD results are 39.43%, 5.05%, and 14.19% for the drain current IDS, gate-drain capacitance CGD and gate-source capacitance CGS, respectively, and the average relative errors are 0.46%, 0.69%, and 0.51%, respectively. Moreover, the well-trained ANN is successfully converted into Verilog-A programs and demonstrates excellent compatibility with widely used HSPICE tools. Based on the models, the basical units including the STI, NTI, PTI, T-NAND and T-NOR are simulated. The designed ternary logic circuits eliminate the need for passive elements and are compatible with conventional CMOS processes.  Conclusions  This paper proposes a novel TDMOSFET. It maintains compatibility with CMOS fabrication processes, thereby simplifying the manufacturing workflow and reducing both cost and integration complexity. The proposed ternary logic circuits based on the TDMOSFETs achieve ternary operation without increasing the number of transistors, relying on passive components, or requiring multi-valued supply voltages., offering significant reference value for future research. Furthermore, the “TCAD→ANN Modeling→Verilog-A Language→HSPICE Simulation” methodology established in this work can also be applied to the investigation of other emerging semiconductor devices.
MG-MoE: Routed Multi-Granularity Expert Ensemble
XIAN Fengyu, JIAN Haifang, XIE Zihui, DU Jun, ZHANG Yuanyuan, NING Xin, DONG Miaomiao, WANG Hongchang
Available online  , doi: 10.11999/JEIT260219
Abstract:
  Objective  Fine-Grained Image Recognition (FGIR) aims to distinguish visually similar subcategories that differ only in subtle local patterns. It must also remain robust to large intra-class variations caused by pose changes, occlusion, illumination shifts, and complex backgrounds. In real-world scenarios, these challenges are further intensified by long-tailed category distributions. Rare or difficult classes are more likely to overfit spurious contextual cues and suffer from unstable decision boundaries. Therefore, a conditional computation paradigm is needed, in which complementary inductive biases are separated into specialized expert branches and adaptively combined for each sample. This work aims to develop a routed multi-granularity mixture-of-experts framework that improves discriminative performance under controllable inference cost. It also enhances robustness for difficult samples and long-tailed categories through adaptive sparse expert activation.  Methods  A Multi-Granularity Mixture-of-Experts (MG-MoE) model is proposed. It is a routed ensemble architecture composed of a shared backbone, four heterogeneous experts, and a learnable router that predicts input-conditioned expert weights (Fig. 2). The experts are designed with complementary inductive biases to address key factors in FGIR. MPSA emphasizes global structure and contour-level semantics. PMG captures fine local details through multi-granularity part modeling. TransFG focuses on pose and deformation modeling. PIM improves robustness in cluttered backgrounds through background suppression. To limit interference and reduce unnecessary computation, MG-MoE adopts sparse fusion. Only the Top-K experts, with K=2 by default, contribute to the final prediction during inference. To improve routing stability and generalization, a two-stage optimization strategy is designed. In the first stage, dynamic cluster-level training is performed. A cluster-level soft teacher distribution is constructed from validation-set statistics and imposed through Kullback-Leibler (KL) divergence regularization. This process stabilizes routing behavior and promotes effective expert specialization. In the second stage, residual fine-tuning is conducted. The feature-driven routing mechanism is kept unchanged, while the classification heads of the Top-2 experts associated with each cluster are selectively unfrozen. The router and expert heads are then jointly optimized with grouped learning rates. This design reduces fusion bias and strengthens discrimination for difficult samples and long-tailed categories.  Results and Discussions  MG-MoE achieves strong performance on standard FGIR benchmarks. On CUB-200-2011, it obtains 92.89% Top-1 accuracy. This result is higher than those of representative expert backbones used individually, including MPSA (91.23%), PIM (91.17%), and TransFG (90.49%). It also outperforms the multi-granularity baseline PMG (88.32%) (Table 1). On the Bird-1445 sampled set, MG-MoE achieves 96.80% Top-1 accuracy and consistently improves over strong baselines (Table 2). These results indicate that routed multi-expert specialization remains effective in data-limited and highly similar fine-grained scenarios. The efficiency-accuracy trade-off is summarized in Table 3. With Top-2 sparse routing, MG-MoE reaches 92.89% accuracy with a compute budget of 143.9 GFLOPs. It avoids dense expert activation during inference by selecting only the Top-2 experts for each sample, thereby achieving a favorable balance between accuracy and efficiency. Ablation experiments show that increasing K beyond 2 does not yield consistent gains, which suggests that indiscriminate fusion can dilute discriminative evidence. Top-2 fusion produces the best performance, whereas Top-1 fusion is more sensitive to routing errors and larger K values may introduce noise and reduce accuracy (Table 4). The role of expert diversity and composition is also analyzed. Two- and three-expert variants generally underperform the full four-expert configuration, indicating that each inductive bias contributes to different fine-grained difficulty factors. In contrast, adding homogeneous experts without new functional diversity brings diminishing or negative gains, which is consistent with increased routing ambiguity and limited expert complementarity (Table 5). These results support the use of a compact set of heterogeneous experts combined with sparse routing. To interpret the learned specialization, category-wise routing statistics are visualized. The expert-category heatmap shows that MPSA receives dominant routing weights across many categories, reflecting the central role of global structure in fine-grained discrimination. PIM and TransFG show higher activation for specific difficult categories, which is consistent with their roles in background suppression and pose and deformation modeling (Fig. 3). Finally, t-SNE visualizations illustrate the qualitative effect of expert fusion on class separability. Shared backbone features show stronger inter-class entanglement among visually similar subcategories. In contrast, fused outputs form clearer clusters with better between-class separation and within-class compactness, indicating a more reliable decision space shaped by routed expert aggregation (Fig. 4).  Conclusions  MG-MoE is a multi-granularity routed mixture-of-experts framework for fine-grained recognition. By combining four complementary experts, Top-2 sparse fusion, and a two-stage optimization strategy for stable routing and calibrated fusion, MG-MoE improves recognition accuracy on CUB-200-2011 and the Bird-1445 sampled set. It also provides interpretable evidence of expert specialization (Table 1, Table 2, Fig. 3, Fig. 4). Ablation results confirm that controlled Top-2 fusion and heterogeneous expert design are key to the observed performance gains. Overly dense fusion or homogeneous expert expansion provides limited benefit (Table 4, Table 5).
Energy-Aware and Attention-Driven Edge–End Collaborative Inference and Resource Allocation
LIU Yiming, TIAN Jie, LI Tiantian, ZHOU Xiaotian, ZHANG Haixia
Available online  , doi: 10.11999/JEIT260086
Abstract:
  Objective   The perception and decision-making capabilities of mobile intelligent applications have been significantly improved by the development of Deep Neural Networks (DNN). Nevertheless, these applications often have high computational loads, which greatly strain mobile terminals with constrained latency and energy consumption. Due to limited wireless bandwidth and edge computing power, severe resource contention is inevitable in multi-user concurrent scenarios, even though mobile edge computing (MEC) lowers latency pressure by moving tasks from the edge to the terminal. More importantly, because edge servers have limited resources, strict budget control of their long-term energy consumption is necessary to ensure the system's dependability and long-term operational efficiency.An energy-aware and attention-driven edge-to-edge collaborative inference and resource allocation technique is proposed in this paper. The majority of research ignores the system's long-term energy consumption limitations in favor of maximizing instantaneous performance. This approach greatly lowers the average end-to-end latency of multi-user inference tasks under various load scenarios and greatly increases the utilization efficiency of general computing resources while closely adhering to the long-term energy budget.  Methods   Based on the DNN collaborative inference model in an edge environment, this paper models the problem as minimizing the long-term average end-to-end processing latency of all user inference tasks, under the premise of satisfying the long-term energy consumption budget constraint of the ES system. This optimization problem and energy consumption constraint both involve long-term averages and stochasticity, and the DNN model partitioning and general computing resource allocation decisions are highly coupled. Therefore, this paper adopts Lyapunov optimization theory to transform this complex problem into a more manageable single-slot deterministic optimization problem. In order to jointly finish the DNN partitioning and general computing resource allocation decisions within each time slot, this paper develops a Joint Collaborative Inference and Resource Allocation Algorithm (JCIRA). There are three steps in the algorithm (Algorithm 1). To determine the ideal DNN model partitioning point for the task, the first step balances the estimated latency against energy penalties based on the comprehensive cost function. Task urgency, remaining computational load, data volume, and global energy deficit state are all mapped into high-dimensional feature vectors in the second stage, which presents a joint attention mechanism based on the key-query-value paradigm. To accomplish cooperative distribution of general computing resources, the matching degree is computed using a scaled dot product attention mechanism. In order to improve the system's utilization rate of general computing resources and lower the average end-to-end latency of multi-user inference tasks while closely adhering to long-term energy consumption constraints, the third stage creates a closed-loop decision-making process by monitoring execution progress and updating queue status.  Results and Discussions   In this simulation experiment, the task arrival process is random and has a Poisson distribution. In this experiment, three distinct deep neural network models were employed: ResNet18, MobileNetV2, and EfficientNet-B0. When compared to other algorithms, JCIRA's actual energy consumption under various load scenarios consistently falls below the budget threshold, demonstrating the proposed algorithm's ability to manage long-term energy consumption constraints (Fig. 2). Furthermore, JCIRA's virtual queue length is always brief (Fig. 3). Among all schemes that meet the energy consumption constraints, JCIRA's latency is consistently below the 300 ms QoS threshold (Fig. 4). It performs better and successfully resolves the problem of other algorithms' insufficient processing power under heavy loads. Finally, the algorithm maintains the highest task completion rate even with 1300 concurrent tasks in a high load scenario (Fig. 5).  Conclusions   To address the challenges of limited computing resources and long-term energy consumption constraints in DNN collaborative inference within a MEC environment, this paper proposes an energy-aware and attention-driven edge-end collaborative inference and resource allocation method. The long-term energy consumption hard constraint is separated into a low-complexity single-slot deterministic optimization subproblem using Lyapunov optimization theory. In order to accomplish systematic scheduling of DNN model partitioning and computing resources, a joint optimization algorithm called JCIRA is created. Simulation results show that the suggested approach significantly lowers the average end-to-end latency of multi-user inference tasks under different load conditions and efficiently optimizes the use of computing resources while guaranteeing strict satisfaction of long-term energy consumption constraints.
Efficient Non-Orthogonal Multiple Access Scheme Based on Modified Alamouti Code Design
WAN Dehuan, HUANG Ronglan, JI Fei, LIU Jingxian, LIANG Yaokun, LÜ Lu, LI Xingwang, YUE Xinwei
Available online  , doi: 10.11999/JEIT260567
Abstract:
  Objective  Existing Alamouti-coding-based Non-Orthogonal Multiple Access (NOMA) schemes adopt an equal-number symbol transmission mode for both cell-center and cell-edge users, which overlooks the significant channel disparity between the two types of users. This leads to two drawbacks: on one hand, the superior channel condition of cell-center users is not fully exploited, resulting in wasted transmission resources; on the other hand, the equal-number transmission inevitably weakens the performance of Space-Time Block Coding (STBC) in suppressing intra-group multi-user interference. Therefore, a novel design that adapts to user channel differences is urgently needed to improve spectral efficiency and interference mitigation capability.  Methods  In light of the significant channel disparity between cell-center and cell-edge users, this paper proposes an unequal-number symbol transmission scheme for Alamouti coding in NOMA. Specifically, when constructing Alamouti group transmission codes, the number of symbols required for the cell-edge user is made smaller than that for the cell-center user. By reducing the number of transmitted symbols for the cell-edge user, two benefits are achieved: first, the multi-user interference imposed on the cell-center user is directly reduced; second, under a total power constraint, reducing the number of symbols for the cell-edge user equivalently increases the per-symbol transmission power, thereby significantly improving the received signal-to-interference-plus-noise ratio (SINR). Furthermore, the paper derives perfect closed-form solutions for the achievable sum rate and outage probability of the proposed scheme, and validates its effectiveness via numerical simulations.  Results and Discussions  The theoretical derivations yield closed-form analytical expressions for the achievable sum rate and outage probability, providing an accurate basis for system performance evaluation. Numerical simulation results demonstrate that, compared with existing equal-symbol Alamouti coding schemes, the proposed unequal-symbol Alamouti coding scheme effectively reduces the intra-group interference from the cell-edge user to the cell-center user, while significantly enhancing the SINR of the cell-edge user through power reallocation. Under typical channel parameters, the system achieves a notable sum-rate gain and a substantial reduction in outage probability, confirming the high efficiency of the proposed scheme.  Conclusions  The proposed unequal-number symbol transmission scheme based on Alamouti coding for cell-edge and cell-center users fully exploits the potential benefits arising from channel disparity. By reducing the number of symbols transmitted by the cell-edge user, the scheme achieves both suppression of intra-group interference and enhancement of the edge user’s transmission power. The theoretical closed-form solutions and simulation results consistently show that the proposed scheme outperforms conventional equal-number transmission schemes, providing an effective new approach for mitigating multi-user interference and optimizing resource allocation in NOMA systems.
Function-Aware Partitioning Driven Hierarchical Circuit Representation Learning
YE Juyang, CHEN Qilin, WANG Yaohua
Available online  , doi: 10.11999/JEIT260645
Abstract:
  Objective  One of the core challenges in applying machine learning techniques to electronic design automation (EDA) lies in learning high-quality circuit representations from large-scale gate-level netlists. As modern digital integrated circuits scale to tens of millions of gates, existing methods based on graph neural networks (GNNs) and graph transformers (GTs) suffer from excessive computational and memory overhead, rendering them impractical for industrial-scale designs. The fundamental issue is the lack of a proper tokenization mechanism for netlists—unlike natural language, where subword tokenization effectively compresses long sequences, the circuit domain lacks an analogous decomposition strategy that preserves functional semantics while reducing the effective graph size. This work aims to bridge this gap by introducing a hypergraph-partitioning-based circuit tokenizer that decomposes massive netlists into functionally cohesive sub-circuits, termed circuit elements, thereby enabling scalable and fine-grained representation learning.  Methods  This paper proposes a function-aware hypergraph partitioning driven framework for large-scale circuit representation learning. Inspired by the tokenization paradigm of large language models, the framework first models a gate-level netlist as a directed hypergraph, where gates are nodes and signals are hyperedges that can connect multiple gates. A novel optimization objective, the Functional Independence Ratio (FIR), is introduced to guide the partitioning process. FIR incorporates circuit structural priors and a bus recognition correction mechanism that identifies bus-structured signals based on structural similarity (gate type purity across predecessor/successor levels) and spatial similarity (topological distance variance within candidate groups). The bus recognition module corrects the effective interface count, ensuring that functionally cohesive modules are not penalized for using wide buses. An iterative greedy refinement procedure accepts only moves that strictly decrease FIR, converging to a locally optimal partition. On top of the partitioned circuit elements, a two-stage self-supervised pretraining framework is designed. In the first stage, a masked autoencoder with edge prediction tasks is applied to the coarse-grained circuit-element graph, learning global inter-element dependencies. The graph transformer encoder is then frozen and circuit element embeddings are saved. In the second stage, within each circuit element, a contrastive learning scheme is employed at the gate level. Positive pairs are constructed via Boolean equivalence transformations (e.g., associativity, De Morgan's laws), which preserve the Boolean function while altering the gate-level structure. Negative pairs are drawn from functionally different circuit elements within the same batch. The training jointly optimizes node-level and local-global alignment losses. A feature-wise modulation mechanism injects circuit-element-level context into gate-level representations, enabling the same gate type to acquire different embeddings depending on its functional context.  Results and Discussions  Extensive experiments are conducted on circuits collected from multiple sources, including ITC99, EPFL, OpenCores, and three RISC-V SoC designs (Rocket, BOOM, and OpenC910), with the largest design containing 22.3 million gates. For the tokenizer evaluation, FIR-based partitioning is compared against Mt-kahypar using the Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) metrics, with the original module hierarchy serving as ground truth. Across all three RISC-V designs, FIR achieves AMI improvements of 24.6% to 27.1% and ARI improvements of 25.5% to 35.8% over Mt-kahypar, demonstrating consistent and substantial gains in functional coherence. On downstream tasks, the proposed method consistently outperforms state-of-the-art baselines including DeepGate4 and NetTAG across all comparable datasets. For full-circuit function recognition, the proposed method achieves F1 scores of 0.896, 0.861, 0.817, and 0.773 on ITC99, EPFL, OpenCores, and Rocket, respectively, representing improvements of 10.0% to 15.4% over the strongest baseline. Crucially, on BOOM (8.12M gates) and OpenC910 (22.3M gates), the proposed method is the only approach capable of completing training without running out of memory, achieving F1 scores of 0.771 and 0.748 for full-circuit tasks and 0.723 and 0.679 for gate-level tasks, respectively. Ablation studies with six model variants reveal a clear division of labor: the module-level pretraining stage dominates full-circuit performance (20.9% F1 drop when removed), while the gate-level pretraining stage dominates gate-level performance (26.1% F1 drop when removed). The FIR partitioning objective contributes 8.6% and 11.7% F1 improvements at the full-circuit and gate levels, respectively. The bus recognition module and the two-stage architecture also show consistent positive contributions across all metrics.  Conclusions  This paper presents a novel framework that addresses the scalability challenge in circuit representation learning by introducing a function-aware hypergraph partitioning tokenizer and a two-stage self-supervised pretraining architecture. The key insight is that by abstracting the intermediate circuit element level between individual gates and the full circuit, one can achieve both scalability to multi-million-gate designs and fine-grained gate-level discriminative capability. The proposed FIR objective effectively captures functional cohesion during partitioning, and the two-stage pretraining framework decouples global context learning from local representation refinement. Experimental results demonstrate state-of-the-art performance on function recognition tasks and, more importantly, the unique ability to scale to industrial-sized designs where existing methods fail. Future work includes extending the framework to other EDA tasks such as logic synthesis and physical design, exploring more aggressive hierarchical strategies for billion-gate designs, and adapting the bus recognition mechanism to non-standard cell libraries.
Difference-aware Adaptive Prompt Learning and Dense Alignment for Weakly Supervised Building Change Detection
CHEN Yanxia, MA Longlong, CHEN Yanhua, HUANG Yuchun
Available online  , doi: 10.11999/JEIT260595
Abstract:
  Objective  Building change detection from bi-temporal high-resolution remote sensing images is an important task for urban planning, land resource management, illegal construction monitoring, and disaster damage assessment. Existing fully supervised change detection methods usually achieve high detection accuracy by relying on pixel-level annotations. However, obtaining pixel-level labels for large-scale remote sensing images is labor-intensive and time-consuming, which limits their practical application in large-scale monitoring scenarios. Image-level weakly supervised change detection reduces annotation costs by using only image-level labels that indicate whether an image pair contains changes. Nevertheless, the lack of spatial supervision makes it difficult to accurately locate changed regions. Current weakly supervised methods generally rely on class activation maps (CAMs) to generate pseudo labels, but CAMs tend to highlight only the most discriminative regions and may ignore complete change areas or introduce background noise. Vision-language models provide a new way to introduce semantic priors into weakly supervised learning. However, directly applying Contrastive Language-Image Pre-training (CLIP) to change detection is still challenging. Fixed text prompts are difficult to adapt to the difference semantics of bi-temporal images, and the original CLIP objective mainly focuses on global image-text alignment rather than pixel-level localization. To address these problems, this paper proposes a Difference-aware Adaptive Prompt Learning and Dense Alignment method for weakly supervised building Change Detection, named DAPL-CD.  Methods  The framework introduces CLIP-based cross-modal semantic knowledge into image-level weakly supervised building change detection. For a pair of bi-temporal remote sensing images, a shared CLIP visual encoder is first used to extract visual representations from the two temporal images. The local visual features are fused along the channel dimension to obtain bi-temporal difference features, which contain semantic information related to changed and unchanged regions. Based on the difference characteristics of building change detection, a difference-aware adaptive prompt learning strategy is designed. Instead of manually designed fixed text templates, learnable context vectors are inserted into text prompts while preserving category-related semantic words. The generated foreground and background text embeddings are used as semantic prototypes to provide adaptive semantic guidance for change localization. Furthermore, a pixel-text dense alignment mechanism is introduced to extend CLIP’s global image-text alignment capability to local feature matching. The initial CAM generated by the classification branch is used to obtain preliminary foreground and background regions. Then, visual-text positive and negative sample pairs are constructed between local difference features and text embeddings. An InfoNCE-based dense alignment loss is employed to pull matched visual and textual features closer and push mismatched features apart. Finally, the classification branch and segmentation branch are jointly optimized using classification loss, global alignment loss, dense alignment loss, and segmentation loss. To avoid unreliable pseudo labels during early training, the segmentation loss is introduced after the CAM quality becomes relatively stable.  Results and Discussions  Experiments are conducted on WHU-CD and LEVIR-CD, two public benchmark datasets for building change detection. During training, only image-level labels are used, while pixel-level annotations are used only for evaluation. Overall Accuracy (OA), F1-score, and Intersection over Union (IoU) are adopted as the main evaluation metrics. Because changed buildings usually account for a small proportion of remote sensing images, OA can be strongly affected by the dominant unchanged background pixels. Therefore, F1-score and IoU are emphasized to evaluate the detection quality of changed regions. Quantitative comparison results show that DAPL-CD achieves 94.7% OA, 82.8% F1-score, and 70.6% IoU on the WHU-CD dataset, and obtains 92.3% OA, 68.0% F1-score, and 51.5% IoU on the LEVIR-CD dataset, outperforming the compared weakly supervised change detection methods in terms of F1-score and IoU (Table 1). Visual comparison results further demonstrate that the proposed method produces more complete responses for large-scale building changes and more continuous predictions for small and scattered changed buildings (Fig.3, Fig.4). Ablation experiments verify the effectiveness of the proposed adaptive prompt learning and pixel-text dense alignment mechanisms. The baseline model, which uses fixed text prompts without foreground or background alignment, achieves an F1-score of 63.2% and an IoU of 46.2%. Introducing both foreground and background alignment improves the two metrics to 66.3% and 49.6%, respectively, demonstrating that dense semantic matching between local visual features and text prototypes enhances the discrimination of changed regions. After adaptive prompt learning is incorporated, the F1-score and IoU increase to 68.0% and 51.5%, indicating that learnable prompts alleviate the semantic mismatch between fixed text descriptions and bi-temporal difference features. Under the learnable-prompt setting, foreground alignment alone achieves an F1-score of 64.4% and an IoU of 47.5%, whereas background alignment alone obtains 65.3% and 48.5%, respectively. Combining the two alignment branches achieves the best F1-score and IoU, demonstrating that foreground and background semantic constraints provide complementary guidance for change localization (Table 2). The CAM results show that dense alignment produces stronger and more complete responses over actual changed regions while suppressing irrelevant background activations (Fig. 5).  Conclusions  A weakly supervised building change detection framework based on difference-aware adaptive prompt learning and pixel-text dense semantic alignment is proposed. By introducing CLIP-based vision-language semantic priors, the proposed method effectively transforms text-level semantic knowledge into local change localization capability. The adaptive prompt learning strategy improves the representation of change-related semantic descriptions, while the dense alignment mechanism establishes direct correspondence between local difference features and foreground/background semantic prototypes. Experimental results on WHU-CD and LEVIR-CD demonstrate that DAPL-CD achieves competitive performance under image-level supervision and improves the completeness and accuracy of changed building localization. The proposed framework provides an effective solution for reducing annotation requirements in large-scale remote sensing change detection. Future research will focus on improving pseudo-label reliability, reducing dependence on large-scale pre-trained models, and extending the method to multi-temporal or multi-spectral remote sensing data.
Research on LEO Constellation Interference Prediction and Detection Algorithm Driven by Collaborative Spatial Feature Mapping and Temporal Transformer
YANG Boyu, QIU Kun, CHEN Zhe, ZHAO Jin, GAO Yue
Available online  , doi: 10.11999/JEIT260368
Abstract:
  Objective  The continuous putting into use of large-scale low Earth orbit (LEO) satellite constellation groups has caused a serious crowding of orbital space resources. The wide employment of frequency repeating technique by satellite operation persons has seriously brought challenges to the frequency spectrum resources of current LEO satellite network systems. At the same time, the inner quick motion and changing topological structure of LEO satellites bring about violent changes in satellite electric power lead to serious co-frequency interference between satellites. Existing detection methods for interference between satellites mainly depend on compulsory rules formulated by the International Telecommunication Union (ITU), which employ the Interference-to-Noise Ratio (I/N) to serve as the primary indicator. Nevertheless, the link attenuation formulae and satellite position computations that are contained in the calculation procedure bring about a sharp rise in total calculation cost, hence it is hard to satisfy the real-time interference check demands of large-scale satellite constellation working. The recent interference detection algorithms which are based on deep learning are still limited to taking all visible satellites as input, thus they cannot carry out constraint on the spatial sparsity of interference feature distribution. For solving these problems, this paper puts forward a LEO constellation interference prediction and detection method which is based on spatial feature mapping and temporal Transformer, therefore it greatly decreases calculation complexity while it realizes high-efficiency interference prediction and detection.  Methods  The proposed spatial-temporal Transformer framework efficiently models dynamic co-frequency interference in large-scale heterogeneous multi-constellation LEO networks. By analyzing ITU interference formulations, we extract three core physical properties: permutation invariance, spatial sparsity, and temporal continuity. Consequently, the traditional physical iteration is reframed as a spatio-temporally decoupled feature mapping problem. A dedicated spatial mapping module, leveraging attention and parallel pooling mechanisms, compresses raw satellite parameters into robust permutation-invariant features. This effectively highlights key interference sources while filtering redundant background nodes. Finally, a temporal Transformer captures the dynamics of interference trajectories, achieving high-accuracy prediction and significantly reducing computational complexity for real-time monitoring applications.  Results and Discussions  The interference prediction and detection algorithm for LEO constellation uses space feature mapping and a time Transformer to carry out interference detection and forecast by making use of high-fidelity data. According to ITU regulations, one high-mobility LEO satellite scene has been simulated, which includes thousands of real satellites from Starlink and OneWeb constellations. Parameter analysis results show that setting reasonable feature dimensions and encoder layers can achieve an optimal balance between feature extraction accuracy and computational cost (Fig. 4 and Fig. 5). By setting model parameters such as sampling step size, prediction accuracy and engineering real-time performance can be comprehensively considered (Fig. 6 and Fig. 7). Simulation results show that compared to baseline methods, the predicted trajectory of the proposed method closely matches the baseline ground truth (Fig. 8), with the absolute deviation controlled within 0.5 dB at a 90% cumulative probability (Fig. 9). Furthermore, the algorithm maintains an extremely high interference detection rate with low false alarms (Fig. 10), and the Root Mean Square Error (RMSE) remains stable at 0.45 dB within a 20.0 s prediction window, significantly outperforming traditional regression and neural network models (Fig. 11). In terms of computational efficiency, the algorithm in this paper takes 14.2 ms for a single inference, which is only 4.1% of the time required by the traditional ITU-based physical iteration method (Table 3), effectively decoupling computational overhead and constellation scale.  Conclusions  For the problem that LEO satellite network topology has extremely strong dynamic property, and traditional physical iteration interference checking methods have limitations brought by sharply risen computation complexity, this paper puts forward a low-Earth orbit constellation interference checking algorithm which combines spatial feature mapping and temporal Transformer, therefore it effectively makes computation cost and constellation size decouple. The results of simulation show that, when the comparison is made with baseline methods, this algorithm possesses fine long-term dynamic tracking capabilities, and at the same time it keeps conformity with ITU evaluation rules. Experiments which carry on equipment analysis with different parameters have proven that the algorithm we put forward can obtain a balance between feature extraction accuracy and computation spending. By utilizing the long-distance dependence modeling ability of time-domain Transformer, the RMSE still keeps steady at 0.45 dB at identical time step. Through the method of eliminating unnecessary background satellite nodes, the single inference time has been reduced to 14.2 ms. By means of manifold experiments, which contain error distribution assessment and model parameter analysis, the algorithm put forward by us obtains the best precision and full decoupling of calculation cost from constellation scale, hence it gives a highly flexible engineering scheme for future super-large constellation arrangements.
Random-Linear-Network-Coding-Based Cooperative Reliable Transmission Protocol for Underwater Acoustic Communication Networks
ZHANG Zhilin, PU Zhanqing, ZHU Yunan, LI Xueying, TIAN Jie, HUANG Haining
Available online  , doi: 10.11999/JEIT260648
Abstract:
  Objective  Reliable data delivery in underwater acoustic communication networks is constrained by high packet error rates, long propagation delays, limited bandwidth, and topology variations. In single-source dual-sink multi-hop transmission, the same data generation must be reliably delivered to two destination nodes, and single-hop packet losses can accumulate across multi-hop forwarding and joint recovery at the two destinations, making reliable delivery more difficult. Existing reliability-enhancement mechanisms, such as retransmission, redundant forwarding, forward error correction, and multipath redundant transmission, usually rely on predetermined forwarding structures or fixed redundancy configurations. These mechanisms have limited ability to exploit the complementary coded information distributed among multiple relay nodes, and therefore still suffer from insufficient joint recovery capability and high redundancy overhead. To address this problem, this paper proposes a network-coded cooperative reliable transmission protocol, namely the Network-Coded Cooperative Reliable Transmission Protocol for Underwater Acoustic Communication Networks (NCCRTP), for underwater acoustic communication networks.  Methods  NCCRTP operates on a generation basis and employs random linear network coding (RLNC) over the Galois field \begin{document}$ \text{GF}({2}^{8}) $\end{document}. To reduce coding overhead, each packet carries a code identifier (CodeID) rather than the full global coding vector, and relay nodes recover the corresponding coding vector from a shared local dictionary. During hop-by-hop forwarding, NCCRTP generates forward candidates under a residual-hop decreasing constraint and adaptively selects among three transmission modes: SINGLE, COOP, and BRANCH. SINGLE maintains a shared forwarding process toward the two destinations, COOP enables two relay nodes to jointly utilize linearly independent coded packets, and BRANCH splits the transmission toward different destinations. For each candidate structure, NCCRTP estimates link success probability, computes the required transmission budget, and uses a two-hop structural utility evaluation to select the forwarding mode with a better tradeoff between recovery capability and transmission cost.  Results and Discussions  Simulation results show that NCCRTP achieves the highest joint packet delivery ratio (JPDR) under both regular and random deployments. In the controlled comparison with cooperative uncoded transmission (CU), single-branch uncoded transmission (SU), and single-branch coded transmission (SC), NCCRTP consistently outperforms the schemes using only cooperative forwarding or only RLNC, indicating that the reliability gain comes from the joint effect of distributed relay cooperation and linearly independent coded-packet recovery (Fig. 5). When the packet error rate increases or the transmission depth grows from 3 to 7 hops, NCCRTP maintains a higher JPDR, demonstrating better robustness under lossy multi-hop conditions (Figs. 5(a) and 5(b)). In random deployments, vector-based forwarding (VBF) and focused beam routing (FBR) are each combined with packet replication (REP) or RLNC to form the VBF+REP, VBF+RLNC, FBR+REP, and FBR+RLNC schemes (Figs. 6 and 7). In medium-to-high packet error rate scenarios, its JPDR increases by up to approximately 50%, while the equivalent transmission cost per successful joint delivery is reduced by up to approximately 40% (Fig. 6). These results indicate that NCCRTP improves dual-sink reliability not by simply increasing redundant transmissions, but by adaptive forwarding-structure selection, link-quality-driven transmission budgeting, and joint utilization of linearly independent coded packets.  Conclusions  This paper addresses the reliability and redundancy-overhead challenges in single-source dual-sink underwater acoustic multi-hop transmission by designing a cooperative transmission structure that enables random linear network coding to exploit distributed reception and complementary coded information among relay nodes. The proposed NCCRTP protocol adaptively selects SINGLE, COOP, and BRANCH forwarding modes according to residual-hop constraints, link-quality-driven transmission budgeting, and two-hop structural utility evaluation. In addition, a lightweight coding-vector representation based on CodeID is introduced to reduce the header overhead caused by carrying full global coding vectors. The protocol is evaluated under both regular and random deployments, and the results show that: (1) NCCRTP achieves the highest joint packet delivery ratio among all compared schemes, demonstrating stronger joint recovery capability for the two destination nodes; (2) under medium-to-high packet error rates or multi-hop transmission conditions, NCCRTP improves the joint delivery ratio by up to approximately 50%; (3) the equivalent transmission cost per successful joint delivery is reduced by up to approximately 40%, indicating that the reliability gain mainly comes from adaptive structure selection, transmission-budget control, and joint utilization of linearly independent coded packets rather than excessive redundant transmissions. Future work will extend NCCRTP to more complex multi-source multi-sink multi-hop scenarios, and further investigate its implementation and performance under node mobility and realistic underwater acoustic channel dynamics.
Secrecy Performance Analysis of Multi-Tag Bistatic Backscatter Communication Systems With Outdated CSI and Link Correlation
LIU Yingting, TANG Yong, LI Xingwang
Available online  , doi: 10.11999/JEIT260823
Abstract:
  Objective  Due to feedback delay, the channel state information (CSI) used in the tag selection phase may become outdated during the actual data transmission phase, leading to a mismatch between the selected tag and the actual optimal tag. Meanwhile, most existing studies on outdated CSI are based on the independent and identically distributed (i.i.d.) channel assumption, which may fail to accurately reflect the heterogeneous characteristics of practical links caused by different propagation distances. In addition, when the eavesdropping node is close to the destination, the legitimate and eavesdropping links may experience correlated fading. Ignoring these factors may lead to theoretical results that deviate from the actual system performance. Accordingly, under independent and non-identically distributed (i.n.i.d.) channel conditions, this paper studies the secrecy performance of a multi-tag BBC system subjected to the joint impact of outdated CSI and the correlation between the legitimate link and the eavesdropping link.  Methods  In this paper, candidate tags are ranked by their backscatter-link channel gains, and the tag with the largest gain is selected for transmission. An outdated CSI model is introduced to characterize the mismatch between the CSI used for tag selection and the actual CSI during data transmission. Since tag selection depends only on the backscatter-link CSI, outdated CSI is modeled only for this link. Meanwhile, a correlated Rayleigh fading model is adopted to capture the statistical dependence between the legitimate and eavesdropping links. Under i.n.i.d. Rayleigh fading, order statistics are used to derive the PDF of the selected tag’s outdated backscatter-link channel gain in (30)–(36) and the joint PDF of the legitimate and eavesdropping channel gains in (10)–(13). Based on these results, closed-form and high-transmit-power asymptotic expressions for the secrecy outage probability are obtained, and the secrecy performance is analyzed in terms of the gain ratio between the legitimate and eavesdropping links.  Results and Discussions  Monte Carlo simulations validate the analytical and asymptotic results. The results show that outdated CSI significantly degrades secrecy performance, as feedback delay may cause the CSI used for tag selection to differ from the actual CSI during transmission, so the selected tag may no longer provide the largest backscatter-link gain. For a fixed gain ratio between the legitimate and eavesdropping links, a secrecy outage floor emerges at high transmit power (Fig. 2), indicating that increasing transmit power alone cannot eliminate the performance bottleneck. In contrast, increasing the gain ratio effectively mitigates this floor and enables a secrecy diversity order of 1 (Fig. 4). Meanwhile, under the considered system model, correlation between the legitimate and eavesdropping links also improves secrecy performance (Fig. 3) by reducing the likelihood that the legitimate link experiences deep fading while the eavesdropping link remains strong. Moreover, despite outdated CSI, the proposed gain-order-based tag selection scheme remains effective and consistently outperforms random tag selection (Fig. 3).  Conclusions  This paper develops a secrecy-performance analysis framework for multi-tag BBC systems with outdated CSI and correlated legitimate and eavesdropping links. Closed-form SOP and high-transmit-power asymptotic expressions characterize the effects of CSI outdated, link correlation, and the legitimate-to-eavesdropping gain ratio. The results identify outdated CSI as a major source of secrecy degradation, highlighting the need for low-latency feedback. Enhancing the legitimate-link gain advantage and employing backscatter-link gain-based tag selection effectively improve system secrecy performance.
A Parametric Architecture Description Framework for Embedded FPGAs and Multi-objective QoR-driven Architecture Exploration
ZHOU Jing, ZHANG Shengbing, CHEN Lei, FENG Hanxu, WANG Shuo, TIAN Chunsheng
Available online  , doi: 10.11999/JEIT260609
Abstract:
  Objective  Architecture parameters of Commercial off-the-shelf Field Programmable Gate Arrays (FPGAs) are fixed by the vendor and reused across products. Embedded FPGAs (eFPGAs) leave that choice to the architect, so the LUT input size K, the cluster size N, the interconnect topology, and the heterogeneous tiles are all configurable for the target application. This makes architecture exploration a task in design flow, where tens to hundreds of architectures may have to be drafted and compared before one is settled on. The way architectures are described today holds such exploration back: every architecture is written out by hand and re-aligned across the architecture files of multiple backend toolchains after each parameter change, no formal model keeps these files consistent, and measured QoR data spanning real application domains and heterogeneous tile combinations are rarely shared. To loosen these frictions, the work reported here treats eFPGA architecture description as a research object. The description is given a typed form on which multi-view generation and field-level checking are formally defined, and this model is realized as an open framework that drives the mainstream backend toolchains from one source description. A parameter-sweep design space exploration is then carried out on the framework, and a multi-domain QoR dataset is released alongside as a public benchmark.  Methods  HorizonArch, the parametric framework proposed in this work, organizes architecture parameters into three independence layers (Table 1). L0 captures process invariants that depend only on the technology node; L1 holds coupled parameters such as K, N, segment length, switch block type, and channel connectivity, where one change forces updates across many fields in the architecture files of different backend toolchains; and L2 holds independent parameters that the architect can set on their own. Architecture construction is formalized as one operator B that maps a parameter vector p to a complete architecture object (Fig. 4), governed by five formal rules: parameter completeness (R1), fragment independence (R2), type compatibility (R3), explicit coupling (R4), and static checkability (R5). Each backend artifact is then produced by an independent derivation function that consumes the same source object, so adding a new tool requires only one new view rather than edits across all existing ones. Field-level validation runs at load time, and cross-field references are checked before any backend is invoked. The class structure (Fig. 3) splits the architecture into synthesis, circuit, and layout views, with each semantic element declared only once. Three extension levels keep the framework open: G1 adds a black-box module, G2 enlarges the value set of an existing coupled parameter, and G3 introduces a new coupled parameter together with its constrained value set. Built on this framework, a parameter-sweep design space exploration method scans the (K, N) grid and the heterogeneous tile combinations, and assesses each configuration using three QoR metrics: area, critical-path delay (CPD), and area-delay product (ADP), where ADP is a derived composite metric calculated from area and CPD.  Results and Discussions  End-to-end validation confirms that one source description drives VPR, OpenFPGA, and Yosys consistently; COFFE is validated through interface integration and SPICE startup (Table 4, Table 5), and the G1, G2, and G3 extension experiments all pass the cross-field checks. The design space exploration study then reveals a clear divergence among the three QoR metrics on the (K, N) plane (Fig. 5, Fig. 6). Area is dominated by K and favors small K, and most circuits place their area optimum at K=4 N=4. CPD, in contrast, is mildly negatively correlated with both K and N, and its optimum clusters near (8, 10) and (7, 10). The relative range across the twenty (K, N) points (Table 7) is much larger for area and ADP than for CPD, so parameter choice gives a much bigger lever on cost than on raw delay. A study of default configurations (Table 8) shows that K=4 N=4 reaches the per-circuit ADP optimum on 68% of circuits and stays within about six percent on average, while its CPD deviates by close to fifty percent; K=8 N=10 takes the opposite role, with its CPD deviation cut to 5.9% and 30.8% of circuits hit exactly, but its area and ADP deviations soar past 670% and 480%. A random-forest cross-domain surrogate reaches about 65% top-5 accuracy, so a parameter sweep is still needed when the design target is tight.  Conclusions  This work presents HorizonArch, a formal multi-backend parametric architecture description framework for eFPGA exploration. The framework drives VPR, OpenFPGA, and Yosys from one source description and exposes an extensible COFFE interface. The parameter-sweep exploration carried out on the framework shows that area and CPD favor opposite directions in the (K, N) plane, which means eFPGA architecture parameters should be chosen against an explicit design target rather than a default value. An open QoR dataset covering five application domains is released alongside as a public benchmark. Future work will complete the COFFE SPICE rewriting component, refit the routing-area coefficient with measured data, and explore more efficient exploration strategies on top of the framework.
Adaptive Fusion Detection for Dual-Radar with Non-Identical Clutter Distributions
ZHOU Baoyi, YANG Yong, YANG boyu
Available online  , doi: 10.11999/JEIT260616
Abstract:
  Objective  Small sea-surface targets, characterized by low radar cross-sections, generate echoes easily marked by intense sea clutter, resulting in extremely low signal-to-clutter ratios (SCR). Constrained by single operating frequency bands and fixed observation angles, single radar systems present notable performance bottlenecks, whereas multi-radar collaborative detection offers an effective solution to this limitation. Existing multi-radar fusion detection algorithms are derived under two ideal premises: identical clutter distributions and equal SCRs across radar nodes, assumptions that rarely hold in practice. In practice, discrepancies in radar parameters (frequency band, range resolution, and grazing angle) result in distinct statistical characteristics of sea clutter. Furthermore, target RCS fluctuates with observation azimuth and operating frequency, yielding inconsistent SCRs for the same target across different radars and inducing severe model mismatch in traditional fusion detectors. To address the coexistence of heterogeneous clutter distributions and unequal SCRs, this paper proposes a Neyman-Pearson (NP) criterion-based dual-radar adaptive fusion detector, termed NP-AFD, for Rayleigh and Weibull sea clutter backgrounds.  Methods  Amplitude distribution fitting is conducted on measured S-band and X-band sea clutter datasets using five mainstream models: Rayleigh, Lognormal, Weibull, Gamma, and K-distribution. Fitting accuracy is evaluated via the mean square error (MSE) of the probability density function (PDF) and complementary cumulative distribution function (CCDF), as listed in Table 1. Based on the fitting results, local optimal test statistics are derived separately for Rayleigh and Weibull clutter backgrounds. As the two radars operate independently, the joint likelihood ratio equals the product of their individual likelihood ratios. The optimal fusion statistic is formulated as a weighted sum of the two local statistics, where adaptive weights are determined by real-time estimated SCRs, and radar channels with higher SCRs are assigned larger weights. Closed-form analytical expressions linking the false alarm probability and detection probability to the decision threshold are also derived.  Results and Discussions  Monte Carlo simulations over 105 independent trials verify the validity of all derived closed-form expressions. Under identical SCRs for both radars, the simulated detection curves of single radars and NP-AFD show strong agreement with theoretical curves (Fig. 3). Compared with decision-level OR and AND fusion, NP-AFD consistently achieves the highest detection probability, with OR fusion ranking second and AND fusion delivering the worst. Performance gain analysis shows that NP-AFD maintains a positive gain over the better-performing single radar across all tested Weibull shape parameters. In contrast, OR fusion exhibits negative gain at low SCRs under small Weibull shape parameters, while AND fusion maintains negative gain across the full SCR range (Fig. 4). The simulated false alarm rate is stably controlled around the preset value of 10–3 (Fig. 5). Furthermore, two-dimensional joint evaluation is performed with independently varying SCRs of the two radars, and detection probability contours are plotted with the two radar SCRs as coordinates (Fig. 6). The results demonstrate that NP-AFD requires lower SCR combinations to achieve the same detection probability, compared with OR and AND fusion. Experiments on measured sea clutter data further verify that NP-AFD yields the highest detection probability among all compared methods (Fig. 7, Fig. 8), and its detection performance shows agreement with theoretical predictions (Fig. 9, Fig. 10).  Conclusions  This paper addresses the challenges of non-identical clutter distributions and varying SCRs in multi-radar collaborative detection. Based on the Neyman-Pearson criterion, a dual-radar adaptive fusion detection method, NP-AFD, is derived for Rayleigh and Weibull clutter backgrounds. The proposed method provides closed-form expressions for fusion weights, decision threshold, and detection probability. Theoretical derivations, simulations, and experiments on measured sea clutter data consistently demonstrate that NP-AFD outperforms OR fusion and AND fusion under arbitrary SCR combinations, delivering superior detection performance and robustness.
Communication Performance and Fault Degradation Analysis of Boundary-Interface Configurations in 2.5D Chiplet Systems
HOU Shuaikang, LIU Qinrang, LV Ping, LIU Zhengyu, XU Yuhang, LI Peijie, GUO Wei
Available online  , doi: 10.11999/JEIT260633
Abstract:
  Objective  Chiplet-based integration has emerged as an important approach for constructing large-scale heterogeneous systems. In a 2.5D chiplet system, inter-chiplet packets travel from a source node to a boundary interface in the source chiplet, traverse an interposer network, and enter the destination chiplet through a destination-side interface. Because packaging resources, micro-bump count, routing resources, and power and area budgets are limited, interfaces can be deployed only at a subset of boundary routers. Their number and distribution affect access distance, service-region formation, interface-load distribution, and traffic remapping after failures. Existing studies have addressed die-to-die standards, interposer architectures, placement, routing, simulation, and fault tolerance, but the independent structural impact of boundary-interface configuration remains insufficiently characterized. This work investigates how interface count, location, service-region partitioning, and failures affect communication performance and fault-induced degradation in 2.5D chiplet–interposer networks.  Methods  A graph-level structural model is developed for a 2.5D chiplet–interposer network. Each chiplet is modeled as an n×n mesh, and selected boundary routers serve as active interfaces. Each node is mapped to its nearest active interface, and nodes mapped to the same interface form a service region. Average interface-access distance, abstract end-to-end hop count, interface-load variation, maximum load ratio, and fault-induced degradation are evaluated. Three chiplet scales, n=4, 8, and 16, are considered, with interface budgets of k=2, 4, 6, k=4, 8, 12, and k=8, 16, 24, respectively. Five layouts are analyzed: Uniform-corner, Balanced-edge, Two-edge-centered, Single-edge-clustered, and Greedy-DL. Uniform, Hotspot, Transpose, Tornado, and Neighbor traffic patterns are considered. For fault analysis, each active interface is removed in turn, and affected nodes are remapped to their nearest healthy interfaces. A cycle-accurate gem5/Ruby GARNET model comprising four 4×4-mesh chiplets and a 4×4 interposer mesh is constructed. Uniform and Hotspot traffic are evaluated from 0.01 to 0.15 flits/node/cycle. The latency-based saturation injection rate is the first sampled rate at which average packet latency exceeds 50 cycles. This two-level evaluation separates structural effects from cycle-level network behavior and enables interface count, placement, traffic pattern, and fault location to be compared under consistent topology and routing assumptions for mechanism-oriented analysis.  Results and Discussions  Graph-level results show that increasing the number of boundary interfaces reduces average interface-access distance and abstract end-to-end hop count, while the marginal benefit diminishes as boundary coverage becomes sufficient (Fig. 3). Under a fixed interface budget, different interface locations induce distinct service-region partitions and interface-load distributions (Table 2, Fig. 4). For an 8×8 chiplet with k=8 under Uniform traffic, Greedy-DL achieves the smallest average interface-access distance of 3.125, while Balanced-edge provides a favorable compromise, with an access distance of 3.500 and a maximum load ratio of 1.250. Single-edge-clustered produces the largest access distance and abstract end-to-end hop count because its interfaces are concentrated on one boundary. Under Hotspot traffic, Two-edge-centered reduces interface-load variation to 0.016, but its access distance remains higher than those of Balanced-edge and Greedy-DL, confirming a trade-off between distance and load balance (Table 2, Fig. 5). Interface failures cause only moderate changes in average access distance but substantial load concentration at the remaining healthy interfaces. Under Hotspot traffic, the worst maximum-load-ratio degradation is 23.94% for Balanced-edge and 93.75% for both Two-edge-centered and Single-edge-clustered (Table 3), indicating that vulnerability is governed mainly by service-region migration and load reconcentration.The gem5 results further support the graph-level observations. When the interface count increases from k=1 to k=4, low-load latency decreases from 24.50/24.44 cycles to 18.05/18.09 cycles under Uniform and Hotspot traffic, respectively, while the latency-based saturation injection rate increases from 0.06 to 0.14 (Fig. 6). Across the k=4 layouts, Balanced-edge achieves low latency and hop count, whereas Single-edge-clustered has the highest baseline cost. Under Hotspot traffic, Balanced-edge reduces low-load latency by about 14.3% and average hop count by about 18.7% compared with Single-edge-clustered. Greedy-DL provides short low-load paths, but its saturation injection rate is 0.13, lower than the 0.14 achieved by Balanced-edge and several other layouts, showing that minimizing access distance alone does not guarantee stronger medium- and high-load behavior (Fig. 7). For fault validation, all 16 single-interface failure locations are traversed for each of three layouts, yielding 48 scenarios. The results show increases in latency and hop count, together with an earlier onset of congestion. The worst-case latency-based saturation injection rate decreases to 0.10–0.11, and Balanced-edge exhibits smaller average and worst-case degradation than the more concentrated layouts (Fig. 8, Table 5).  Conclusions  Boundary-interface configuration is a key structural parameter in 2.5D chiplet interconnect design. It affects interface-access distance, service-region formation, interface-load distribution, and fault-induced traffic remapping. Increasing interface count improves communication efficiency, but with diminishing benefit. Under the same interface budget, interface placement determines whether traffic remains balanced or becomes concentrated after mapping and remapping. The graph-level metrics support low-cost screening and mechanism analysis, while gem5/Ruby GARNET simulations provide cycle-accurate validation under buffering, arbitration, and flow-control effects. Therefore, interface count, location, load balance, and fault degradation should be evaluated jointly in 2.5D chiplet–interposer network design.
An Adaptive Kalman Speech Enhancement Method Driven by Burst Noise Suppression and Dual-Time-Scale Perception
CHEN Bo, ZHENG ZeRui, SUN Chao, WANG ZheMing, SHEN Ying
Available online  , doi: 10.11999/JEIT260636
Abstract:
  Objective  The traditional Auto-Regressive (AR) Kalman speech enhancement algorithm has three critical drawbacks under non-stationary noise: difficult adaptive noise model update, easy model mismatch caused by burst noise, and lack of environment-aware covariance adjustment. These defects degrade enhancement performance and cannot satisfy practical speech communication demands. To tackle the above issues, this paper proposes an improved adaptive Kalman method for better speech quality and intelligibility under complex noise.  Methods  First, an environment deviation ratio based on Energy Entropy Ratio (EER) is constructed to measure statistical deviation between the current frame and background. A dual-time-scale EER tracker is built to capture instantaneous fluctuations and steady background statistics respectively, and their difference generates an adaptive intensity factor for joint adjustment of process and observation noise covariances. Second, combining burst noise features, a two-stage discrimination scheme is presented: energy mutation threshold combined with spectral flatness detects impulsive noise; a Speech Modelling Metric derived from linear prediction residual variance further separates speech from medium-energy burst noise and avoids AR model contamination.  Results and Discussions  Experiments on NOIZEUS dataset show the proposed method outperforms classic AR-Kalman and J1-sensitivity based improved Kalman in STOI, PESQ and SegSNR. It gains better adaptability and stability under non-stationary and burst noise. Dual-time-scale perception and accurate burst identification accelerate environmental adaptation and reduce speech distortion.  Conclusions  The burst-suppressed dual-time-scale adaptive Kalman algorithm solves inherent defects of traditional AR-Kalman in complex noise. EER-based tracking realizes environment-aware covariance tuning, while the two-stage judgment greatly improves burst noise detection accuracy. Objective results verify its strong robustness and provide a feasible scheme for practical non-stationary speech enhancement.
An Anomalous Traffic Detection Method Combining Stream Data Compression and Self-Supervised Graph Learning
XIA Jiqiang, ZHAO Jianjin, WANG Zihao, TIAN Le, HU Yuxiang, LI Menglong
Available online  , doi: 10.11999/JEIT260118
Abstract:
  Objective  With the continuous growth of network traffic scale and the increasing sophistication of attack methods, efficient and intelligent anomalous traffic detection has become essential for protecting critical information infrastructure. However, existing detection methods still face severe challenges when deployed in large-scale networks. On one hand, the analysis of raw packet sequences and deep learning-based end-to-end models incurs substantial computation and storage overhead, which is unaffordable in line-rate processing scenarios. On the other hand, flow records are usually treated as independent samples, and the topological structure and context information of inter-host communications are ignored, resulting in the lack of a global view when distributed and correlated threats are handled. Moreover, supervised learning schemes rely heavily on large amounts of labeled data, which can hardly be satisfied in practical deployment and weakens the generalization ability against unknown threats. Therefore, it is of great significance to develop an anomalous traffic detection method that supports efficient flow feature extraction under limited resources and achieves high-precision detection without relying on labeled data.  Methods  In the training phase, the graph encoder takes both the original communication graph and augmented negative samples as inputs to learn edge embeddings. The encoder outputs these embeddings to a discriminator. The discriminator estimates mutual information by contrasting edge embeddings of positive and negative samples against a global graph summary. The training objective maximizes scores for positive samples while minimizing them for negative ones. Through this self-supervised optimization, the encoder parameters are refined to enhance the discriminative power of edge embeddings. This iterative process continues until convergence via gradient descent. After training, the encoder parameters are fixed. The resulting edge embeddings then serve downstream anomaly detection tasks. In the inference phase, the feature extractor converts traffic into a communication graph. The trained graph encoder then generates corresponding edge embeddings. A lightweight classifier takes these embeddings as input for end-to-end anomaly detection, producing the final classification results.  Results and Discussions  Comprehensive experiments are conducted on four public datasets, i.e., CAIDA, CIC-IDS2018, UNSW-NB15, and TON-IoT. For feature extraction, under identical memory configurations, the average relative error (ARE) and per-flow weighted mean relative error (WMRE) of counter-type features measured by the optimized MFSketch are reduced by 31.5% and 31.0% on average compared with the fixed-structure baseline, and those of bitmap-type features are reduced by 36.1% and 34.9%, respectively (Fig. 4). Meanwhile, high throughput is maintained on datasets with different traffic skewness, and an average of about 12Mpps is reached on the CAIDA dataset (Fig. 4(c)). For detection accuracy, SketchGNN combined with PCA, HBOS, or IF stably achieves an accuracy no lower than 95.2%, a macro-F1 no lower than 90.1, and a weighted-F1 no lower than 96.7 on CIC-IDS2018 and UNSW-NB15, which outperforms Kitsune, Whisper, Anomal-E, and TS-IDS, whose accuracy either stays below 90% or fluctuates sharply across datasets (Fig. 5, Fig. 6). For detection efficiency, HBOS provides the most stable and highest throughput among the three classifiers (Fig. 7(a)), and the end-to-end packet-level equivalent throughput of SketchGNN reaches 640Kpps, which is about 17 times that of Kitsune (37Kpps) and is comparable in magnitude to the DPDK-accelerated Whisper (1.3Mpps) (Fig. 7(b)). In addition, the performance variation across different datasets is kept within 3%, demonstrating robust generalization to normal traffic fluctuations and diverse flow-level anomalous behaviors.  Conclusions  To address the challenges of high feature-extraction overhead, limited utilization of traffic context, and heavy reliance on labeled data, this paper proposes SketchGNN, an anomaly detection framework that integrates flow data compression with self-supervised graph learning. MFSketch, a dynamically configurable sketch, efficiently extracts and accurately measures diverse flow features under limited resources. Self-supervised graph neural networks then model host communication graphs and learn representations, enabling efficient anomaly detection without labeled data. Experimental results show that MFSketch can dynamically optimize its data structure based on traffic distribution, ensuring high-throughput, high-precision feature inputs for downstream detection tasks. The edge embeddings generated via graph-based self-supervised learning achieve higher detection accuracy across various unsupervised classifiers compared to baseline methods. The future work will further explore hybrid detection mechanisms that combine Deep Packet Inspection with programmable data planes to further enhance the ability to identify anomalous traffic at the application layer.
A Reconfigurable Parallelised Coprocessor Design for the RISC-V-Based Grain Cryptographic Algorithm
NAN Longmei, WANG Haoyu, DU Yiran, LI Wei, CHEN Tao
Available online  , doi: 10.11999/JEIT260391
Abstract:
  Objective  To overcome the performance bottleneck of Grain cryptographic algorithms on general-purpose processors (GPPs), as well as the inflexibility and high overhead of application-specific inte-grated circuit (ASIC) implementations, this work integrates dedicated cryptographic hardware accel-erators into a RISC-V coprocessor via a custom instruction extension approach. A reconfigurable and parallelisable implementation architecture is proposed specifically for the Grain algorithm family, lev-eraging the RISC-V coprocessor interface. Corresponding custom instructions are designed to enable flexible and efficient execution of Grain-80, Grain-128, Grain-128a, and Grain-128AEAD on the same hardware platform. The proposed architecture achieves a favourable trade-off between processing throughput and hardware flexibility, making it suitable for resource-constrained embedded systems.   Methods  This paper employs a unified feedback shift register architecture to support flexible swi-tching between the Grain-80, Grain-128, and Grain-128a, Grain-128AEAD algorithms, while enabling parallelisation granularity to be flexibly configured between 1 and 8 steps, thereby further enhancing throughput and resource utilisation, When extending specialised instructions for the Grain cryptogra-phic algorithm, the software and hardware functional modules are analysed, dividing the entire crypt-ographic process into software and hardware components to efficiently complete the encryption proc-edure. The coprocessor implementation proposed in this study features a streamlined architecture, ac-hieving reconfigurable parallel realisations of all four Grain algorithms with minimal hardware resou-rce expansion.   Results and Discussions  By invoking specialised instructions for the reconfigurable parallelised Grain cryptographic algorithm, three distinct cryptographic algorithms can be flexibly implemented. Comp-aring results with nonextended instructions (Table 7) demonstrates that specialised instructions enh-ance both the execution rate of cryptographic algorithms and reduce the number of required instruc-tions. Contrasting with results from other literature (Table 8) reveals the advantages of this approach in resource reuse and design flexibility. Compared to purely software implementations, this approach achieves significant optimisation in both the number of executed instructions and the number of ope-rational cycles. When contrasted with general-purpose reconf-igurable hardware implementations, this solution demonstrates superior resource utilisation.   Conclusions  To address the demand for agile deployment and resource-efficient implementation of cryptographic algorithms in lightweight embedded systems, this paper designs a hardware-software co-operative, reconfigurable parallelisation scheme for accelerating the Grain cryptogra-phic algorith-m. This approach leverages the RISC-V coprocessor instruction extension mechanism. Through a unif-ied shift register architecture, configurable feedback tap selection network, and feedback logic tailored for a deterministic algorithm set, the solution enables dynamic switching and parallel processing of four algorithms—Grain-80, Grain-128, Grain-128a and Grain-128AEAD—on a single hardware platfo-rm. This approach effectively reduces hardware resource consumption while maintaining functional flexibility. The present work primarily focuses on the reconfigurable parallelisation of the Grain algor-ithm family. Future research will explore universal reconfigurable architectures for nonlinear Boolean functions, incorporating reconfigurable units such as lookup tables (LUTs) or programmable logic ar-rays. By flexibly adapting to multiple stream cipher algorithms through configuration information, th-is approach aims to enhance algorithm compatibility and scalability within hardware modules while maintaining high throughput and low resource overhead.
SkipSync: Accelerating Instruction Sampling for LLM Workload
CAI Luoshan, ZHOU Yaoyang, WANG Kaifan, LIU Tianyi, SUN Ninghui, BAO Yungang
Available online  , doi: 10.11999/JEIT260397
Abstract:
  Objective   The rapid evolution of Large Language Models (LLMs) has significantly increased the demand for efficient design and evaluation of domain-specific accelerators. While sampling-based performance evaluation method effectively reduces the cost of cycle-accurate simulation, they rely on Instruction Set Simulators (ISS) to execute full workloads for profiling and checkpoint generation. In LLM scenarios, frequent changes in model structures, inference frameworks, and accelerator instruction sets invalidate checkpoint reuse and substantially increase the overhead of ISS-based functional simulation, which becomes the dominant bottleneck in the evaluation workflow. Existing ISS acceleration techniques, such as virtualization and dynamic binary translation (DBT), either require ISA compatibility or incur high implementation complexity, making them unsuitable for rapidly evolving LLM workloads. Therefore, a flexible and efficient ISS acceleration approach is required to support fast and accurate sampling-based evaluation for LLM accelerators.  Methods   This paper proposes SkipSync, an ISS acceleration method for LLM inference workload sampling. The key insight is that the control flow of most LLM operators is independent of intermediate tensor values and determined by static parameters. Therefore, SkipSync introduces Skip Execution to bypass the execution stage of accelerator instructions within selected operators while preserving instruction fetch and decode stages to maintain profiling correctness. To guarantee correct control flow of subsequent execution, a lightweight Host–Simulator Synchronization mechanism is further introduced to synchronize host-computed results back to the ISS. Lightweight synchronization primitives and custom instructions are designed to integrate SkipSync into existing sampling workflows with minimal implementation effort.  Results and Discussions   SkipSync is implemented on QEMU and supports RISC-V matrix and vector extensions as accelerator instructions. Experimental results show SkipSync significantly reduces ISS simulation overhead while preserving accuracy. Compared with the baseline, it achieves an average speedup of 4.64× in the Prefill stage and 5.51× in the Decode stage (Fig. 5), as it eliminates the dominant cost of vector and matrix instructions execution, which accounts for over 80% of runtime (Table 1). SkipSync outperforms DBT-based approaches (Fig. 6), while offering better extensibility to support new accelerator instructions. Notably, the synchronization mechanism introduces only limited overhead, adding approximately 8.5% additional runtime relative to ideal skip-only execution (Fig. 8). This confirms that the data synchronization does not negate the acceleration benefits. Meanwhile, SkipSync maintains high sampling fidelity, with an average error of approximately 2.55% and performance results nearly identical to baseline execution (Fig. 9).  Conclusions  This paper presents SkipSync, an extensible ISS acceleration method for sampling-based performance evaluation of LLM inference workloads. By combining Skip Execution and Host–Simulator Synchronization, SkipSync effectively alleviates the ISS bottleneck in LLM sampling workflows while preserving correctness and sampling fidelity. Experimental results demonstrate that SkipSync achieves significant speedup with low synchronization overhead and negligible impact on sampling accuracy. The proposed method provides a practical solution for efficient accelerator evaluation required by rapidly evolving LLM ecosystems. Future work will extend SkipSync to more complex LLM inference optimization scenarios and explore automatic identification of skippable operators to further improve usability.
Intelligent Detection of DSSS Signals Under False-Alarm Rate Constraints Based on a Noise Score-Pool Threshold Calibration Mechanism
ZHANG Tao, TANG Xiaomei, SUN Guangfu
Available online  , doi: 10.11999/JEIT260414
Abstract:
  Objective  To address the performance degradation of GNSS signal detection in weak-signal environments and the lack of false alarm rate (FAR) control in existing deep learning models, this study investigates a detection method under a constant false alarm rate (CFAR) constraint to enhance receiver reliability and practicality.  Methods  This study models the detection of direct sequence spread spectrum signals as a binary classification problem and proposes a comprehensive DL-based detection framework. The methodology is centered on three core innovations. First, an adaptive one-dimensional residual neural network (1D-ResNet-18) is designed to suit the characteristics of I/Q sampled time-series data. Key modifications include adjusting the input convolution kernels to 1×3 and removing the initial maximum pooling layer to prevent the loss of fine-grained features inherent in weak signals. Second, a "noise score pool threshold calibration" mechanism is introduced. By inputting a large volume of pure noise samples into the trained network, an empirical distribution of confidence scores for the "signal present" category is constructed. Decision thresholds are then dynamically determined based on the quantiles corresponding to preset FAR levels. Third, an unnormalized data preprocessing strategy is adopted, as it is demonstrated that preserving the original signal amplitude information is beneficial for network learning under low signal-to-noise ratio (SNR) conditions. The model's performance was rigorously validated using a simulated GPS L1 C/A signal dataset under various SNRs, FAR settings, and non-ideal colored noise environments.  Results and Discussions  Experimental results demonstrate significant performance gains achieved by the proposed method. At a FAR of 0.01, the detection probability approaches 100% at an SNR of –8 dB, marking a substantial improvement over traditional techniques. A systematic comparison of different FAR settings (0.0001, 0.001, and 0.01) indicates that while the detection probability curve predictably shifts as the FAR decreases, overall performance remains consistently high. Furthermore, the unnormalized data preprocessing strategy consistently outperformed the normalized strategy, introducing a performance gain of approximately 1 dB in the low SNR range. Compared with traditional autocorrelation detection methods, the deep learning model exhibits a significant overall detection performance improvement of 3 to 4 dB across various false alarm rates. Notably, the method also displayed robust generalization in untrained, non-ideal colored noise environments.  Conclusions  The proposed method effectively bridges data-driven deep learning with classical detection theory, resolving the lack of FAR control in neural networks. By balancing high sensitivity with precise false alarm management, this work provides a practical and robust framework for navigation signal processing in complex environments.
Efficient Hyperdimensional Computing Accelerator Design for Chinese Text Classification Task
YU Tianyang, WU Bi, LIU Weiqiang
Available online  , doi: 10.11999/JEIT260556
Abstract:
  Objective  With the proliferation of edge computing in IoT, smart wearables, and offline terminals, low-latency and high-privacy Chinese text analysis on local devices has become a core requirement. While neural network-based and Transformer-based language models achieve excellent accuracy, their massive parameter scales and computational demands make them unsuitable for power- and storage-constrained edge devices. Hyperdimensional Computing (HDC), an emerging brain-inspired paradigm, represents text through 2K-10K dimensional hypervectors and employs lightweight encoding-query mechanisms instead of complex multi-layer networks, offering a hardware-friendly approach for edge-side text tasks. However, existing HDC research has focused exclusively on alphabetic writing systems such as English, where a limited set of letters suffices for N-gram encoding. For Chinese, with its thousands of commonly used characters, directly applying existing methods would cause prohibitive storage overhead from the explosion of base hypervectors, fundamentally undermining the lightweight advantage of HDC. This paper aims to overcome this bottleneck by proposing a novel character encoding method and a dedicated HDC framework with hardware acceleration for Chinese text classification.  Methods  Drawing inspiration from the structural characteristics of Chinese characters as ideographic symbols, this paper proposes a hyperdimensional encoding method based on character glyph structure. Specifically, the method reverse-engineers the encoding process of the Wubi input method, decomposing each Chinese character into its constituent strokes and radicals and mapping them to equivalent letter sequences for efficient hyperdimensional encoding. A retrieval library covering 3,500 commonly used Chinese characters, as defined in the General Standard Chinese Characters Table issued by the Ministry of Education of the People's Republic of China, is constructed. For each character, the Wubi library is queried to obtain a letter sequence of length 3 or 4, where each letter position corresponds to a base hypervector. Through cyclic shift and binding operations applied to the base hypervectors of consecutive letters, the character is encoded into a character-level hypervector that simultaneously captures both letter identity and sequential order information, effectively preventing confusion arising from different letter permutations. This approach reduces the storage overhead associated with base hypervectors by more than 95% compared to the method of directly assigning independent hypervectors to each Chinese character (Table 1). Building upon this character encoding scheme, the HDChinese framework is developed to support both training and inference procedures. During inference, the sentence-level hypervector, obtained by bundling all character hypervectors, queries the class hypervectors via cosine similarity, which is efficiently realized through XNOR-based inner product computations. During training, class hypervectors are derived by bundling sentence hypervectors belonging to the same category and applying a sign-based binarization operation. A subsequent retraining phase iteratively fine-tunes the non-binary class hypervectors using misclassified samples with a learning rate of 0.25, thereby mitigating the influence of outlier hypervectors on the cluster center. Furthermore, a dedicated hardware accelerator architecture tailored to HDChinese is designed. The architecture comprises three main modules: an encoding module, a querying module, and a class hypervector update module, which can be selectively activated to support inference-only, training, and retraining modes respectively (Fig. 2). The accelerator is prototyped on an Ultra96v2 FPGA development board featuring a Xilinx ZYNQ SoC.  Results and Discussions  Four open-source datasets covering both binary and multi-class tasks are used (Table 2). The Wubi-based encoding outperforms the Pinyin-based method in accuracy while saving 31.46% storage (Table 4), and ignoring rare characters (average frequency below 0.5%) has negligible accuracy impact (Table 4). A dimension of 2K is adopted for hypervectors as higher dimensions yield no significant gains (Fig. 3). FPGA measurements show HDChinese achieves a model size of approximately 16 KB—over 99% smaller than kNN, SVM, and Random Forest (Table 6). Inference time is reduced by 6.67%-47.25% and total training time by 4.91%-77.36% compared to SVM and Random Forest, while accuracy after retraining reaches comparable levels (Fig. 5, Table 6). Additionally, the KB-scale model enables operation without external memory chips, unlike comparison algorithms requiring over 400 MB of runtime memory. At 100 MHz, power consumption is 0.278 W, representing the optimal throughput-power trade-off (Table 5).  Conclusions  This paper proposes a hyperdimensional encoding method based on Chinese character glyph structure, resolving the incompatibility between existing HDC methods and Chinese text. The HDChinese framework and its dedicated hardware accelerator achieve over 99% model complexity reduction while maintaining competitive accuracy, with training time reductions of 4.91%-77.36% and inference time reductions of 6.67%-47.25%. The proposed approach effectively balances accuracy with extremely low overhead, providing an ideal hardware solution for Chinese text analysis on edge devices.
Space-time-coding metasurface Enabled Integrated Design of Radar Communication and Electromagnetic Stealth
ZHANG Ming, WANG Zhe, WANG Boya, YANG Lin, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260536
Abstract:
  Objective  To address the complexity and limited modulation of existing reconfigurable metasurfaces integrating radiation and stealth, this study proposes a single-layer space-time-coding metasurface that dynamically switches between beam scanning and RCS reduction. By periodically modulating meta-atom states in the time domain, the phase responses of incident and harmonic waves are differentiated, enabling dynamic integration of radiation and stealth without complex feeding networks or multilayer structures. The design achieves precise beam scanning and excellent RCS reduction, providing a simple, low-cost solution for integrating radar communication and electromagnetic stealth.  Methods  This metasurface adopts a metal-dielectric-metal structure, with each meta-atom integrating a PIN diode. Electromagnetic simulations are conducted using CST Microwave Studio. At the optimal operating frequency, the meta-atom exhibits a 1-bit tunable reflection phase and a near-lossless co-polarized reflection amplitude. By feeding space-time-coding sequences into the metasurface model, dynamic control of beam scanning and RCS reduction is achieved. A prototype is fabricated using PCB technology, and its performance is validated through echo scattering measurements conducted with a vector network analyzer and a frequency offset option.  Results and Discussions  Simulations and measurements confirm that the proposed space-time-coding metasurface enables precise beam steering and excellent RCS reduction. At 10.0 GHz, the meta-atoms provide a 180° reflection phase difference and near-lossless amplitude. In radiation mode, harmonic beams are steered to angles of +2°, +14°, +32°, +46°, –15°, –28°, and –44°, with side-lobe levels for the ±3rd to fundamental harmonics suppressed to –10.59 dB. In scattering mode, the RCS of the fundamental echo is reduced by 11.56 dB compared to a copper plate. Additionally, a –10 dB RCS reduction is achieved from 9.9 to 10.1 GHz, peaking at 14.83 dB at 9.95 GHz. Experimental results validate the design’s integrated radiation-scattering capability and dynamic control flexibility.  Conclusions  This study presents a design method for a space-time-coding metasurface that enables dynamically integrated manipulation of radiation and scattering. By employing meta-atoms integrated with PIN diodes, the proposed method periodically modulates the operating states of the meta-atoms on a single-layer metasurface to achieve differentiated distributions of the wavefront phases of various harmonics. The proposed scheme significantly reduces design complexity and manufacturing costs. The successful realization of beam scanning and RCS reduction highlights the great potential of this technology in the integrated design of radar communication and electromagnetic stealth.
A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar
HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang
Available online  , doi: 10.11999/JEIT260695
Abstract:
  Objective  Traditional optimization-based power allocation algorithms for collocated MIMO radar have two fundamental limitations. First, they optimize tracking performance only for the next time step and therefore lack a full-time-horizon view of the power allocation process. This myopic strategy cannot achieve optimal multi-target tracking accuracy over extended periods, particularly when target trajectories vary substantially. Second, these algorithms rely on iterative nonlinear constrained optimization, resulting in high computational complexity. Therefore, they cannot satisfy the real-time requirements of dynamic battlefield environments where target states change rapidly. To address these limitations, this paper proposes a Reinforcement Learning (RL)-driven power allocation algorithm. Unlike conventional methods, the proposed approach formulates the power allocation problem as a Markov Decision Process (MDP) that maximizes long-term cumulative tracking accuracy. The algorithm adaptively allocates limited transmit power among multiple beams according to the current system state, balancing immediate tracking performance with long-term cumulative tracking accuracy.  Methods  The Posterior Cramér-Rao Lower Bound (PCRLB) is employed to quantify the theoretical lower bound of the tracking error for each target. The state space is constructed by combining the motion states (position and velocity) of all targets with the normalized PCRLB from the previous allocation step. The action space consists of discrete transmit power levels for each beam, subject to the total power budget and individual beam power constraints. All feasible power allocation vectors are enumerated and encoded to reduce the action-space dimensionality. The reward function is defined as the negative weighted sum of the normalized PCRLB, encouraging the agent to minimize tracking errors. The power allocation process is formulated as an MDP and solved using the Dueling Double Deep Q-Network (D3QN) algorithm. The D3QN framework incorporates three major enhancements: (1) a double-network architecture comprising an online Q-network and a target Q-network to improve training stability; (2) a dueling architecture that decomposes the Q-value into a state-value function and an action-advantage function to improve action discrimination; and (3) off-policy learning with experience replay to improve the use of historical trajectories. An ε-greedy strategy is adopted for exploration, with ε gradually decreasing during training. After offline training, the learned network directly generates real-time transmit power allocation decisions from the current system state without iterative optimization.  Results and Discussions  Simulations are conducted using three targets following the Constant Velocity (CV) model. Fixed power allocation yields the lowest tracking accuracy because of inefficient resource utilization. The traditional optimization method, which minimizes the instantaneous tracking error, achieves moderate tracking performance but remains myopic. When the discount factor \begin{document}$ \gamma =0 $\end{document}, the D3QN algorithm achieves performance comparable to that of the traditional optimization method because both optimize only immediate rewards. In contrast, when \begin{document}$ \gamma =0.99 $\end{document}, the D3QN algorithm significantly improves full-time-horizon tracking accuracy. The resulting power allocation strategy allocates more transmit power to distant, low-Signal-to-Noise Ratio (SNR) targets at earlier stages while reducing redundant power assigned to nearby high-SNR targets. The training curves show that \begin{document}$ \gamma =0.99 $\end{document} achieves a higher steady-state cumulative reward, although convergence exhibits greater oscillation because of the increased difficulty of estimating long-term returns. Furthermore, the trained D3QN network generates transmit power allocation decisions almost instantaneously, whereas the traditional optimization method must solve a constrained optimization problem at every time step, providing a substantial real-time computational advantage.  Conclusions  This paper proposes an RL-driven power allocation algorithm for collocated MIMO radar multi-target tracking that overcomes the myopic behavior and high computational complexity of conventional optimization methods. The proposed algorithm constructs the state space and reward function using the PCRLB, models the power allocation process as an MDP, and solves it using the D3QN algorithm. Simulation results demonstrate that, with an appropriate discount factor (\begin{document}$ \gamma =0.99 $\end{document}), the proposed approach significantly improves full-time-horizon tracking accuracy. This improvement results from the agent’s ability to learn a long-term optimal policy that proactively allocates transmit power to future distant, low-SNR targets. Furthermore, the trained network enables real-time decision-making through direct forward propagation, substantially reducing computational latency compared with iterative optimization. This work provides a new approach for intelligent radar resource management in complex battlefield environments.
Hierarchical Attention Mechanism-based Path Planning for Multi-UAV Inspection
FEI Bowen, XING Wenjie, LIU Daqian
Available online  , doi: 10.11999/JEIT260192
Abstract:
  Objective  In modern power inspection, the use of multiple Unmanned Aerial Vehicles (UAVs) for cooperative inspection is an efficient but challenging task. Existing multi-UAV path planning methods often have limited cooperative scheduling capability. They also fail to accurately capture the topological relationships among heterogeneous nodes, especially those between device nodes and charging stations under strict energy constraints. To address these limitations, this paper proposes Hierarchical Attention mechanism-based Path Planning for multi-UAV Inspection (HAPPI). The objective is to minimize the total flight distance of the UAV fleet while ensuring that all device nodes are inspected and all UAVs safely return to the base station under energy and visit-count constraints.  Methods  The multi-UAV power inspection problem is first formulated as a combinatorial optimization problem with energy constraints. It is then modeled within a Markov Decision Process (MDP) framework. To solve this problem, HAPPI adopts an encoder-decoder architecture with a customized hierarchical attention mechanism. The encoder uses a multi-level attention design to model three types of node relationships. Self-attention among device nodes is used to learn spatial proximity and visit-order preferences. Cross-attention between device nodes and charging stations is used to model energy supply-demand relationships. Self-attention among charging stations is used to explicitly capture the topological structure of the charging-station network. This hierarchical design enables the model to distinguish functional differences and dependencies among heterogeneous nodes. The decoder integrates the global graph embedding, the embedding of the last visited node, and the current remaining energy of the UAV to generate a context vector. A single-head attention mechanism is then used to compute compatibility scores for all candidate nodes. A masking strategy excludes infeasible nodes, including visited nodes, unreachable nodes, nodes that would prevent the UAV from reaching a charging station, and premature returns to the base station. The final node is selected from a probability distribution generated by softmax, which supports both greedy and sampling decoding strategies. The policy network is trained using reinforcement learning, and a baseline network is used to stabilize training. Policy-gradient optimization is used to minimize the expected total path length (Fig. 2).  Results and Discussions  Extensive simulations are conducted on three problem scales: T20C2, with 20 device nodes and 2 charging stations; T60C6, with 60 device nodes and 6 charging stations; and T100C10, with 100 device nodes and 10 charging stations. The training results show that HAPPI achieves faster convergence and a lower final cost than the baseline Attention Model (AM) and Heterogeneous Attention-based Deep Reinforcement Learning (HADRL) methods (Fig. 4). In the comprehensive performance comparison, HAPPI with sampling obtains the shortest total path lengths on T60C6 and T100C10, with values of 6.21 and 8.41, respectively. It outperforms five classical metaheuristic algorithms and two deep reinforcement learning baselines on these two larger-scale scenarios (Table 1). On T20C2, HADRL with sampling achieves the shortest path length, whereas HAPPI remains highly competitive. Overall, HAPPI reduces the total path length by approximately 12% on average compared with the baseline methods. The visualization results show that HAPPI generates routes with fewer route crossovers and a more balanced workload among UAVs, improving safety and efficiency (Fig. 6 and Fig. 7). The single-UAV path length distribution further confirms the superior load-balancing capability of HAPPI across all problem scales (Fig. 8).  Conclusions  This paper presents HAPPI, a hierarchical attention mechanism-based deep reinforcement learning method for cooperative path planning in multi-UAV power inspection scenarios with multiple charging stations. By explicitly modeling spatial relationships among device nodes, energy dependencies between device nodes and charging stations, and the internal topology of the charging-station network, HAPPI improves information aggregation and constraint satisfaction. Experimental results across different problem scales show that HAPPI achieves higher planning quality, greater computational efficiency, and stronger generalization than heuristic and learning-based comparison methods. Future work will extend this framework to multi-objective optimization that considers time, risk, and energy trade-offs, and will further validate the method using real-world inspection data.
Deep Side-Channel Attack Method Integrating Convolutional Block Attention Mechanism and Triplet Metric Learning
XU Yang, LI Kaibin, HE Xingxing
Available online  , doi: 10.11999/JEIT260140
Abstract:
  Objective  Side-Channel Attack (SCA) is one of the primary threats to the physical security of cryptographic chips, and deep learning methods for secret key recovery have attracted considerable attention in the field of SCA. However, existing deep learning-based side-channel attack methods have limited capability to focus on critical leakage intervals during feature extraction, particularly for long traces with high-dimensional noise. Therefore, irrelevant background noise interferes with feature extraction, leading to reduced attack efficiency and slower convergence of Guessing Entropy (GE). To address these limitations, a deep side-channel attack method integrating the Convolutional Block Attention Module (CBAM) and triplet loss is proposed to improve the extraction of weak leakage features under complex noise conditions and enhance secret key recovery efficiency.  Methods  CBAM is embedded into a Convolutional Neural Network (CNN) to construct an adaptive feature extraction network. CBAM consists of a Channel Attention Module (CAM) and a Spatial Attention Module (SAM). CAM adaptively recalibrates feature-channel weights to emphasize leakage-related features with a high Signal-to-Noise Ratio (SNR), whereas SAM identifies Points Of Interest (POI) in the temporal domain and suppresses background noise outside the leakage intervals. After attention-based feature refinement, triplet loss is adopted as the optimization objective to optimize the distribution of embedding features, encouraging compact intra-class clusters while maximizing inter-class separation in the embedding space. Finally, a multivariate Gaussian template attack is performed using the optimized embedding features to recover the secret key. The overall framework is illustrated in (Fig. 2).  Results and Discussions  The proposed method is evaluated on two public benchmark datasets, ASCAD and AES_HD, using GE and the minimum number of attack traces required for GE to converge to 1 (\begin{document}$ {T}_{{\mathrm{GE0}}} $\end{document}) as evaluation metrics. On the ASCAD dataset, the proposed method requires only 144 attack traces in the ASCAD_f (HW) scenario, representing a 51.0% reduction compared with the conventional CNN model. In the ASCAD_f (ID) scenario, only 61 attack traces are required, corresponding to a 68.0% reduction. In the ASCAD_r dataset, GE converges with 176 attack traces under the HW leakage model and 137 attack traces under the ID leakage model, outperforming representative methods, including RL-SCA and Metric Learning (Table 2 and Fig. 3). On the low-SNR AES_HD dataset, the proposed method requires 1,219 attack traces, outperforming MHA and NLS while maintaining smooth and stable GE convergence (Table 2 and Fig. 4). Furthermore, desynchronization experiments demonstrate that the proposed method maintains accurate localization of effective leakage points under severe desynchronization noise, indicating strong robustness to time-domain jitter (Table 3). Ablation studies further verify the synergistic effect of the proposed architecture and confirm the effectiveness of its core components (Table 4).  Conclusions  A deep side-channel attack method integrating CBAM and triplet-loss-based metric learning is proposed. The CBAM module enables the network to adaptively focus on leakage-related features, improving feature extraction over conventional CNN-based methods. Triplet loss enhances the discriminability of embedding features, thereby improving template matching accuracy. Experimental results on the ASCAD and AES_HD datasets demonstrate that the proposed method substantially reduces the number of attack traces required for successful secret key recovery and accelerates GE convergence. The proposed method consistently outperforms existing mainstream approaches under fixed-key, random-key, and low-SNR conditions. Future work will focus on improving robustness under more severe desynchronization conditions and enhancing generalization in small-sample scenarios.
A WiFi Multi-link Collaborative Human Tracking Method for Smart Home
PAN Houcheng, CAI Yushuang, YAO Junmei, ZHANG Tingting
Available online  , doi: 10.11999/JEIT260267
Abstract:
  Objective  WiFi-based indoor passive human tracking has attracted increasing attention for smart home applications because Internet of Things (IoT) devices are widely interconnected through existing WiFi infrastructures. Existing approaches estimate the Angle of Arrival (AoA) or Time of Flight (ToF) of the target-reflected path to achieve tracking. However, these methods are fundamentally limited by the difficulty of separating weak target-reflected signals from dense multipath propagation when commercial WiFi devices provide only a small antenna array and limited bandwidth under existing communication protocols. Therefore, Doppler Frequency Shift (DFS)-based approaches that exploit multiple links have become more practical. Dead Reckoning (DR) is widely adopted in these methods, but dynamic link selection for signal fusion remains challenging. Furthermore, the performance of DR-based methods degrades substantially when device-location errors are present. Existing methods also generally employ a single-transmitter–multi-receiver architecture, in which Access Point (AP) broadcasts downlink WiFi signals to STAtions (STA). It complicates data aggregation and limits the use of multiple receive antennas of AP's. To address the challenges of multi-link signal fusion and device-location uncertainty, a Particle Filter (PF)-based method is proposed. Furthermore, a multi-transmitter–single-receiver architecture is adopted, in which multiple STAs transmit uplink packets to a single AP. This architecture simplifies data aggregation while exploiting the AP’s multi-antenna capability.  Methods  In the IEEE 802.11 standard, the wireless channel is estimated at the receiver using pilot signals embedded in WiFi packets, and Channel State Information (CSI) is continuously obtained. Human motion perturbs multipath propagation and produces time-varying changes in CSI. Therefore, CSI serves as the primary sensing signal because it implicitly captures target motion. In practice, raw CSI is affected by amplitude and phase impairments. Accordingly, signal preprocessing is first performed before Doppler extraction. Subsequently, multi-link DFS measurements are fused using PF, in which the target state is sequentially propagated and updated according to likelihoods derived from DFS measurements. This probabilistic framework naturally enables dynamic multi-link fusion because unreliable links receive lower weights rather than being deterministically discarded. Moreover, the effects of device-location errors are incorporated into the measurement-noise model, improving robustness to device-location uncertainty. When prior knowledge of device locations is unavailable, device self-localization is first performed by jointly estimating the Line-of-Sight (LoS) AoA and ToF between the AP and each STA. Preliminary device self-localization experiments (Fig. 5) demonstrate a median STA localization error of approximately 0.56 m (Table 1), providing reliable initialization for the subsequent tracking algorithm.  Results and Discussions  A prototype system is implemented using a multi-transmitter–single-receiver architecture with commercial Intel AX200/AX201 WiFi cards. CSI is collected over a 20 MHz channel centered at 5.24 GHz using PicoScenes, with IEEE 802.11ac packets transmitted at 100 or 200 Hz. Each CSI sample contains measurements from 57 subcarriers. Meanwhile, Fine Time Measurement (FTM) is performed on Channel 11 at 2.4 GHz using iw and hostapd. In the prototype system, each STA transmitter is equipped with a single antenna, while the AP receiver employs two antennas. Experiments are conducted using one AP and three or four STAs (Fig. 7). WiTraj and PITrack, both based on DR, are used as baseline methods. When device self-localization is required, the proposed method achieves a median tracking error of 0.47 m, compared with 1.78 m for WiTraj and 1.77 m for PITrack (Fig. 9), representing an accuracy improvement of approximately 70% over conventional DR-based methods. Error-injection experiments with random device-location perturbations show that, unlike conventional DR-based methods, which are highly sensitive to device-location errors, the proposed method remains robust to device-location uncertainty (Fig. 10). Additionally, computational complexity analysis shows that, although execution time increases with the particle count (Table 3), tracking performance reaches a stable level beyond a moderate particle count (Fig. 11). Therefore, real-time operation can be achieved by selecting an appropriate particle count. Finally, experiments in complex environments demonstrate that the proposed method consistently achieves higher tracking accuracy than the baseline methods (Fig. 12).  Conclusions  A PF-based method is proposed for multi-link collaborative passive human tracking using commercial WiFi devices. By probabilistically fusing measurements from multiple links, robust tracking is achieved even when device-location errors are present. When device locations are unavailable, device self-localization is achieved by combining CSI-based estimation with FTM measurements, providing the geometric information required for tracking. A prototype system based on a multi-transmitter–single-receiver architecture is developed to simplify data aggregation while exploiting the AP’s multi-antenna capability. Experimental results demonstrate that the proposed method achieves sub-0.5 m median tracking error when device-location errors are present, representing an improvement of approximately 70% over conventional DR-based methods. The proposed method also exhibits strong robustness to device-location uncertainty and supports real-time implementation when an appropriate particle count is selected. Future work will extend the framework to multi-person tracking and evaluate its performance under more realistic deployment conditions.
Decoupled Learning for Long-tailed Oracle Bone Character Recognition Based on Adaptive Difficulty Sampling
SUN Junwei, GUAN Suyan, CHEN Xinyu, WANG Kun, CAI Yuanqiang
Available online  , doi: 10.11999/JEIT260327
Abstract:
  Objective  Oracle Bone Character (OBC) recognition is challenged by an extreme long-tailed distribution and substantial intra-class variation. Conventional deep learning methods are often dominated by head classes, whereas existing approaches tend to overfit tail classes or fail to account for differences in learning difficulty across classes. To address these limitations, a two-stage decoupled learning framework is proposed to improve the recognition of tail and difficult classes while preserving the discriminative capability of head classes.  Methods  The proposed framework decouples feature representation learning from classifier optimization. In the first stage, the backbone network is trained using a mixed data augmentation strategy that combines CutMix and RandAugment with Label-Distribution-Aware Margin (LDAM) loss to learn robust feature representations and alleviate the effect of intra-class variation. In the second stage, the backbone network is frozen, and only the classifier is optimized. An adaptive difficulty sampling strategy is proposed to dynamically assign sampling weights according to historical and current class-level training difficulty. The classifier is further optimized using a Class-Balanced LDAM (CBL) loss, which combines class-balanced weighting with LDAM to refine decision boundaries for long-tailed classification.  Results and Discussions  Experiments on the highly imbalanced OBC306 dataset demonstrate that the proposed method achieves an overall accuracy of 94.34% and an average class accuracy of 89.89%. Compared with the Inception-v4 baseline, the proposed method improves the average class accuracy by 19.61%. Comparisons with representative long-tailed OBC recognition methods further demonstrate superior overall performance. Comprehensive ablation studies verify the effectiveness of the mixed data augmentation strategy and the adaptive difficulty sampling strategy in improving the recognition of rare and difficult characters. Parameter sensitivity analysis and qualitative error analysis further confirm the robustness and effectiveness of the proposed framework.  Conclusions  The proposed two-stage decoupled learning framework effectively addresses long-tailed OBC recognition by balancing the learning priorities of head, tail, and difficult classes. The mixed data augmentation strategy improves feature robustness, whereas the adaptive difficulty sampling strategy and the Class-Balanced LDAM loss jointly optimize classifier learning and refine decision boundaries without degrading head-class recognition performance. The proposed framework provides an effective solution for the digital recognition of Oracle Bone Characters and offers technical support for low-resource ancient character recognition.
Construction of a DNA Strand Displacement Memristor and Its Filter Circuit Characteristics
WANG Yanfeng, CHEN Guanzhou, SUN Ce, SUN Junwei
Available online  , doi: 10.11999/JEIT260283
Abstract:
  Objective  Filter circuits are widely used in modern control and signal-processing systems for noise suppression and signal integrity enhancement. Conventional Resistor-Capacitor (RC) filters are widely applied, but their fixed parameters limit adaptability and miniaturization in emerging molecular and nanoscale computing platforms. To address these limitations, DNA Strand Displacement (DSD) technology is integrated with memristor theory to develop tunable multistable molecular filter circuits. This study aims to design and validate first- and second-order low-pass filter circuits based on the dynamic response and state-dependent behavior of a DSD-based memristor. The proposed filters are designed to improve frequency selectivity, parameter adaptability, and system stability compared with traditional filter architectures. This approach is intended for molecular signal processing, integrated biocircuits, and adaptive filtering systems that require compact size and reconfigurability.  Methods  The method consists of four stages. First, core DSD reaction modules, including sine, cosine, integration, addition, and multiplication modules, are designed to construct a programmable multistable memristor model. Second, square-wave and sinusoidal input signals are generated through DSD reactions to evaluate the memristor response under different frequencies and amplitudes. Third, the memristor is embedded into low-pass filter structures to construct first- and second-order DSD-based memristor filter circuits. Fourth, simulations are performed using Visual DSD for molecular dynamics analysis and MATLAB for circuit-level analysis. Circuit performance is evaluated using transfer functions, Nyquist plots, Bode diagrams, and time-domain comparisons with classical RC filters. This combined simulation strategy verifies both molecular feasibility and circuit functionality.  Results and Discussions  The DSD-based memristor exhibits multistable behavior and converges to six stable equilibrium points under different initial conditions (Fig. 8). Its hysteresis characteristics further confirm the state-dependent memory behavior of the designed molecular memristor (Fig. 7). The first-order DSD-based memristor filter circuit provides stable attenuation for square-wave and sinusoidal input signals. Its output amplitudes are consistently higher than those of the traditional RC filter across the tested frequencies (Table 3). The second-order DSD-based memristor filter circuit further reduces signal delay and improves stability, especially under high-frequency inputs (Table 4). Frequency-response analyses show that the cutoff frequency can be dynamically tuned by adjusting DSD reaction rates and initial concentrations (Figs. 9 and 11). Time-domain simulations further confirm the filtering performance of the first- and second-order circuits (Figs. 10 and 12). Reliability analysis indicates that lower initial copy numbers increase stochastic molecular noise, whereas higher initial copy numbers make the output distribution closer to the deterministic response and improve the probability of successful filtering. These results verify the feasibility of DSD-memristor integration for adaptive molecular filtering.  Conclusions  A DSD-based memristor with multistable characteristics and its corresponding first- and second-order low-pass filter circuits are designed and validated. Compared with traditional RC architectures, the proposed filters show improved output stability, parameter tunability, and frequency adaptability. By combining DSD technology with memristor theory, this study provides a reconfigurable molecular-scale filtering framework for signal-processing applications. The results provide a basis for future work on adaptive molecular circuits, intelligent filtering, and nanoelectronic system design. Further studies should focus on experimental validation, real-time tuning strategies, sequence optimization, anti-interference design, signal amplification, and circuit integration.
Non-Orthogonal PSWFs Signal Detection Method Based on Adaptive Time-Space Feature Fusion
CHEN Wenhua, MAO Zhongyang, LU Faping, SUN Ye, GAO Yixuan
Available online  , doi: 10.11999/JEIT260024
Abstract:
  Objective   To address B5G/6G demands for high spectral efficiency and transmission reliability, non-orthogonal Prolate Spheroidal Wave Function (PSWF) modulation has drawn extensive research interest for its strong time-frequency energy concentration. However, severe mutual interference in PSWF signal multiplexing degrades conventional detection performance in complex channels. Limited by ideal channel assumptions or single-modal feature extraction, existing methods cannot fully exploit signal temporal-spatial information and lack adaptability to dynamic interference. This work proposes an intelligent detection architecture with adaptive temporal-spatial feature fusion (ATSFF) for accurate, robust non-orthogonal PSWF signal detection.  Method   A dual-path parallel framework extracts complementary features from time and transform domains. A Gated Recurrent Unit (GRU) extracts deep temporal features and captures long-range dependencies from 1D received signals. The other path converts 1D signals into 2D representations via the Gramian Angular Difference Field (GADF), and extracts hierarchical spatial features using ResNet50. An adaptive probability-weighted fusion mechanism dynamically adjusts feature contributions based on sub-network prediction uncertainty to generate robust detection outputs.  Results and Discussion   Simulations on a 32-class non-orthogonal PSWF dataset (Fig. 2) show the proposed ATSFF outperforms conventional coherent detection, cross-term suppression detection, Approximate Message Passing Interleave Division Multiple Access (AMP-IDMA) and Temporal Sparse Bayesian Learning Least Squares (TMSBL-LS) across the full Signal-to-Noise Ratio (SNR) range. t-SNE visualization (Fig. 4) confirms better inter-class separability and intra-class compactness of fused features. At a bit error rate of 4×10–5, it achieves 0.2 dB gain (Fig. 6) over coherent detection. ATSFF has higher network overhead and weaker real-time performance than linear detection, yet it supports GPU batch inference with fixed single-sample computation cost, suiting accuracy-sensitive communication scenarios.  Conclusions   Targeting interference issues in non-orthogonal PSWF signal detection, this paper presents an adaptive temporal-spatial feature fusion detection method. It realizes dual-modal feature extraction via GRU and ResNet50, and adopts a prediction-uncertainty-based adaptive weighted fusion mechanism to integrate dual-path strengths, significantly improving detection accuracy and robustness. Simulations validate its superior performance and stable channel adaptability, providing an efficient solution and design reference for intelligent detection of non-orthogonal PSWF signals.
A Multi-Dimensional Scenario-Based Evaluation Method for Deep Learning Side-Channel Analysis Using a Multi-Attribute Decision Model
GU Zepeng, CHEN Lin, CAI Juesong, YAN Yingjian
Available online  , doi: 10.11999/JEIT260198
Abstract:
  Objective  Deep Learning Side-Channel Analysis (DL-SCA) has substantially improved the effectiveness of attacks against protected cryptographic implementations. However, the transition of DL-SCA models from research to practical deployment is limited by the lack of systematic, fair, and scenario-specific evaluation methods. Existing evaluations mainly rely on Guessing Entropy (GE) and Success Rate (SR), while overlooking practical factors such as resource overhead and environmental adaptability. Moreover, inconsistent hyperparameter optimization leads to unfair model comparisons and provides limited quantitative guidance for model selection under different deployment constraints, including resource-constrained devices, high-noise environments, and real-time applications. This paper proposes a systems engineering-based evaluation framework that enables comprehensive, quantitative, and scenario-specific assessment of DL-SCA models.  Methods  A multi-dimensional, scenario-based evaluation framework is developed using systems engineering principles. First, a hierarchical evaluation index system is established, comprising three criteria—attack effectiveness, resource overhead, and environmental adaptability—and six evaluation metrics: GE, SR, training time (TC), peak memory consumption (MC), model complexity (MoC), and noise robustness (Rob). Second, a standardized evaluation process based on the V-model is designed to ensure fair comparison. Each candidate model, including a Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), and CNN-LSTM hybrid model, undergoes independent hyperparameter optimization using grid search before multi-dimensional performance evaluation. Third, a hybrid Criteria Importance Through Intercriteria Correlation-Analytic Hierarchy Process (CRITIC-AHP) Multi-Attribute Decision-Making (MADM) framework is developed. The CRITIC method derives objective weights from the statistical characteristics of the evaluation data, whereas the AHP method incorporates scenario-specific preferences through pairwise comparison matrices. The objective and subjective weights are fused to generate scenario-specific weights. Finally, a Multi-dimensional Attack Performance Metric (MAPM) is defined as the weighted sum of normalized evaluation metrics using the fused weights, providing a composite score for each model under a specific deployment scenario.  Results and Discussions  The proposed framework is validated using the ASCAD fixed-key dataset. After independent hyperparameter optimization, the three model architectures are evaluated using all six metrics. The CRITIC method produces the objective weight vector W critic = [0.17, 0.19, 0.15, 0.21, 0.14, 0.14]. Four representative deployment scenarios—Resource-Constrained, High-Performance, High-Noise, and Real-Time—are then defined, and the corresponding AHP preference weights are fused with the objective weights to generate the final scenario-specific weights. For example, MC receives the highest weight (0.52) in the Resource-Constrained scenario, whereas Rob dominates the High-Noise scenario with a weight of 0.57. The resulting MAPM scores (Table 9, Fig. 9, and Fig. 10) clearly differentiate the strengths of the evaluated models and demonstrate the scenario-specific decision capability of the proposed framework. CNN achieves the highest score in the High-Performance scenario (0.894), MLP ranks first in the Real-Time scenario (0.758) because of its shortest training time, and the CNN-LSTM hybrid model performs best in the High-Noise scenario (0.863) because of its superior noise robustness despite higher resource overhead. These results demonstrate that no single model is optimal across all deployment scenarios and that MAPM provides a clear and quantitative basis for model selection under specific deployment constraints.  Conclusions  This paper proposes a systems engineering-based, multi-dimensional evaluation framework to address the major limitations of current DL-SCA model assessment. By integrating a hierarchical evaluation index system, a standardized V-model evaluation process, and a hybrid CRITIC-AHP Multi-Attribute Decision-Making (MADM) framework, the proposed method quantitatively balances the trade-offs among attack effectiveness, resource overhead, and environmental adaptability. Experimental results obtained using the ASCAD benchmark demonstrate that the framework provides clear, quantitative, and scenario-specific guidance for model selection. The proposed Multi-dimensional Attack Performance Metric (MAPM) provides a practical decision basis for selecting DL-SCA models under diverse deployment constraints, narrowing the gap between academic attack development and practical model deployment. Future work will extend the framework to additional model architectures and datasets, improve evaluation automation, and validate its effectiveness in practical deployment environments.
Spatial-domain Anti-jamming for Unmanned Systems Under Limited Prior Information
PAN Zihao, ZHANG Bangning, ZHEN Pan, ZHU Bowen, WANG Ning, GUO Daoxing
Available online  , doi: 10.11999/JEIT260296
Abstract:
  Objective  Unmanned systems play an increasingly important role in emergency response, public safety, intelligent transportation, and other mission-critical applications. Reliable communications in complex electromagnetic environments are essential for autonomous operation. However, communication links are directly exposed to open, non-cooperative electromagnetic environments and are therefore vulnerable to intentional jamming and unintentional interference. In practical scenarios, prior information regarding the desired signal, jamming sources, and multipath propagation is often unavailable, substantially degrading the performance of conventional spatial-domain anti-jamming methods. To address this challenge, this paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information.  Methods  The proposed method first applies a spatial smoothing algorithm to the received signals to decorrelate coherent multipath components. Capon spatial spectrum estimation is then performed to detect the Direction Of Arrival (DOA) of potential incident signals. Spectrum peaks corresponding to individual incident signals are subsequently identified. A Covariance Matrix Reconstruction (CMR)-based beamforming algorithm is then applied by traversing all detected spectrum peaks to sequentially extract the signal associated with each peak, thereby separating the mixed signals. After signal separation, a signal classification method based on spectral similarity and time delay is employed. Kullback-Leibler (KL) divergence between the spectrum of each separated signal and the reference spectrum is calculated to identify jamming signals. The remaining communication signals are further classified into direct-path and multipath signals according to their relative time delays. Finally, different processing strategies are applied according to the identified signal type. Specifically, multipath signals are either suppressed as interference or coherently combined with the direct-path signal after time-delay and phase alignment.  Results and Discussions  Two simulation scenarios, including jamming only and combined jamming and multipath, are designed to evaluate the proposed method in terms of the output Signal-to-Interference-plus-Noise Ratio (SINR), beam pattern, Bit Error Rate (BER), and Error Vector Magnitude (EVM). Simulation results demonstrate that, under the jamming-only scenario, the proposed method achieves performance close to the theoretical optimum. The output SINR increases with the input Signal-to-Noise Ratio (SNR) at a fixed Jamming-to-Signal Ratio (JSR) (Fig. 3(a)) and remains nearly unchanged as JSR increases at a fixed SNR (Fig. 3(b)), indicating stable jamming suppression capability. The recovered time-domain waveform and spectrum remain highly consistent with the transmitted signal (Fig. 4). The BER curve nearly overlaps that of the optimal beamformer (Fig. 5). At \begin{document}$ {E}_{\rm b}/{N}_{0}=10\;{\mathrm{dB}} $\end{document}, the recovered Quadrature Phase-Shift Keying (QPSK) constellation closely matches the ideal constellation, achieving an EVM of –11.52 dB (Fig. 6). Under simultaneous jamming and multipath conditions, the proposed framework flexibly suppresses or exploits multipath signals. Compared with multipath suppression, multipath utilization further improves both the output SINR and BER (Fig. 7(a) and Fig. 7(b)). The corresponding beam pattern forms a beam toward the multipath direction rather than a null, demonstrating effective multipath exploitation (Fig. 7(c)).  Conclusions  This paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information. Using only the received mixed signals, the proposed framework estimates the directions of arrival, separates incident signals, and classifies them as direct-path, multipath, or jamming signals. Appropriate suppression or preservation strategies are then applied according to the identified signal type. Therefore, the framework flexibly suppresses or exploits multipath signals while preserving the direct-path signal and mitigating jamming. Simulation results demonstrate the effectiveness of the proposed method in terms of output SINR and demodulation accuracy, confirming reliable jamming suppression and communication performance even when prior information regarding the desired signal, jamming sources, and multipath propagation is unavailable. Future work will investigate the effects of array perturbations, intelligent jamming, and heterogeneous communication modes on the proposed framework and extend it to more complex unmanned-system communication environments.
THz Ultra-Massive MIMO Channel Estimation via a Noise-Conditioned Fixed-Point Network
ZHANG Huawei, NIU Yaning, JIANG Zhanjun, LIU Yingting
Available online  , doi: 10.11999/JEIT260420
Abstract:
  Objective  Terahertz (THz) Ultra-Massive Multiple-Input Multiple-Output (UM-MIMO) systems are expected to support future high-capacity wireless communications. However, accurate channel estimation remains challenging under hybrid near-/far-field propagation and Array-of-SubArrays (AoSA) architectures, where limited Radio-Frequency (RF) chains, low Signal-to-Noise Ratio (SNR), noise uncertainty, and structural perturbations degrade compressed observations. Existing compressed sensing, Bayesian inference, and deep unfolding methods generally rely on fixed statistical assumptions, which limit their cross-SNR generalization under varying noise conditions and statistical mismatches. To address these limitations, this paper proposes a Noise-Conditioned Fixed-Point Network (FPN-NCAS) for robust THz UM-MIMO channel estimation. The proposed method aims to improve estimation accuracy, cross-SNR generalization, robustness, and iterative stability by incorporating noise-aware nonlinear recovery.  Methods  FPN-NCAS is developed within an Orthogonal Approximate Message Passing (OAMP)-based fixed-point unfolding framework. A coarse noise power estimate is obtained from repeated pilot differences and injected into the nonlinear recovery module as an explicit conditioning variable. After each linear update, the vector-domain estimate is reshaped into an AoSA-aligned tensor to exploit subarray-level structural priors. The nonlinear recovery chain consists of three components. Token-Gate performs lightweight subarray-level reliability pre-calibration to suppress unreliable responses under low-SNR and structurally inconsistent conditions. Block Shrink serves as the core noise-conditioned block-sparse proximal operator, in which a smooth dual-threshold mechanism balances strong denoising at low SNR with structural preservation at medium-to-high SNR. G-HMTD(Gated Hybrid Multi-scale Transformer Denoiser) further refines residual errors by combining Local Multi-scale Enhancement and Global Context Modeling. Bridge relaxation and nonlinear residual scaling are also incorporated to improve inter-stage stability.  Results and Discussions  Simulation results demonstrate that FPN-NCAS consistently outperforms LS, OAMP, ISTA-Net+, FPN-OAMP, and FPN-OTFN over the 0~20 dB SNR range (Fig. 6). At SNR = 0 dB, FPN-NCAS achieves Normalized Mean Square Error (NMSE) gains of approximately 3.0 dB and 1.9 dB over FPN-OAMP and FPN-OTFN, respectively. At SNR = 20 dB, the gains increase to approximately 4.5 dB and 2.6 dB (Fig. 6(a)). The convergence curves show that FPN-NCAS reaches a stable plateau after approximately three layers at SNR = 5 dB and five layers at SNR = 15 dB, demonstrating stable fixed-point iterative behavior (Fig. 6(b) and Fig. 6(c)). Analysis of repeated pilots shows that four repeated pilot pairs introduce only 3.13% additional pilot overhead while reducing the relative standard deviation of the coarse noise power estimate to 25.00%. Under moderate noise power mismatch, NMSE degradation remains within 0.3 dB. The proposed method also maintains strong robustness under colored Gaussian noise, impulsive noise, near-/far-field distribution shifts, variation in the number of propagation paths, AoSA subarray shuffling, and RF amplitude/phase mismatch (Fig. 7 and Fig. 8). Ablation studies indicate that Block Shrink provides the largest performance gain, whereas Token-Gate and G-HMTD further improve performance through structural calibration and residual refinement (Fig. 9).  Conclusions  This paper proposes FPN-NCAS for noise-conditioned fixed-point channel estimation in THz UM-MIMO systems. By integrating repeated-pilot-based noise conditioning, AoSA-aware feature reshaping, Token-Gate calibration, Block Shrink recovery, and G-HMTD refinement, the proposed method improves NMSE performance, robustness, and iterative stability under different SNR conditions and non-ideal scenarios. The improved performance is achieved at the cost of higher inference complexity. Future work will focus on lightweight implementations and extensions to wideband, multi-user, and hardware-impaired THz communication systems.
A Nested Multi-scroll Memristive Hopfield Neural Network and Its Hardware Implementation
WANG Zhe, WAN Qiuzhen, ZHOU Pan, RAO Huhui
Available online  , doi: 10.11999/JEIT260516
Abstract:
  Objective  In recent years, memristors have been employed to emulate neuronal synapses with dynamically adjustable synaptic weights, enabling the construction of Memristive Hopfield Neural Networks (HNNs). Compared with conventional HNNs, Memristive HNNs more accurately reproduce the nonlinear dynamical behavior of biological neural systems. Multi-scroll attractors have attracted considerable attention in secure communication because of their complex topological structures and strong state-space ergodicity. However, previous studies have primarily focused on conventional multi-scroll attractors with single structural patterns, whereas multi-scroll attractors with special structures remain largely unexplored. Therefore, this paper proposes a nested multi-scroll Memristive HNN system that generates nested multi-scroll attractors, thereby overcoming the limitations of conventional single-structure multi-scroll attractors.  Methods  A Four-Dimensional (4D) Memristive HNN system is constructed from a three-neuron HNN by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of neuron 2. Equilibrium-point and stability analyses are performed to investigate the regulatory effects of the Memristive self-connected synapse coupling strength and system initial conditions on the system dynamics. The number of multi-scroll attractors is regulated by adjusting the memristor control parameters. Building on this framework, a Multi-level Logic Pulse current (IMLP) is introduced to construct a nested multi-scroll Memristive HNN system. The proposed system generates nested multi-scroll attractors with enhanced dynamical complexity. Finally, the MATLAB numerical simulation results are validated through Multisim circuit simulations and Field-Programmable Gate Array (FPGA)-based hardware experiments.  Results and Discussions  The results demonstrate that regulating the Memristive self-connected synapse coupling strength enables the proposed 4D Memristive HNN system to exhibit period-doubling bifurcations and chaotic behavior, as illustrated by the bifurcation diagrams and Lyapunov exponent spectra (Fig. 3). Various types of coexisting attractors are generated under different coupling strengths (Fig. 4). By adjusting the memristor control parameters, multi-scroll attractors with different numbers of scrolls are generated through one-directional extension (Figs. 58). After the introduction of the IMLP, the proposed nested multi-scroll Memristive HNN system generates nested multi-scroll attractors while preserving the controllable scroll-number extension property (Figs. 810). Spectral Entropy (SE) analysis demonstrates that the IMLP increases the dynamical complexity of the proposed system compared with the original 4D Memristive HNN system (Figs. 9 and 10). The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments confirms the physical realizability of the proposed nested multi-scroll Memristive HNN system (Figs. 1214).  Conclusions  A 4D Memristive HNN system is constructed by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of a three-neuron HNN. Equilibrium-point and stability analyses reveal the regulatory effects of the Memristive self-connected synapse coupling strength and the evolution of coexisting attractors associated with different initial conditions. The results show that the system enters chaos through the period-doubling route to chaos and generates single-scroll and double-scroll chaotic attractors. The number of multi-scroll attractors is continuously increased by adjusting the memristor control parameters. Furthermore, introducing the IMLP produces a nested multi-scroll Memristive HNN system capable of generating nested multi-scroll attractors with increased dynamical complexity. The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments validates the physical realizability of the proposed nested multi-scroll Memristive HNN system.
Research on Ka-band Enhanced Active Load Modulation Ultra-wideband High-efficiency Doherty Power Amplifier
YANG Lin, YAN Chengyu, WANG Yanping, ZHANG Ming, WANG Baozhu, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260514
Abstract:
  Objective  The Ka-band has become a key frequency band for satellite communications, placing stringent requirements on the millimeter-wave power amplifier, a core component of the transmitter, to provide high efficiency, compact size, and broadband operation. To maximize spectral efficiency, millimeter-wave satellite communication signals typically exhibit a high Peak-to-Average Power Ratio (PAPR), making high back-off efficiency particularly important. Although the Doherty Power Amplifier (DPA) is widely adopted because of its high efficiency under power back-off conditions, its operating bandwidth is inherently limited. In addition, both saturated efficiency and back-off efficiency degrade substantially at millimeter-wave frequencies. Therefore, extending the operating bandwidth while maintaining high efficiency remains a major challenge for millimeter-wave DPAs used in satellite communication transmitters.  Methods  An ultra-wideband enhanced active load modulation technique is proposed to overcome the trade-off between bandwidth and back-off efficiency in DPAs. The proposed method achieves optimal load impedance modulation for both the carrier and peaking amplifiers over an ultra-wide frequency range by introducing an Impedance Tunable Bias Network (ITBN) and a dual-drive impedance control mechanism. These techniques improve load modulation while extending the load modulation bandwidth, thereby enhancing both efficiency and bandwidth in the millimeter-wave DPA architecture. Furthermore, a broadband phase-compensation technique is integrated into an unequal power division network to achieve sufficient load modulation across the entire operating band while accurately compensating for the phase difference between the carrier and peaking paths. The proposed ultra-wideband phase-compensated power division network further extends the high-efficiency operating bandwidth while reducing the chip area.  Results and Discussions  To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was designed and fabricated using a 0.15-μm GaN process. Across the 24~33 GHz frequency band, corresponding to a relative bandwidth of 31.6%, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB with a gain flatness of ±1 dB. The measured saturated output power is 28.9~31.0 dBm, with a Power-Added Efficiency (PAE) of 25.2%~33.5% at saturation and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating over similar frequency bands, the proposed design achieves the highest reported small-signal gain and saturated PAE while maintaining high saturated output power and high 6 dB back-off PAE. Furthermore, it occupies the smallest chip area among reported three-stage DPA MMICs.  Conclusions  An ultra-wideband enhanced active load modulation method is proposed to achieve sufficient active load modulation through multi-frequency impedance tuning across the entire operating band. A novel ultra-wideband phase-compensated unequal power division network is also proposed to reduce the chip area while maintaining accurate phase compensation. To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was fabricated using a 0.15-μm GaN process. Measurement results demonstrate that, over the 24~33 GHz frequency band, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB, a saturated output power of 28.9~31.0 dBm, a saturated PAE of 25.2%~33.5%, and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating in similar frequency bands, the proposed design achieves a maximum relative bandwidth of 31.6%, the highest reported small-signal gain and saturated PAE, high saturated output power, and high 6 dB back-off PAE, while occupying the smallest chip area among reported two- and three-stage DPA MMICs. These measurement results validate the proposed method and demonstrate its strong potential for millimeter-wave satellite communication transmitters.
Iterative Parameter Estimation Method for Energy Detection Threshold in Ambient Backscatter
QU Wenfeng, YE Yinghui, SHI Liqin, LU Guangyue
Available online  , doi: 10.11999/JEIT260418
Abstract:
  Objective  In Ambient Backscatter Communication (AmBC) systems, a low-complexity Energy Detector (ED) is commonly employed at the reader to recover symbols transmitted by the tag. The detection performance of ED depends strongly on the accurate setting of the detection threshold, which is determined by the average received signal power corresponding to tag symbols “1” and “0”. Existing parameter estimation methods assume that the two symbols are transmitted with equal probability and therefore divide the sorted received signal power samples into two equal groups. However, because the number of transmitted symbols is finite, the actual numbers of symbols “1” and “0” are generally unequal. Therefore, equal partitioning introduces sample misclassification, causing the estimated threshold to deviate from its optimal value and reducing detection performance. To address this limitation, an iterative threshold parameter estimation method is proposed to reduce the parameter estimation bias caused by sample misclassification and improve the accuracy of detection threshold estimation.  Methods  An iterative threshold parameter estimation method is proposed to overcome the sample misclassification introduced by conventional sorting-based grouping. Because the initial detection threshold obtained by the sorting-based grouping method provides reliable decisions for most received samples, these initial decisions are used as the basis for sample reclassification. The received signal samples are then reclassified to iteratively update the threshold parameters, progressively refining the detection threshold. The proposed method is evaluated through simulations under three representative ambient radio-frequency source conditions: complex Gaussian, Phase-Shift Keying (PSK), and Quadrature Amplitude Modulation (QAM) sources.  Results and Discussions  Simulation results show that, over a wide range of Signal-to-Noise Ratio (SNR) values, the proposed iterative method substantially reduces the Bit Error Rate (BER) compared with the conventional sorting-based grouping method and approaches the theoretical lower bound obtained with perfect parameter estimation. At a given SNR, the proposed method improves BER by approximately 0.1, 1.5, and 1.3 orders of magnitude under complex Gaussian, PSK, and QAM sources, respectively (Fig. 2). These results demonstrate that the proposed iterative method effectively corrects sample misclassification and reduces the performance loss caused by parameter estimation bias. Moreover, most of the performance gain is achieved after only one iteration, indicating rapid convergence with minimal additional computational overhead. Under different numbers of sampling points, BER improvements of approximately 0.5, 1.6, and 1.1 orders of magnitude are achieved under complex Gaussian, PSK, and QAM sources, respectively (Fig. 3). These results indicate that using the initial decisions for sample reclassification effectively reduces the estimation bias introduced by fixed equal partitioning, thereby improving detection performance under limited-sample conditions. Under different Relative Channel Difference (RCD) values, BER improvements of approximately 0.56 and 0.7 orders of magnitude are achieved under complex Gaussian and QAM sources, respectively (Fig. 4). As the RCD increases, the separation between the received signal power distributions becomes more pronounced, improving the accuracy of the initial decisions and enabling more reliable sample reclassification. This positive feedback process further refines the parameter estimates and improves detection performance.  Conclusions  An iterative threshold parameter estimation method is proposed to address the sample misclassification introduced by conventional sorting-based grouping in ED. The proposed method uses the initial decisions to reclassify the received signal samples and iteratively update the threshold parameters. In addition, closed-form expressions for the detection threshold and BER under QAM sources are derived. Simulation results demonstrate that the proposed method effectively reduces parameter estimation bias with only one iteration while maintaining robust performance under limited-sample and varying channel conditions. Significant BER improvements are achieved with minimal additional computational overhead, making the proposed method well suited for practical, high-reliability AmBC systems.
LLMA-GCN: A Semantically Enhanced Hierarchical Spatiotemporal Graph Convolutional Network for Skeleton-based Action Recognition
JIA Guimin, ZHOU Xilong
Available online  , doi: 10.11999/JEIT260154
Abstract:
  Objective  Current methods that introduce the Large Language Model (LLM) into skeleton-based action recognition face three main limitations. Semantic guidance remains decoupled from spatial topology learning, temporal modeling lacks hierarchical semantic support, and traditional classification paradigms have limited generalization. To address these issues, this paper proposes LLMA-GCN, a semantically enhanced graph convolutional framework. The proposed framework integrates LLM-derived semantic prior knowledge with graph convolution to improve spatiotemporal feature learning and action classification.  Methods  LLMA-GCN uses visual skeleton data and semantic inputs in a dual-branch framework. Frozen LLMs and prompt engineering are used to precompute the joint semantic adjacency matrix and action text prototypes. The framework consists of three main components: an LLM-based hybrid graph topology learning strategy, a hierarchical visual sequence encoder based on the LLM Refinement Block (LRB), and an action text-prototype-guided decision learning mechanism. These components enable semantic guidance in graph topology learning, hierarchical spatiotemporal feature extraction, and visual-text alignment.  Results and Discussions  Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD I show that LLMA-GCN achieves competitive or superior performance compared with state-of-the-art methods. Ablation studies confirm the key roles of the hybrid graph topology, the LRB, and the action text-prototype-guided decision learning mechanism. Model-complexity analysis further indicates the potential of the proposed framework for practical application.  Conclusions  By fusing the joint semantic adjacency matrix and the physical adjacency matrix, LLMA-GCN enables action semantics to directly guide graph convolution and improves semantic perception. The LRB embeds semantic information into hierarchical spatiotemporal feature extraction, which strengthens the modeling of complex actions. The action text-prototype-guided decision learning mechanism further shifts skeleton-based action recognition from purely visual classification to text-prototype-guided visual-text alignment. Overall, LLMA-GCN provides a robust and generalizable framework for skeleton-based action recognition through deep fusion of visual and semantic features.
Low-Complexity Phase Ambiguity Resolution DOA EstimationAlgorithm for Composite Hierarchical Receiving Array Structure
CHEN Yiwen, DONG Yangze, CHEN Xiahua, LING Wenchang, XIONG Yiwen
Available online  , doi: 10.11999/JEIT260447
Abstract:
  Objective  Direction Of Arrival (DOA) estimation is a key technique for sonar target localization. As the demand for high-precision DOA estimation in complex environments continues to increase, the number of array elements used for estimation is steadily growing, leading to massive arrays. Although larger arrays improve DOA estimation accuracy and resolution, they also impose a substantial computational burden on conventional DOA estimation algorithms. To address this issue, a low-complexity composite hierarchical receiving array structure is constructed, and two fast phase ambiguity resolution algorithms are proposed: Composite HierArchical Global Nearest-Neighbor Matching (CHA-GNNM) and Composite HierArchical Cross-Correlation Covariance Merging (CHA-CCM).  Methods  The CHA-GNNM algorithm constructs multiple candidate solution sets by exploiting the auto-covariance and cross-covariance relationships among the subarrays within each group. The true solution in each candidate solution set is identified through nearest-neighbor matching based on source consistency, and the final DOA estimate is obtained through multilevel coherent combining. This approach achieves phase ambiguity resolution and angle matching with relatively low computational cost. However, because the correlation information among all array elements is not fully exploited, some estimation performance is sacrificed. To improve DOA estimation performance, the CHA-CCM algorithm reorganizes the composite hierarchical structure into evenly partitioned groups, which are regarded as several large subarrays. Multiple large candidate solution sets are first constructed from the cross-correlation relationships among these groups. Each group is then divided into multiple small subarrays, from which additional candidate solution sets are generated using the corresponding auto-covariance and cross-covariance relationships. A coarse DOA estimate is obtained through coprime clustering, followed by a more accurate initial DOA estimate derived from the small candidate solution sets. This initial DOA estimate is subsequently used to eliminate spurious solutions from the large candidate solution sets, yielding the final DOA estimate. Combined with a low-complexity covariance block-processing strategy, this approach avoids computationally expensive operations while improving DOA estimation accuracy.  Results and Discussions  Simulation results demonstrate that both proposed algorithms substantially reduce the computational burden as the number of array elements increases, while effectively achieving phase ambiguity resolution through the proposed composite hierarchical receiving array structure (Fig. 5). Compared with the conventional Root-MUSIC algorithm, CHA-GNNM achieves coarse DOA estimation with nearly four orders of magnitude lower computational complexity (Fig. 7), making it suitable for applications with stringent real-time requirements. In contrast, CHA-CCM requires only a modest increase in computational cost (Fig. 7) while achieving DOA estimation performance close to the Cramér-Rao Lower Bound (CRLB) above a certain signal-to-noise ratio threshold (Fig. 6). Therefore, a favorable balance is achieved between DOA estimation accuracy and computational complexity.  Conclusions  To address the rapid increase in computational complexity associated with massive arrays, a composite hierarchical receiving array structure is constructed for efficient DOA estimation. By hierarchically grouping the array elements, the proposed structure provides a new framework for low-complexity DOA estimation. Based on this structure, two fast DOA estimation algorithms are developed. Both algorithms achieve effective phase ambiguity resolution with low computational complexity by exploiting the structural differences among array groups and the consistency of observations from the same source across different groups, thereby enabling rapid DOA estimation. CHA-GNNM primarily exploits the phase relationships among subarrays to perform phase ambiguity resolution and angle matching through a simple computational procedure, making it suitable for applications requiring high computational efficiency and real-time processing. Because the cross-correlation information among all array elements is not fully exploited, some estimation performance is reduced under challenging signal conditions. To overcome this limitation, CHA-CCM reorganizes the composite hierarchical receiving array into evenly partitioned groups while preserving the low-complexity advantage of the hierarchical structure. Group-level cross-correlation information is further exploited so that the intrinsic relationships among different groups are more fully utilized. In addition, the signal processing procedure is simplified by eliminating unnecessary computational steps, thereby improving the robustness and accuracy of DOA estimation while maintaining manageable computational complexity. Compared with CHA-GNNM, CHA-CCM incurs only a small increase in computational cost and achieves a better balance between computational complexity and DOA estimation performance. Overall, the proposed composite hierarchical receiving array structure and the two fast DOA estimation algorithms provide an effective solution for efficient DOA estimation in massive arrays. CHA-GNNM is more suitable for applications with stringent real-time requirements, whereas CHA-CCM is better suited for applications requiring higher DOA estimation accuracy and robustness. The proposed structure achieves efficient phase ambiguity resolution and accurate DOA estimation and provides both theoretical significance and practical value for engineering applications of massive array signal processing.
A Hierarchical Cross-layer Closed-loop Learning Framework andCoordination Mechanism for Complex Multi-agent Systems
ZHANG Long, HUANG wenbo, LEI Zhen, FENG Xuanming, WANG Ying
Available online  , doi: 10.11999/JEIT260143
Abstract:
Complex Multi-Agent Systems (MAS) in dynamic and uncertain environments face challenges in unified modeling, adaptive coordination, and interpretable effectiveness evaluation. Existing methods usually address individual decision-making, inter-agent coordination, and high-level policy evolution separately. This separation leads to fragmented decision chains and weak cross-layer coupling. It also makes it difficult to explain how local learning gains are transformed into global effectiveness improvements under mission variation, observation disturbance, and structural damage. To address this issue, a Hierarchical Cross-layer Closed-loop Learning (HCCL) framework is proposed. The framework couples individual autonomy, system-level coordination, and system-of-systems learning to build a computable path from local policy optimization to overall effectiveness enhancement.   Methods   HCCL adopts a unified three-layer architecture. At the individual autonomy layer, each agent is modeled as a Partially Observable Markov Decision Process (POMDP) to describe decision-making under partial observability. At the system-level coordination layer, multi-agent coordination is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and represented by a dynamic directed weighted coordination graph. A Graph Neural Network (GNN) is used to encode interaction dependencies, structural coupling, and joint value information. At the system-of-systems learning layer, a Meta-Decentralized Partially Observable Markov Decision Process (Meta-Dec-POMDP) is established to describe task-context adaptation and rule evolution. A cross-layer closed-loop mechanism is further designed. In the bottom-up behavior induction pathway, local state and capability features are aggregated into graph-level structural representations and supplied to the upper rule-learning process. In the top-down rule-shaping pathway, learned high-level rules are converted into control parameters and fed back to lower layers to regulate local policies and coordination relationships. Simulations are conducted under baseline, mission-variation, observation-disturbance, and structural-damage scenarios. The full HCCL model is compared with a non-closed-loop model and an upward-induction-only model. Interface ablation studies are also performed to analyze the contributions of cross-layer feature reporting, structural induction, and rule shaping.   Results and Discussions   The full HCCL model consistently outperforms the comparison models and ablated variants. In the baseline scenario, it achieves a task success rate of 88.6% and a comprehensive system effectiveness of 0.842. Under mission variation, it reduces the adaptation process to 16±2 rounds. Under structural damage, it achieves a recovery rate of 81.4% and restores coordination-structure stability to 0.742 within 20 steps. These results indicate that HCCL improves task performance, adaptation speed, and structural recovery. Ablation results show that removing any cross-layer interface reduces performance, while removing the top-down rule-shaping pathway causes the largest loss. This result indicates that upward structural perception alone is insufficient for sustained system-level improvement. The effectiveness gain mainly arises from closed-loop coupling between bottom-up behavior induction and top-down rule shaping, rather than from simple hierarchical stacking.   Conclusions   The HCCL framework is proposed for complex MAS by integrating POMDP-based individual autonomy modeling, Dec-POMDP- and graph-based coordination modeling, and Meta-Dec-POMDP-based rule evolution. Through bottom-up behavior induction and top-down rule shaping, HCCL provides a computable and interpretable path from local learning to overall effectiveness enhancement. Experimental results verify its advantages in task completion, adaptation, recovery, and coordination stability under multiple disturbances. Future work will focus on larger-scale heterogeneous systems, communication-constrained networking, online continual adaptation, and data-driven evaluation in realistic environments.
Off-grid Blind Near-Field Integrated Sensing And Communication: Algorithm Design and Lower Bound
YUAN Zhengdao, GUO Qinghua, HUANG Chongwen, GAO Dawei, MEI Fengtong, LIAO Guisheng
Available online  , doi: 10.11999/JEIT260404
Abstract:
  Objective  With the widespread deployment of extra-large-scale antenna arrays in 6G networks, user terminals are increasingly located in the near-field region. Existing Near-Field Integrated Sensing And Communication (NF-ISAC) algorithms face key challenges, including off-grid power leakage, severe model mismatch, and strong pilot dependence. These limitations make them unsuitable for low-overhead, high-performance 6G transmission. This paper aims to design an off-grid blind NF-ISAC algorithm and derive the theoretical performance bound for near-field sensing.  Methods  To overcome the limitations of analytical geometric steering vectors and adapt to more accurate electromagnetic propagation characteristics without closed-form expressions, an amplitude-phase separation method is first proposed. This method decomposes the nonlinear near-field steering vector into amplitude and phase terms, enabling high-precision characterization of the steering vector using a single-hidden-layer neural network. Second, the NF-ISAC problem is formulated as a constrained matrix factorization problem, and a corresponding factor graph model is constructed. The trained neural network is embedded into the factor graph as a function node. Message passing through the embedded neural network is then achieved, enabling joint blind coordinate sensing, channel estimation, and signal detection in a pilot-free manner. Finally, the Cramér-Rao Lower Bound (CRLB) for multi-user near-field joint distance and angle sensing in polar coordinates is derived based on the neural-network-fitted steering vector.  Results and Discussions  Extensive Monte Carlo simulations are conducted to evaluate the performance of the proposed algorithm. The simulation results show that the proposed algorithm achieves millimeter-level position sensing. Compared with existing mainstream algorithms, it improves both communication Bit Error Rate (BER) and sensing accuracy. The proposed algorithm achieves a 2~3 dB gain in sensing accuracy over the state-of-the-art near-field off-grid algorithm, and its performance is closest to the derived theoretical CRLB. These results indicate that the proposed algorithm effectively mitigates off-grid power leakage and model mismatch.  Conclusions  The proposed off-grid blind NF-ISAC algorithm overcomes the pilot dependence and model mismatch of existing NF-ISAC schemes. It achieves integrated high-precision sensing and reliable communication for near-field users in a pilot-free manner. The derived CRLB provides a theoretical benchmark for evaluating the sensing performance of NF-ISAC systems. This work provides technical support for the design of 6G NF-ISAC systems.
A Lightweight Spatial-Spectral Dual-Branch Transformer Network for Classifying Polarized White Blood Cell Hyperspectral Images
YANG Yushi, YAN Jiaxuan, XIE Yi, QIU Lijia, HUANG Danfei
Available online  , doi: 10.11999/JEIT260124
Abstract:
  Objective  White Blood Cell (WBC) classification and morphology are essential in routine blood analysis and provide important information for disease diagnosis and health assessment. Current WBC classification mainly relies on hematology analyzers and manual microscopic examination. Hematology analyzers cannot acquire cellular images, limiting classification accuracy when abnormal cellular characteristics are present, whereas manual microscopic examination depends on operator experience and is susceptible to human error. Although automated WBC classification based on deep learning and computer vision has attracted considerable attention, methods using stained images are sensitive to staining conditions and image quality. In addition, conventional hyperspectral imaging has limited ability to distinguish WBC subtypes with highly similar morphological and spectral characteristics. To address these limitations, this study combines Polarized Hyperspectral Imaging (PHSI) with deep learning and proposes a lightweight classification framework for polarized hyperspectral WBC images, providing an efficient and reliable approach for clinical decision support.  Methods  A polarized hyperspectral microscopic imaging system is established to acquire images at multiple polarization angles. Based on Stokes Vector Theory, polarization parameters are calculated to generate multidimensional data cubes containing both light intensity and polarization-state information. Degree of Linear Polarization (DOLP) images are then computed to construct a PHSI dataset of WBCs. To exploit the multidimensional characteristics of PHSI, a Lightweight Spatial-Spectral Dual-Branch Transformer Network (LSDBT) is proposed. After preprocessing, the input data are fed into a dual-branch feature extraction module that separately extracts local spatial features and joint spatial-spectral features. An adaptive scaling factor is introduced to fuse the two feature streams and balance their contributions, enabling effective utilization of the multidimensional information contained in PHSI. A lightweight Swin Transformer backbone performs effective global feature modeling while reducing computational complexity. Global average pooling and a fully connected layer are used for classification. Model performance is evaluated using Overall Accuracy (OA), Precision, Recall, Specificity, and F1-Score. Ablation studies, comparative experiments, and feature visualization are conducted to validate the proposed method.  Results and Discussions  The DOLP spectra and PHSI visualizations of monocytes, lymphocytes, and neutrophils (Figures. 5 and 6) demonstrate different polarization characteristics that reflect the selective absorption and scattering of light by their internal structures. Compared with conventional intensity images, PHSI improves image contrast and provides additional polarization information that enhances discrimination among WBC types. The proposed LSDBT achieves an OA of 99.29% on the test set, with consistently high classification performance across all cell categories (Table 1). Analysis of the adaptive scaling factor shows that classification performance first improves and then decreases slightly as the scaling factor increases, with the optimal value of 3 providing the best balance between spatial and spatial-spectral features (Figure. 8). Ablation experiments (Tables 2 and 3) demonstrate that the dual-branch feature extraction module substantially improves classification performance, whereas the lightweight design greatly reduces computational complexity and model parameters with only a marginal reduction in accuracy. Compared with conventional hyperspectral imaging, the PHSI dataset achieves higher classification accuracy with all evaluated classifiers, indicating that polarization information provides complementary physical features that improve discrimination among WBC types (Table 4). Comparisons with representative methods show that LSDBT achieves the best overall classification performance across multiple evaluation metrics (Table 5). Furthermore, t-SNE visualization (Figure. 9) shows compact intra-class distributions and clear separation among different cell types, confirming the strong discriminative capability of the learned features. Although LSDBT does not have the lowest computational cost among the compared methods, it achieves the best balance between classification performance, model size, and computational efficiency (Table 6).  Conclusions  This study proposes a lightweight dual-branch Transformer network for polarized hyperspectral WBC classification. To the best of our knowledge, this is the first study to combine PHSI with deep learning for WBC classification. Comparative experiments with conventional hyperspectral imaging validate the superiority of PHSI for WBC classification. The proposed LSDBT integrates spatial and spatial-spectral information through a dual-branch feature extraction module and performs efficient global feature modeling using a lightweight Swin Transformer backbone. The network maintains high classification performance while substantially reducing computational complexity and model parameters. These results demonstrate that LSDBT provides an accurate and computationally efficient solution for automated WBC classification and supports the application of PHSI in cellular microscopic analysis and clinical auxiliary diagnosis.
Research on Adaptive Hybrid Beamforming Method for Massive MIMO LEO Satellite Communication Systems
XIAN Yongju, HUANG Xiaolong, XING Zhitong, LI Yun
Available online  , doi: 10.11999/JEIT260458
Abstract:
  Objective  With the growing demand for high-capacity and high-spectral-efficiency transmission in LEO satellite communications, massive MIMO has become a promising enabling technology. However, conventional fully digital beamforming is difficult to implement in practice due to the strict constraints on power consumption, hardware complexity, and payload cost of satellite platforms. Although partially connected hybrid beamforming can reduce hardware complexity, the conventional fixed subarray structure lacks flexibility and suffers from performance degradation, especially under low-resolution PSs constraints. To address this issue, this paper investigates adaptive antenna-RF chain mapping for hybrid beamforming design in LEO satellite massive MIMO systems.  Methods  This paper first establishes a system model for LEO satellite multi-user downlink massive MIMO hybrid beamforming and formulates a joint optimization problem with the objective of maximizing spectral efficiency. Considering the constant-modulus discrete phase constraints of low-resolution PSs, antenna-RF chain connection constraints, and transmit power constraint, the resulting problem is highly non-convex. To solve it, the original problem is transformed into an equivalent WMMSE formulation, and auxiliary variables are introduced to decouple the coupled variables. Based on this reformulation, a double-layer iterative optimization framework is developed by combining the BCD method and the PDD method. For adaptive antenna-RF chain mapping, the mapping problem is reformulated as a capacity-constrained linear assignment problem, and an optimal adaptive mapping method based on the Hungarian algorithm is proposed. Furthermore, to reduce the computational burden in large-scale antenna array scenarios, a low-complexity adaptive mapping method based on antenna priority sorting is developed.  Results and Discussions  Simulation results show that the proposed methods exhibit good convergence behavior. Specifically, the spectral efficiency increases rapidly in the initial iterations and then gradually converges, while the constraint violation decreases continuously, confirming the effectiveness of the proposed iterative optimization framework (Fig. 3). In terms of spectral efficiency, the proposed adaptive mapping methods consistently outperform the conventional fixed subarray and the existing greedy dynamic subarray scheme over different transmit powers and antenna scales. Among them, the Hungarian-based method achieving better spectral efficiency, whereas the antenna priority sorting-based method attains near-optimal performance with significantly reduced computational complexity (Figs. 4 and 5). As the number of PSs quantization bits increases, the system performance gradually approaches that of the continuous-PSs case, demonstrating the effectiveness of the proposed low-resolution PSs-based design (Fig. 6). In terms of energy efficiency, the proposed methods also outperform the conventional fully digital beamforming, fully connected, fixed subarray, and greedy dynamic subarray hybrid beamforming under different transmit powers and antenna scales (Figs. 7 and 8).  Conclusions  This paper proposes an adaptive hybrid beamforming design for LEO satellite massive MIMO systems under low-resolution PS constraints. By combining WMMSE reformulation with a PDD-BCD based optimization framework, joint design of digital precoding, analog precoding, and adaptive antenna-RF chain mapping is achieved. Simulation results demonstrate that the proposed methods provide superior performance in both spectral efficiency and energy efficiency. In particular, the Hungarian-based method provides better system performance, while the antenna priority sorting-based method achieves a favorable trade-off between performance and computational complexity. The proposed design provides an effective solution for high-performance hybrid beamforming in hardware-constrained LEO satellite massive MIMO systems.
Impact of Wireless Priors on the Computation and Energy Cost of MU-MIMO Precoding Learning
CONG Pengyu, HAN Shengqian, DENG Mingyu, LIU Shengjie, YANG Chenyang, SHEN Songhui
Available online  , doi: 10.11999/JEIT260388
Abstract:
  Objective  This paper studies the learning of downlink MU-MIMO (Multiuser Multi-Input Multi-Output) precoding policy from the perspective of computational complexity and energy consumption. Traditional numerical optimization achieves strong performance but incurs rapidly growing computational complexity with the number of base-station’s antennas and served users, leading to high inference latency and inference energy consumption. In recent years, deep learning has been widely used to reduce online computational cost, yet existing evaluations typically rely on training/inference time or FLOPs and lack direct energy/power measurements. More importantly, the computational cost of a deep model is tightly coupled with the network architecture, which should be designed to effectively exploit the prior knowledge of the precoding policy. The objective of this paper is thus to develop a network architecture that matches the multi-dimensional permutational properties of the policy, and to analyze how policy priors influence the complexity and energy through comprehensive hardware-based energy/power measurements and simulations.  Methods  We formulate the MU-MIMO precoding policy as a mapping from multi-user channel information to the optimal precoding matrix under a transmit power constraint. The optimal policy exhibits multi-dimensional joint permutation equivariance and invariance with respect to user indices, receive-antenna indices, and base-station antenna indices. To exploit this prior, we propose an Attention-based Graph Neural Network (AGNN) built on a hypergraph structure, whose update and aggregation procedures are designed to satisfy the required equivariance/invariance properties. An attention mechanism modeling inter-user interference is introduced to improve generalization to varying number of users. For broadband precoding, the input layer aggregates multi-subcarrier channel information into an expanded input representation. To quantify energy cost, we build a mixed CPU/GPU measurement framework that collects energy and power for GPU, CPU, and DRAM during training and inference. Simulations use 3GPP TR 38.901 UMa channel data with different antenna array sizes and bandwidth settings. We compare the proposed AGNN against numerical baselines (ZFBD+Greedy pairing) and two Transformer-based architectures, where only one-dimensional permutation properties are satisfied.  Results and Discussions  : The paper reveals two main findings. First, partial exploitation of policy priors leads to poor performance and high cost. In the MU-MISO scenario, the Transformer variants with one-dimensional permutation equivariance yield lower spectral efficiencies than the numerical baseline ZFBD+Greedy and much larger model sizes and inference FLOPs than AGNN. In contrast, AGNN, designed to match the multi-dimensional permutation properties, achieves higher spectral efficiency while reducing inference FLOPs by about one order of magnitude. Hardware measurements further show that AGNN reduces inference energy and power on both CPU and GPU. Second, in the MU-MIMO scenario with small- and large-scale settings, ZFBD+Greedy increases the system sum rate to 10.9 times, but raise inference FLOPs to 436.4 times, inference time to 20.8 times, and inference energy to 40.8 times. Conversely, AGNN increases the sum rate to 11.5 times while inference FLOPs rise to only 5.3 times; inference time and inference energy become dramatically smaller (0.03 times and 0.22 times, respectively). These results suggest that matching the precoding policy’s multi-dimensional permutation priors is an effective way to reduce the computational cost, including FLOPs, latency, and energy consumption, in large-scale 6G MU-MIMO systems.  Conclusions  This paper investigated how policy priors affect computational complexity and energy consumption in MU-MIMO precoding learning. By analyzing the multi-dimensional joint permutation equivariance/invariance properties of the optimal precoding policy, we designed an attention-based GNN (AGNN) that matches these properties. A hardware-aware measurement platform was built to obtain direct energy/power measurements for training and inference on CPU, GPU, and DRAM components. Simulations on 3GPP TR 38.901 channel datasets showed that Transformer architectures satisfying only one-dimensional permutation properties can lead to both inferior spectral efficiency and significantly higher computational/energy overhead. In contrast, AGNN achieves higher spectral efficiency while substantially reducing inference FLOPs, inference time, inference energy, and training complexity. As system size increases, traditional numerical methods incur rapidly growing energy costs relative to rate gains, whereas the proposed policy-prior learning method maintains low inference time and energy consumption. Overall, exploiting the prior knowledge of MU-MIMO precoding policy for network architecture design is a key enabler for energy- and complexity-efficient high-dimensional precoding optimization in future 6G networks.
Resource Allocation for Multi-UAV Relay Networks in 6G Semantic Communication
XIAO Liming, GUAN Zheng, LIU Jie, YU Jihong, CHEN Liyuan
Available online  , doi: 10.11999/JEIT260520
Abstract:
  Objective  Sixth-generation (6G) mobile networks aim to achieve global seamless coverage through space-air-ground integrated architectures. In this context, Unmanned Aerial Vehicles (UAVs) can serve as mobile aerial relay nodes to support massive ground user access in complex environments. However, traditional data-oriented communication paradigms incur high bandwidth overhead, making them difficult to apply under the limited spectrum resources of UAV networks. In addition, traditional centralized resource allocation methods struggle with real-time implementation due to highly dynamic network topologies and the complex coupling of multi-dimensional resources. To address these challenges, semantic communication has emerged as a new paradigm that extracts the essential meaning of information at the transmitter and reconstructs it at the receiver, thereby reducing redundant data traffic. Although semantic communication provides a feasible solution, existing semantic-driven resource allocation methods primarily focus on static ground networks or single-UAV scenarios, often overlooking the coverage limitations and co-channel interference in multi-UAV collaborative networks. Therefore, an intelligent joint resource optimization model and a distributed resource allocation framework are needed for multi-UAV relay networks. By jointly optimizing multi-dimensional resources, such a framework can improve semantic transmission efficiency and ensure long-term user fairness, thereby supporting intelligent resource scheduling in 6G integrated networks.  Methods  This study formulates a joint resource optimization model for multi-UAV relay networks under the semantic communication paradigm (Fig.1), coupling the number of semantic symbols, UAV trajectories, transmission power, and channel allocation. To evaluate semantic transmission quality, the proposed model introduces a Semantic Communication Quality of Service (SC-QoS) metric, which integrates Semantic Quantization Efficiency with normalized transmission delay. The optimization objective is formulated as a Mixed Integer Nonlinear Programming problem that maximizes the weighted sum of system-wide SC-QoS and long-term user fairness evaluated by Jain's fairness index. To solve this problem, a Two-Stage Hybrid Reinforcement Learning (TS-HRL) framework is proposed (Fig.2). The first stage employs a Capacity-Aware K-means (CA-K-means) algorithm for heuristic UAV pre-deployment. By introducing a dynamic distance compensation term based on residual capacity, this algorithm guides edge users toward low-load UAVs, achieving load balancing among UAV clusters while preserving spatial proximity. The second stage models the dynamic scheduling problem as a Decentralized Partially Observable Markov Decision Process and solves it using a Recurrent Independent Proximal Policy Optimization with Parameter Sharing (R-IPPO-PS) algorithm (Algorithm 1). This algorithm leverages Long Short-Term Memory (LSTM) networks to aggregate historical trajectories and observations, thereby capturing hidden environmental states. Furthermore, the multi-agent parameter-sharing mechanism reduces model complexity with respect to the agent scale, while the preference decoupling strategy converts discrete channel allocation decisions into continuous preference variables to facilitate model training.  Results and Discussions  The proposed TS-HRL framework is evaluated in a dynamic environment with randomly moving ground users. The convergence analysis shows that the proposed method achieves a higher initial reward and reaches a stable state with fewer iterations than the random deployment, memoryless allocation, and Enhanced Independent Soft Actor-Critic (EI-SAC) baselines (Fig.3). By leveraging LSTM-based historical information and CA-K-means topology initialization, the proposed method reduces invalid exploration and improves convergence stability. Compared with traditional bit-based communication, the semantic communication framework increases Semantic Spectral Efficiency (S-SE) by 3.6 times, reduces transmission delay by 92.1%, and improves fairness by 14.3% (Fig.4). Within the semantic communication paradigm, compared with the memoryless and heuristic schemes, the proposed method improves S-SE by 92.7% and 51.9%, reduces transmission delay by 13.5% and 57.7%, and improves fairness by 17.3% and 39.7%, respectively. Although the EI-SAC baseline achieves a fairness index of 0.93, its S-SE remains relatively low. In contrast, the proposed TS-HRL framework maintains a high fairness level while achieving an S-SE approximately 4.3 times higher than that of EI-SAC. Compared with random deployment, the proposed method improves S-SE by 11.3% with only a marginal 1.1% decrease in fairness, demonstrating a better trade-off between transmission efficiency and fairness. As the number of users increases from 10 to 40, most baselines show decreases in S-SE and fairness due to intensified co-channel interference and spectrum constraints (Fig.5). In contrast, the adaptive scheduling and global fairness reward mechanisms of the proposed method mitigate performance degradation and maintain the average transmission delay below 0.1 ms. These results indicate that the proposed method improves overall system performance while ensuring long-term user fairness.  Conclusions  This paper investigates joint resource allocation and trajectory planning in dynamic 6G multi-UAV relay networks by integrating the semantic communication paradigm. A joint optimization model coupling the number of semantic symbols, UAV trajectory, power control, and channel allocation is formulated to maximize system SC-QoS and long-term user fairness. To address the high-dimensional coupling of the formulated problem, a TS-HRL framework is proposed. This framework combines CA-K-means pre-deployment with the R-IPPO-PS algorithm to handle multi-dimensional resource scheduling under partial observability. Simulation results verify the convergence, stability, and scalability of the proposed method under different user densities. Through distributed cooperation among UAVs, the proposed method enhances semantic transmission efficiency and reduces transmission delay while ensuring long-term service fairness among ground users. These findings provide a viable approach for intelligent resource allocation in UAV-assisted semantic communication for future space-air-ground integrated networks.
LEO Satellite Multi-beam Multicast Precoding and User Grouping Joint Optimization Algorithm
GUO Lili, FENG Yimeng, YUAN Peihong, GAO Yue
Available online  , doi: 10.11999/JEIT260375
Abstract:
  Objective  In 6G LEO communication, multicast precoding is utilized to mitigate the significant inter-beam interference introduced by Full Frequency Reuse (FFR). However, traditional precoding algorithms are hindered by cubic computational complexity, making them unsuitable for massive MIMO. Additionally, existing user grouping strategies often fail to meet the fixed group size requirements of the DVB-S2X standard. Therefore, a joint optimization scheme is developed to integrate a low-complexity unsupervised deep learning model for precoding with improved user grouping algorithms to enhance sum rate and fairness.  Methods  To address the precoding challenge, an unsupervised deep learning model based on a hybrid CNN-LSTM architecture is proposed (Fig. 2). The Convolutional Neural Network (CNN) is utilized to extract spatial features from the Channel State Information (CSI), while the Long Short-Term Memory (LSTM) network captures their deep feature correlations. Unlike supervised learning, this model is trained by directly maximizing the sum rate as the loss function, subject to the Per-Antenna Constraint (PAC). For user grouping, two algorithms are developed to comply with the DVB-S2X standard. First, the CK-means algorithm is introduced, which modifies the standard K-means to ensure an equal number of users in each group. Second, FA-MAUG algorithm is proposed which prioritizes users with poor channel quality, thereby enhancing the overall robustness and fairness of the grouping result.  Results and Discussions  The intra-group similarity metric is employed to evaluate user grouping performance, which indicate that the CK-means algorithm achieves a similarity score approximately 0.1 higher than the MAUG algorithm and nearly 0.5 higher than random grouping across various group sizes (Fig. 3). This high similarity translates directly into better beamforming gain. In terms of throughput, the sum rate of the CNN-LSTM precoding combined with CK-means grouping outperforms traditional MMSE across different SNRs and total powers (Fig. 4, Fig. 5). The sum rate based on the proposed CNN-LSTM precoding scheme is on average 48.59% higher than that of the traditional MMSE algorithm and the gain of the CK-means algorithm compared to random grouping reaches 30.12% under different SNR conditions. Additionally, we analyzed the effects of the number of users per group, the number of groups, and the number of antennas on the system rate(Fig. 6, Fig. 7, Fig. 8), verifying the performance of the model in systems of different scales. Additionally, the complexity analysis confirms that the proposed precoding approach reduces the online overhead from the cubic level of traditional MMSE to a linear level relative to the number of antennas, which is critical for real-time deployment in massive MIMO satellite systems.  Conclusions  This paper presents a joint optimization framework for LEO satellite multicast systems, addressing the dual challenges of high computational complexity in precoding and lack of fairness in user grouping. Simulation results demonstrate that the proposed solution significantly enhances the system sum rate—by over 48.59% in typical SNR scenarios—compared to traditional methods. The approach provides a robust and scalable solution for the coordinated management of multi-beam interference in future 6G satellite communications. Future work will further explore the impact of hardware impairments and highly dynamic channel conditions on the generalization capabilities of the proposed model.
Physics-Aware Reconstruction for Millimeter-Wave Radar Gait Recognition Under Complex Wearing Scenarios
HUANG Ling, QIU Liying, WANG Jiacheng, HAN Penglin, ZHOU Qingdi, YAN Huimei
Available online  , doi: 10.11999/JEIT260522
Abstract:
  Objective  Millimeter-wave (MMW) radar gait recognition has demonstrated considerable potential in the domain of non-contact biometric identification, primarily due to its inherent advantages in privacy preservation and resilience to variable lighting conditions. However, a significant challenge persists when applying these systems to unconstrained real-world environments where subjects may wear complex clothing, such as long coats or carry backpacks. These external covariates introduce non-stationary, high-frequency spectral components that result in severe spectral aliasing and masking of the intrinsic micro-Doppler (m-D) signatures of the human body. Traditional deep learning approaches usually treat radar spectrograms as generic image data and neglect the physical relationship between Doppler frequency shifts and human motion. Consequently, clothing-induced interference may overlap with motion-related signals, reducing recognition accuracy. Therefore, it is imperative to develop a framework that integrates radar physics and biomechanical properties to achieve effective signal decoupling and interference suppression. This study aims to provide a robust solution for radar-based gait recognition in complex scenarios by incorporating explicit physical constraints into the neural network architecture.  Methods  To mitigate the adverse effects of clothing-induced interference, this paper introduces PRISM-Net, a physics-aware reconstruction framework for millimeter-wave radar gait recognition. The framework is grounded in the biomechanical characteristics of human motion, specifically the observation that the human torso, representing the primary mass, generates stable and low-frequency Doppler shifts, whereas the movement of limbs produces higher-frequency, periodically alternating spectral components. (1) Physics-aware Frequency Structural Reconstruction: The proposed method diverges from conventional uniform processing by implementing a structural decoupling strategy. Based on the velocity distribution defined by biomechanical properties, the aliased original m-D spectrogram is partitioned into discrete frequency bands. The low-frequency band is dedicated to the torso, capturing stable and identity-persistent features, while the high-frequency band encompasses limb dynamics. This spatial-frequency partitioning enables the network to isolate the spectral regions most susceptible to clothing-induced clutter at the physical layer, thereby suppressing interference propagation. (2) Adaptive Weighted Attention Mechanism (WAM): An adaptive WAM module is integrated to regulate the signal-to-noise ratio (SNR) in the feature space. Given that clothing interference is dynamic and predominantly occupies high-frequency regions, the WAM adaptively evaluates the reliability of different frequency-derived features. In cases where the high-frequency limb features are corrupted by non-stationary noise from swinging garments, the WAM applies soft-thresholding to reduce their response weights. Simultaneously, the contribution of the more reliable torso-related features is enhanced. (3) Experimental Configuration and Protocol: The model is evaluated using the MMRGait-1.0 dataset, which provides a comprehensive set of radar gait signatures under various covariates. To ensure the assessment reflects real-world generalization, an open-set testing protocol is strictly followed. The training set comprises data from 74 subjects, while 47 entirely unseen subjects are used for evaluation. All radar spectrograms are processed to a standard resolution of 224 × 224 pixels. The network is optimized using the AdamW algorithm with a joint loss function consisting of cross-entropy and feature regularization components to ensure both classification accuracy and feature discriminability.  Results and Discussions  The experimental results demonstrate the effectiveness of integrating the biomechanical characteristics of human motion into the recognition process. As shown in Table 1, PRISM-Net achieves an average Rank-1 recognition rate of 89.7% in the 90-degree side-view scenario. In the coat (CT) occlusion scenario, the model maintains an accuracy of 85.1%. This is a 18.1% improvement over traditional lightweight models such as ShuffleNetV2 and a 7.4% improvement over the standard ResNet-18 baseline. The stability of the model was verified through ten independent repeated trials. As illustrated in the box plot in Fig. 1, PRISM-Net shows a low standard deviation of ±0.55%, whereas the baseline model exhibits a higher deviation of ±1.75%. An independent samples t-test results in a p-value of less than 0.001, confirming statistical significance. Ablation studies in Table 2 confirm the contribution of each module. Removing frequency decoupling reduces accuracy in the CT scenario to 75.5%, proving that physical structural separation is necessary to prevent feature distortion. The WAM module further provides a 4.2% accuracy gain by adaptively filtering noise. Regarding computational efficiency, Table 3 shows that PRISM-Net requires 11.33M parameters and 0.75 GFLOPs. Compared with heavy 3D-convolutional models that require over 10 GFLOPs, the proposed method achieves superior accuracy with significantly lower resource consumption. Finally, t-SNE visualization in Fig. 5 shows that PRISM-Net produces more compact intra-class distributions and clearer inter-class margins demonstrating improved feature discriminability.  Conclusions  This research demonstrates that the integration of the biomechanical characteristics of human motion into deep learning architectures improves the robustness of radar gait recognition. The proposed PRISM-Net framework effectively decouples motion components and suppresses clothing-induced noise through physics-aware frequency reconstruction and adaptive weighting. The high recognition accuracy and low computational complexity observed in the MMRGait-1.0 dataset suggest that this approach is suitable for real-time security applications on edge-computing devices.
FedFACO: Personalized Federated Learning Method Based on Fisher Information Matrix for Adaptive Aggregation and Client Collaborative Optimization
JIANG Wei-Jin, LIU Zhi-Hua, CUI Xin-Yu, XU Yu-Sheng, CHEN Shen-You, HU Jia-Long
Available online  , doi: 10.11999/JEIT260344
Abstract:
  Objective  Non-IID data heterogeneity remains one of the major challenges in personalized federated learning, as it often leads to inconsistent local optimization directions, insufficient global knowledge transfer, and degraded model personalization performance. To address these issues, this paper proposes FedFACO, a personalized federated learning method based on Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The proposed method aims to improve the adaptability of federated models to heterogeneous client distributions while maintaining effective knowledge sharing across clients. By introducing an adaptive aggregation mechanism and a collaborative optimization strategy, FedFACO provides a more principled way to balance global generalization and local personalization, which is particularly important in complex Non-IID federated environments.  Methods  FedFACO consists of two key components. First, an adaptive aggregation (AA) mechanism is employed to dynamically adjust fusion weights between global and local models based on client-specific states, generating personalized initializations aligned with local data. Second, a collaborative optimization (CO) mechanism is introduced, combining feature alignment with FIM-based client weighting. This enhances useful global knowledge transfer and suppresses low-quality updates. The FIM is utilized to quantify the information contribution of each update, ensuring reliable aggregation. The method is evaluated on MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet under Non-IID settings, and compared with representative baselines. Convergence behavior, dropout robustness, and sensitivity to low-quality updates are also examined.  Results and Discussions  Experimental results demonstrate that FedFACO consistently outperforms competitive baseline methods across all four benchmark datasets, achieving an average accuracy improvement of approximately 3.1% over mainstream approaches (Fig. 1, Table 2). On the more challenging Tiny-ImageNet dataset, FedFACO reduces the total training time required to reach convergence by approximately 4.8% compared with the best baseline (Table 3). Ablation studies confirm that performance is substantially improved by both AA and CO mechanisms, with their joint application yielding optimal accuracy (Table 6). Furthermore, FIM-guided weighting is shown to accurately quantify contribution quality in client dropout and asynchronous scenarios, significantly enhancing aggregation reliability (Fig. 2Fig. 3). Superior robustness is also demonstrated in malicious client scenarios (Fig. 4).  Conclusions  This paper presents FedFACO, a personalized federated learning method for Non-IID environments via Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The method effectively balances global knowledge sharing and local personalization while enhancing training stability and robustness under heterogeneous client participation. Experimental results validate its effectiveness and superiority in accuracy, convergence efficiency, and robustness. Future work will focus on lightweight Fisher information approximations and adaptive triggering strategies to reduce computational overhead, as well as integration with privacy-preserving and security-defense mechanisms for deployment in resource-constrained and high-security environments.
Channel Estimation for MIMO-OFDM Based on Adaptive Transformer Network in High-Speed Mobile Scenarios
LIAO Xi, HE Xiangni, ZHANG Zhe, WANG Yang
Available online  , doi: 10.11999/JEIT260075
Abstract:
  Objective  Accurate channel state information (CSI) is essential for coherent detection, beamforming, and adaptive resource allocation in multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. In high-mobility environments, large Doppler shifts and multipath propagation jointly cause doubly selective fading, destroy subcarrier orthogonality, and aggravate inter-carrier interference. Consequently, conventional least squares (LS) and linear minimum mean square error (LMMSE) estimators suffer substantial performance degradation. Existing deep learning estimators provide strong nonlinear modeling capability, but often show insufficient adaptability to variations in signal-to-noise ratio (SNR), delay spread, and Doppler shift, or inefficiently incorporate physical channel priors. To address these problems, a Transformer network with adaptive feature modulation, termed AdaFiT, is proposed for channel estimation in high-mobility MIMO-OFDM systems.  Methods  The proposed AdaFiT framework performs channel estimation by jointly exploiting local time-frequency correlations, global dependencies, and explicit channel-aware adaptation. It takes least squares (LS) estimates at pilot positions, together with signal-to-noise ratio (SNR), delay spread, and maximum Doppler shift, as inputs. A separable two-dimensional linear upsampling module first interpolates the sparse pilot estimates to the full OFDM time-frequency grid by independently processing the real and imaginary components along the frequency and time dimensions. A convolutional feature enhancement module then extracts robust local representations: a complex feature mixing layer fuses multi-antenna real and imaginary components, while multi-scale convolutional blocks with channel attention capture short-range time-frequency correlations and suppress noise. Next, a feature-wise linear modulation (FiLM)-based channel-adaptive module embeds the three channel parameters through independent multilayer perceptrons and combines them into a channel-state representation. This representation generates scaling and shifting coefficients to dynamically recalibrate the block-embedded sequences according to varying channel statistics. Finally, the modulated sequences are processed by a Transformer encoder with learnable two-dimensional positional encoding to model long-range dependencies across the time-frequency grid and antenna dimensions. A residual reconstruction structure fuses the global features with locally enhanced representations, yielding accurate channel estimates with global consistency and preserved local details.  Results and Discussions  Simulation analyses are conducted based on the CDL-C and CDL-A channel models. The LS-based bilinear interpolation method, the LMMSE method, the AdaFortiTran model, and the AdaFiT model without the adaptive module are selected as comparison schemes. The mean squared error (MSE) performance of these models is compared under different signal-to-noise ratio (SNR), maximum Doppler shift, and delay spread conditions. Under the CDL-C channel model, AdaFiT achieves the lowest MSE over the entire SNR range (Fig. 3). At an SNR of 0 dB, its MSE is approximately 2.5 dB lower than that of AdaFortiTran, and the performance gain increases to approximately 7 dB at an SNR of 30 dB. Compared with the AdaFiT model without the adaptive module, AdaFiT achieves a maximum MSE gain of approximately 2.1 dB, which verifies the effectiveness of the proposed channel-adaptive feature modulation module. When the maximum Doppler shift increases from 200 Hz to 1400 Hz, AdaFiT consistently maintains the lowest MSE (Fig. 4). In the high-Doppler range of 10001400 Hz, AdaFiT outperforms the AdaFiT model without the adaptive module and AdaFortiTran by approximately 2 dB and 3.5 dB, respectively, demonstrating improved robustness against rapid channel variations. For delay spreads ranging from 100 ns to 700 ns, AdaFiT also maintains the lowest MSE over the entire range (Fig. 5), achieving approximately 3 dB and 5.5 dB gains over the AdaFiT model without the adaptive module and AdaFortiTran, respectively. These results demonstrate that the proposed adaptive feature modulation mechanism effectively improves the adaptability of the model to time-selective and frequency-selective fading.To further evaluate the performance of the AdaFiT model under different channel models, simulation analysis is conducted based on the 3GPP CDL-A channel model. At SNRs of 0–5 dB, AdaFiT outperforms LMMSE by approximately 4–5 dB, while the performance gain reaches approximately 7 dB at SNRs of 20–30 dB (Fig. 6(a)). Compared with AdaFortiTran, AdaFiT achieves an MSE gain of approximately 2 dB under low-SNR conditions, and the gain increases to approximately 5.5 dB under high-SNR conditions. In the maximum Doppler shift experiment, the MSE of AdaFiT remains between approximately –38 dB and –37 dB over the range of 200–800 Hz and is approximately 5 dB lower than that of AdaFortiTran (Fig. 6(b)). Although the MSE of AdaFiT increases when the maximum Doppler shift exceeds 1000 Hz, it still achieves an approximately 6 dB gain over LMMSE at 1400 Hz and remains superior to AdaFortiTran. Over the entire delay spread range, AdaFiT achieves approximately 4–6 dB gain over LMMSE and approximately 4–5.5 dB gain over AdaFortiTran (Fig. 6(c)).  Conclusions  The proposed AdaFiT framework performs channel estimation by jointly exploiting local time-frequency correlations, global dependencies, and an explicit channel-aware adaptation mechanism. Simulation results under the CDL-C and CDL-A channel models show that AdaFiT consistently achieves lower MSE than the LS-based bilinear interpolation method, LMMSE, AdaFortiTran, and the AdaFiT model without the adaptive module under different SNR, maximum Doppler shift, and delay spread conditions. These results verify the effectiveness of the proposed adaptive feature modulation mechanism and demonstrate that AdaFiT maintains high estimation accuracy and stable performance under different channel models, indicating good adaptability to varying channel environments.
A High-Parallelism Simulated Adiabatic Bifurcation Processor for Combinatorial Optimization Problems
LI Renlong, HAO Xin, CHEN Zhuojun, DING Ding
Available online  , doi: 10.11999/JEIT260780
Abstract:
  Objective  Combinatorial optimization problems (COPs) are widely encountered in fields such as network optimization, autonomous driving path planning, VLSI design, and computational biology. The solution space of these problems grows exponentially with problem size, making it impossible for traditional von Neumann architectures (e.g., CPUs) to find high-quality solutions within a feasible time. Quantum-inspired Ising machines have emerged as a promising computing paradigm for accelerating COP solving. However, existing CMOS-based Ising machines still face significant challenges in simultaneously achieving high speed, high energy efficiency, and high accuracy. Discrete-time Ising machines often require thousands of iterations and rely heavily on random number generators, leading to large chip area and long solving times. Continuous-time Ising machines suffer from poor solution quality, frequently getting trapped in local minima. To address these limitations, this paper designs and fabricates a high-parallelism simulated adiabatic bifurcation application-specific processor for combinatorial optimization problems in 65 nm CMOS technology.  Methods  The proposed processor adopts the simulated adiabatic bifurcation (SAB) algorithm, which is inspired by quantum adiabatic optimization. Unlike simulated annealing, SAB does not require Gibbs sampling or random number generators to escape local minima. The algorithm models each spin as a nonlinear oscillator and solves a set of ordinary differential equations to simulate the adiabatic evolution of a classical nonlinear Hamiltonian system exhibiting bifurcation phenomena. The processor builds a hardware architecture that supports fully connected spin topologies and leverages the inherent fully parallel spin update characteristic of the SAB algorithm. To achieve bubble-free iterative computation, a three-stage pipelined spin update strategy is proposed, dividing the update process into coupling coefficient access, momentum update, and position update. To reduce the storage overhead introduced by the fully connected coupling matrix, a folded coupling coefficient storage array is designed, exploiting matrix symmetry to eliminate redundant storage. The momentum update unit and position update unit are implemented using 8-bit fixed-point arithmetic (2 bits for integer, 6 bits for fractional part) to achieve SAB evolution with low hardware overhead. The chip is fabricated in 65 nm CMOS technology, occupying an area of 0.47 × 1.25 mm2 and operating at a 200 MHz clock frequency and 1 V supply voltage.  Results and Discussions  The chip achieves a total power consumption of only 14 mW, with the spin evolution module consuming 58% of the total power. The folded coupling coefficient storage array reduces storage area by 52% compared to full matrix storage, and the rectangular restructuring avoids irregular shapes in physical layout, reducing routing congestion and layout voids (Fig.6). The parallel loading access mechanism allows all coupling coefficients to be read and distributed within a single clock cycle, eliminating the memory access bottleneck inherent in serial reading. For predefined Max-Cut problems configured as 8×8 grid structures (64 nodes) with coupling coefficients quantized to 2-bit precision, the chip converges to the global optimum within only 5 computing cycles, achieving a final Ising energy of –4032 (Fig.9). This energy follows the analytical expression (n4−n2)(n4−n2), confirming that the chip solves Max-Cut problems with the shortest solving time. For larger extended Max-Cut problems (108×108 nodes), the simulated adiabatic bifurcation algorithm achieves 100% accuracy relative to the theoretical ground state (Fig.10). Monte Carlo simulations over 1,000 independent trials on randomly generated Max-Cut problems demonstrate that simulated adiabatic bifurcation achieves an average Hamiltonian of –3617.42, significantly outperforming simulated annealing which achieves only –3352.12 (Fig.11). For 3-SAT problems with 40 clauses and 8 variables, the chip solves instances with clause-to-variable ratios of 3, 4, and 5 in approximately 6 μs, 10 μs, and 16 μs, respectively (Fig.12). These results align perfectly with theoretical phase transition predictions, confirming the processor's effectiveness across varying problem complexities.  Conclusions  This work proposes an simulated adiabatic bifurcation machine that enables fully parallel updates without duplicating spin copies. To improve throughput, a three-stage pipeline strategy is designed that integrates coupling-coefficient access, momentum update, and position update, achieving bubble-free parallel updating and low-latency solving. For sparse coupling coefficients, a folded storage scheme is adopted to significantly reduce memory area overhead. Both momentum and position variables are represented in 8-bit fixed-point format, ensuring sufficient computational accuracy while balancing resource efficiency. Compared with previous fully connected Ising machines, the proposed bifurcation machine achieves 100% solving accuracy, along with higher energy efficiency and lower hardware overhead, demonstrating substantial application prospects in edge-side combinatorial optimization.
A State Prediction Method for Long-Endurance Fixed-Wing UAV Propulsion Systems
LI Sicheng, WANG Lianqing, LI Zhiyong, WANG Guochang, GE Kaihua, CHEN Junfeng, TAN Rongqing
Available online  , doi: 10.11999/JEIT260188
Abstract:
  Objective  Accurate single-step prediction of key propulsion-system states is essential for early fault warning and autonomous health management of long-endurance fixed-wing unmanned aerial vehicles (UAVs). During high-altitude missions lasting more than 24 h, electrical and thermal variables in the propulsion system exhibit strong coupling, multi-time-constant dynamics, and pronounced day-night regime shifts. These characteristics cause short-term disturbances and long-term drifts to coexist, and hinder adaptive feature weighting under time-varying variable sensitivities. General-purpose time-series predictors may therefore fail to meet the accuracy and robustness requirements of multivariate propulsion-state prediction. To address these challenges, a Grouped Squeeze-and-Excitation Multi-scale Temporal Convolutional Network (GEMS-TCN) is developed by enhancing a modern pure-convolution forecasting backbone with multi-scale embedding and grouped channel attention. The aim is to obtain accurate 10 s-ahead single-step forecasts for 16 key propulsion states from 72-dimensional flight telemetry while satisfying the real-time inference requirement of the 1 Hz telemetry cycle.  Methods  Real flight telemetry from a representative long-endurance fixed-wing UAV is used for model construction and evaluation. The data are sampled at 1 Hz for 9 consecutive days, yielding 806,629 time steps and 72 variables. (Fig.2) Sixteen propulsion-related key states, including the bus voltage, control-unit temperature, winding temperature, and power-device temperature of four motors, are selected as prediction targets, and all 72 variables are used as inputs. (Table 1) The raw data are processed through time indexing, interquartile range (IQR)-based anomaly handling, and interpolation, and are then chronologically divided into training, validation, and test sets at a ratio of 8:1:1 to avoid information leakage. (Fig.4) (Fig.5) GEMS-TCN uses a multi-scale embedding layer with parallel one-dimensional convolutions to extract temporal patterns at different receptive-field scales. Stacked GEMS-TCN blocks combine depthwise temporal convolution, grouped convolutional feed-forward networks, and Grouped Squeeze-and-Excitation (GroupSE) modules to recalibrate intra-variable and cross-variable channel responses hierarchically. (Fig.1) The models are trained with Adam and mean squared error (MSE) loss, and are evaluated using mean absolute error (MAE), MSE, and symmetric mean absolute percentage error (SMAPE). PatchTST, ModernTCN, FEDformer, DLinear, and TimesNet are used as comparison models, with ablation and robustness experiments conducted for further verification.  Results and Discussions  On the full 16-dimensional target set, GEMS-TCN achieves a test-set MAE of 0.171 and an MSE of 0.069. (Table 4) Compared with TimesNet, the strongest baseline in overall trend tracking, GEMS-TCN reduces MSE by 28.1% while maintaining comparable MAE and SMAPE, indicating stronger suppression of large prediction deviations. (Table 4) Stable accuracy is obtained in both daytime and nighttime segments, with MAE/MSE values of 0.189/0.085 and 0.150/0.051, respectively, demonstrating robustness to diurnal operating-condition changes. (Table 4) The prediction trajectories show tighter alignment at thrust-transition points, reduced overshoot, and fewer spurious spikes, while low-bias tracking is preserved for slowly varying nighttime temperature profiles. (Fig.7) Ablation results show that multi-scale embedding and GroupSE provide complementary improvements, and their combination achieves the best overall performance among the ablation settings. (Table 6) Under 1% and 5% random dropouts and 60 s continuous missing intervals, the MSE remains below 0.07, indicating tolerance to practical data-loss scenarios. (Table 5) In addition, GEMS-TCN contains 27.88 M parameters and achieves an inference latency of 0.90 ms per sample, which is well below the 1 s sampling interval.  Conclusions  GEMS-TCN provides a practical convolution-based solution for multivariate propulsion-state prediction in long-endurance fixed-wing UAVs. By integrating multi-scale temporal embedding with hierarchical group-wise channel recalibration, the proposed method jointly represents rapid fluctuations and slow evolution, and better captures multi-time-constant dynamics and multivariate coupling in propulsion telemetry. Real-flight experiments demonstrate stable prediction performance across diurnal regimes, state categories, and data-missing scenarios. Ablation results further confirm the effectiveness and complementarity of multi-scale embedding and GroupSE. These findings indicate that structure-aware modeling tailored to propulsion-state evolution can support health monitoring, early fault warning, and autonomous health management of long-endurance fixed-wing UAVs.
Customized Preparation of Single Crystal Tungsten Tips and Surface Reconstruction Mechanism
GUO Jiamei, YIN Shengyi, ZHANG Yongqing, SUN Wanzhong
Available online  , doi: 10.11999/JEIT260328
Abstract:
  Objective  Refractory metal tungsten, particularly single crystal tungsten, serves as a critical material for high-performance field emission cathodes, which are core components in advanced electron microscopes and electron beam lithography systems. The fabrication of single crystal tungsten microtips with precisely controlled geometry and surface cleanliness remains a significant challenge, especially for applications requiring tip radii ranging from sub-100nm for cold field emission to 0.3–1.0 μm for Schottky-type thermal field emission. Currently, the domestic development of high-end electron microscopes in China faces a major bottleneck due to heavy reliance on imported field emission cathodes. Electrochemical corrosion, the mainstream method for tip preparation, suffers from limitations such as difficulty in cutoff timing control, susceptibility to tip bending or passivation, residual surface impurities, and low yield rates. Moreover, the anisotropic nature of single crystal tungsten introduces additional complexity in morphology control. This study aims to establish a controllable fabrication method for single crystal tungsten tips, enabling tailored geometry and surface quality to meet demanding requirements of different field emission applications while significantly improving fabrication yield.  Methods  Single crystal tungsten wires (diameter 0.12 mm, purity 99.95%, (100) orientation) were used as the starting material. The tips were first pre-shaped by electrochemical corrosion in 1 mol/L NaOH solution under a 10 V DC voltage with pulsed control (6 kHz, 100 μs pulse width) for 10 min. Following corrosion, the tips underwent surface cleaning by sequential immersion in ultrapure water and anhydrous ethanol, followed by drying with high-purity nitrogen. Subsequently, the samples were placed in an ultra-high vacuum system (base pressure < 1 × 10–6 Pa) and subjected to high-temperature oxygen treatment. Process parameters were systematically varied, including temperature (15001600 °C), oxygen flow rate (0.015–1.0 sccm), and treatment duration (5–1440 min). During treatment, high-purity oxygen was introduced while maintaining vacuum levels better than 10–4 Pa. Tip morphologies were characterized by scanning electron microscopy, and surface compositions were analyzed by energy dispersive X-ray spectroscopy. Geometric parameters, including half-cone angle and tip radius, were measured following standardized protocols.  Results and Discussions  The combination of electrochemical corrosion and high-temperature oxygen treatment enabled both effective surface purification and controlled morphological reconstruction. Scanning electron microscopy characterization revealed distinct evolutionary pathways depending on oxygen partial pressure (Fig. 3, Fig. 4). Under oxygen-rich conditions (1.0 sccm, 15001600 °C), the tips underwent “blunting reconstruction,” evolving from an initial inverted cone into a characteristic “cylindrical segment + hemispherical cap” structure (Samples 2–5, Table 2). For Sample 3 treated at 1500 °C for 10 min, the tip radius increased to 327 nm with a half-cone angle reduction of 21%; for Sample 4 treated for 15 min, further evolution occurred with tip radius decreasing to 275 nm and cylindrical segment height increasing substantially. With increasing temperature from 1500 °C (Sample 2) to 1600 °C (Sample 5), the half-cone angle change rate increased from 11% to 25%, while tip radius decreased from 380 nm to 172 nm, indicating accelerated kinetics. This anisotropic behavior is attributed to orientation-dependent surface energies of tungsten and preferential oxidation along specific crystallographic planes. Under oxygen-lean conditions (0.015 sccm, 1500 °C, 1440 min), “sharpening reconstruction” occurred, with tip radius decreasing dramatically from 143 nm to 28 nm and half-cone angle reducing from 8° to 1° (Sample 6). In contrast, treatment in vacuum at 1500 °C primarily removed surface contaminants without altering tip geometry (Sample 1). EDS analysis confirmed that treated surfaces were free from impurities other than trace carbon contamination (Table 3), demonstrating the dual functionality of the process in achieving both purification and reconstruction. The findings reveal that oxygen partial pressure serves as the key determinant of reconstruction direction, with oxygen-rich conditions favoring blunting and oxygen-lean conditions promoting sharpening.  Conclusions  A combined process of electrochemical corrosion followed by high-temperature oxygen treatment was successfully developed for the controllable fabrication of single crystal tungsten field emitter tips. The process achieves dual functionality: effective surface purification through high-temperature vacuum treatment, which removes residual surface contaminants from electrochemical corrosion, and controlled morphological reconstruction enabled by the introduction of oxygen under precisely regulated conditions. By systematically adjusting process parameters including treatment temperature, oxygen flow rate, and treatment duration, tip morphology and dimensions can be tailored to meet specific application requirements. The oxygen partial pressure plays a decisive role in determining the reconstruction pathway. Under oxygen-rich conditions (e.g., 1500 °C, 1.0 sccm), blunting reconstruction yields a “cylindrical segment + hemispherical cap” structure ideally suited for Schottky-type thermal field emission cathodes requiring tip radii between 0.3 μm and 1.0 μm. Under oxygen-lean conditions (e.g., 1500 °C, 0.015 sccm), sharpening reconstruction produces nanoscale sharp tips with tip radius as low as 28 nm, suitable for cold field emission cathodes and scanning tunneling microscope probes. This anisotropic reconstruction mechanism is explained by the selective reaction of oxygen atoms with different tungsten crystal planes under high-temperature conditions, where the formation of volatile oxides modifies local surface energy distribution and promotes the exposure of specific crystallographic planes. Control experiments confirmed that such morphological reconstruction does not occur under oxygen-lean or oxygen-free conditions at the same temperatures, further demonstrating the critical role of oxygen in triggering this process. Importantly, this approach provides an effective method to correct imperfections from the initial electrochemical corrosion step, significantly improving fabrication yield and process robustness, thereby offering a viable technical pathway for the domestic fabrication of high-performance field emission cathodes.
Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding
ZHOU Tong, CHEN Hongzhi, XU Jialong, SUN Peng, JIANG Dajie, LIU Jiankang
Available online  , doi: 10.11999/JEIT260379
Abstract:
Currently, the 3rd Generation Partnership Project (3GPP) Technical Specification Group Service and System Aspects Working Group 4 (TSG SA Working Group 4, SA4) has begun researching Artificial Intelligence (AI)-based audio codecs. Based on the Descript Audio Codec (DAC) being studied by SA4, this paper presents a joint optimization scheme for DAC source-channel coding and modulation with total power constraints, a joint optimization scheme for DAC source-channel coding and modulation with constant modulus constraints, and a DAC joint source-channel coding scheme. Furthermore, simulation evaluation is conducted using the typical physical layer parameter configuration of 3GPP Geostationary Earth Orbit (GEO) voice scenario. Finally, preliminary research suggestions are made, with the hope of providing useful insights for the subsequent research on GEO voice.  Objective  In recent years, Joint Source-Channel Coding (JSCC) has gained widespread attention in academia and industry. In academia, the sources studied for JSCC include video, image, and audio. However, the industry has not focused on these sources as in academia, but rather on JSCC where channel state information (CSI) is considered as the source. The reason the industry has not considered JSCC for video and images is due to two main factors. First, for video and images, 3GPP generally does not conduct independent research, but instead directly reuses source compression standards developed by external organizations such as the Moving Picture Experts Group (MPEG) and the Joint Photographic Experts Group (JPEG). Second, it is constrained by the challenges of the source-channel coding architecture. For example, in joint source-channel coding at the application layer, the channel information obtained at the application layer is often outdated, which limits the potential gains of source-channel coding. However, the situation is quite different for audio. Audio coding and decoding are designed independently by 3GPP, and currently, 3GPP SA4 is researching AI-based speech coding and decoding. Moreover, in GEO speech scenarios, where satellites remain largely stationary relative to ground stations and the channel changes slowly, the physical layer CSI can be transmitted to the application layer, overcoming the issue of outdated channel information. This makes the research on JSCC for audio more promising within 3GPP, compared to JSCC for video or image sources.  Methods  This paper firstly uses the typical physical layer parameter configurations of 3GPP GEO voice and the DAC encoder, currently being studied by SA4, as an evaluation baseline, ensuring the fairness of the proposed scheme’s gain evaluation. Based on this, the paper proposes the integration of DAC with channel coding and modulation to improve audio quality in low-bitrate scenarios. First, the total power constrained (TPC) DAC-JSCCM scheme is considered, which offers the maximum potential gain due to the joint optimization of source coding, channel coding, and modulation. Then, considering the high PAPR (Peak-to-Average Power Ratio) impact on the power amplifier, the paper proposes a constant-envelope constrained (CEC) DAC-JSCCM scheme. Finally, since the modulation symbols included in RTP payload require significant changes to the existing protocol, a DAC-JSCC scheme that is more compatible with current protocols is proposed, where bit sequences are included in RTP payload.  Results and Discussions  In the evaluation of baseline scheme 1, the link-level simulation at the physical layer is first conducted to obtain the relationship between SNR and BLER for an input length of 64 bits using a 1/3 Turbo code. Then, assuming that the error rate of an audio frame is the same as the BLER, the transmission error rate of an RTP packet containing L audio frames is calculated. According to the existing processing logic, if one audio frame in an RTP packet is transmitted with an error, the entire RTP packet will be discarded. In real-time communication, VoIP, or speech enhancement scenarios, a PESQ score of 2.5 is generally considered the minimum usable threshold, and voice quality below this score is regarded as unsuitable for professional or commercial services. Therefore, in GEO voice calls, we choose PESQ = 2.5 as the reference point. When PESQ = 2.5, the TPC DAC-JSCCM and CEC DAC-JSCCM provide a coverage gain of approximately 4 dB compared to baseline scheme 1 (Fig.4). The DAC-JSCC scheme offers a coverage gain of 2.5 dB compared to baseline scheme 1 (Fig.4). The coverage gains of DAC-JSCCM and DAC-JSCC come from the application layer performing joint source-channel coding, which avoids the cliff effect caused by using Turbo coding at the physical layer. When the Complementary Cumulative Distribution Function (CCDF) of Peak-to-Average Power Ratio (PAPR) is \begin{document}$ {10}^{-3} $\end{document}, the TPC DAC-JSCCM is 4 dB higher than the QPSK modulation, with sharper signal peaks, higher requirements for power amplifier linearity, and is more prone to distortion. In contrast, the CEC DAC-JSCCM is very close to the QPSK modulated PAPR.  Conclusions  This paper proposes the Total Power Constrained DAC-JSCCM, Constant Modulus Constrained DAC-JSCCM, and DAC-JSCCM schemes, based on the DAC codec currently being researched by SA4. Simulations are conducted using the typical physical layer configuration of the existing 3GPP GEO voice scenario. The simulation results show that, when the PESQ is 2.5 (the minimum usable threshold), compared to DAC combined with existing physical layer transmission technologies, the proposed DAC-JSCCM provides a coverage gain of 4 dB, while the proposed DAC-JSCC scheme provides a coverage gain of 2.5 dB. SA4 is currently researching more advanced GEO voice codecs beyond DAC, and we will extend our research by combining better GEO voice codecs in the future.
Lightweight Semantic Communication System Driven by User Personalization in UAV Networks
WEI Yuxuan, CHEN Xiao, CHEN Qiuyu, JIANG Hao, YANG Zhaohui
Available online  , doi: 10.11999/JEIT260370
Abstract:
  Objective  With the rapid development of the low-altitude economy and 6G intelligent networks, Unmanned Aerial Vehicle (UAV) image communication shows strong potential in target reconnaissance, emergency communication, and intelligent inspection. However, conventional pixel-level transmission cannot meet the requirements of efficient, low-latency, and intelligent communication because UAVs are constrained by limited bandwidth, payload capacity, and onboard computational resources. Semantic communication, which transmits only task-relevant information, provides an effective solution for improving communication efficiency in resource-constrained scenarios. However, current studies on UAV image transmission face several challenges. First, fixed network architectures use unified semantic encoding and transmission strategies for all users and cannot adapt to different personalized requirements. Second, new user access usually requires interest pre-training or model fine-tuning, which increases deployment overhead. Third, most models have high computational complexity. To address these issues, this paper proposes the Lightweight Personalized UAV Semantic Communication (LPUSC) system to balance computational cost, transmission bandwidth, and personalized requirements. The system enables personalized transmission through low-overhead semantic index interaction and a lightweight semantic extraction module, without pre-training for new users. A dual-branch end-to-end network is also designed. In this network, the semantic index transmission network works with the semantic image transmission network trained by a weighted hybrid loss function, thereby supporting high-precision and high-quality transmission of personalized semantic images.  Methods  The proposed LPUSC system adopts a dual-branch architecture for accurate task-driven semantic content transmission. In the semantic index interaction branch, the lightweight object detection model YOLO11s is used to perform semantic perception on UAV-captured visual scenes. Complex image information is compressed into low-dimensional semantic index vectors, which reduces transmission redundancy and communication overhead. On this basis, an end-to-end semantic index transmission network is designed to improve the robustness of semantic index transmission under complex wireless channel conditions. Through the semantic index interaction mechanism, the system accurately identifies targets of user interest and provides prior guidance for subsequent semantic content extraction. In the semantic image transmission branch, the lightweight and high-precision MobileSAM model is adopted for semantic region extraction. This branch uses the target bounding boxes returned by the semantic index interaction branch as box-prompt inputs, enabling pixel-accurate segmentation and extraction of specific semantic targets. To further improve semantic image reconstruction quality, a weighted hybrid loss function is designed. This function integrates Mean Squared Error (MSE), L1-norm loss, Structural Similarity Index Measure (SSIM) loss, gradient loss, perceptual loss, and background suppression loss. These losses jointly optimize pixel accuracy, structural preservation, and fine-detail restoration. Through the joint constraints of multiple loss terms, the proposed system improves semantic region reconstruction and achieves high-quality semantic image transmission.  Results and Discussions  Simulation results validate the proposed LPUSC system in semantic extraction and end-to-end transmission. For semantic extraction, three schemes are compared: YOLO11s-seg, YOLO11s + Segment Anything Model (SAM), and YOLO11s + MobileSAM (Fig. 4). The results show that the detection-segmentation decoupled architecture achieves better semantic boundary localization accuracy. Combined with the quantitative analysis in Table 1, the YOLO11s + MobileSAM scheme reduces resource use while maintaining high extraction accuracy. This confirms its suitability for resource-constrained UAV platforms. For end-to-end transmission, the semantic index vector transmission results (Fig. 5) show that the Bit Error Rate (BER) decreases monotonically as the Signal-to-Noise Ratio (SNR) increases in all three channel environments. The rural environment achieves the best performance, followed by the suburban and urban environments. These differences are mainly caused by variations in scatterer density and link blockage across environments. The proposed transmission network maintains stable BER under different Doppler frequencies, demonstrating its robustness under dynamic channel conditions. For semantic image transmission, the proposed weighted hybrid loss function shows good training stability (Fig. 6), and LPUSC consistently outperforms the Deep Joint Source-Channel Coding (DeepJSCC) and JPEG + Low-Density Parity-Check (LDPC) baselines across the full SNR range (Fig. 7). Specifically, LPUSC achieves SSIM and Peak Signal-to-Noise Ratio (PSNR) gains of 1.3% and 4.8% over DeepJSCC, respectively, and gains of 43% and 79.5% over JPEG + LDPC, respectively. These results indicate that the proposed personalized semantic image transmission network achieves high-quality reconstruction and remains robust to channel variations.  Conclusions  To improve the efficiency and flexibility of UAV image communication, this paper proposes LPUSC, a lightweight personalized semantic communication system. The system uses a dual-branch transmission architecture that integrates lightweight, high-precision object detection and semantic segmentation models. It enables personalized content transmission without interest pre-training. This design satisfies personalized user requirements while maintaining low computational and communication overhead. Simulation results show that the LPUSC system achieves stable and reliable semantic index interaction and outperforms the DeepJSCC and JPEG + LDPC baselines in semantic region reconstruction. The proposed system provides a useful reference for efficient UAV image semantic communication in 6G low-altitude intelligent networks.
Research on Covert Communication Transmission Scheme Combining Relay Selection and Mode Selection over Nakagami-m Fading Channels
HUANG Haiyan, HUANG Yi, ZHANG Ning, LIANG Linlin, ZHANG Xuejun
Available online  , doi: 10.11999/JEIT260287
Abstract:
  Objective  Covert communication enhances the security of wireless communication systems by concealing both transmitted information and communication activities from unauthorized detection. However, practical wireless channels exhibit random and uncertain propagation conditions. The Nakagami-m fading channel, which can characterize a wide range of channel conditions, provides a realistic framework for evaluating the performance of covert communication. Relay-assisted transmission has attracted considerable attention because it improves transmission reliability over fading channels. Moreover, relay selection and transmission mode selection substantially affect system performance. Therefore, investigating their combined effect on covert communication over Nakagami-m fading channels is of both theoretical and practical significance for the design of next-generation secure wireless communication systems.  Methods  This paper proposes a covert communication system incorporating relay selection and transmission mode selection. The source node transmits covert information to the destination node through multiple relays, while a warden monitors transmissions from both the source and relay nodes. A friendly jammer transmits interference signals to degrade the warden’s detection capability. Four transmission schemes are considered: optimal relay selection with fixed Half-Duplex (HD) or Full-Duplex (FD) operation, optimal relay selection with random transmission mode selection, random relay selection with optimal transmission mode selection, and joint optimal relay and transmission mode selection. Closed-form expressions for the warden’s detection error probability under both HD and FD optimal relay selection are derived over Nakagami-m fading channels. Closed-form expressions for the transmission outage probability, asymptotic transmission outage probability, and covert rate are also derived for all transmission schemes. The theoretical analysis is validated through MATLAB simulations.  Results and Discussions  Simulation results demonstrate that an optimal detection threshold exists that minimizes the detection error probability (Fig. 2). As the detection threshold or jamming power increases, the warden’s ability to detect covert communication decreases, causing the detection error probability to approach one (Figs. 2 and 3). Under the same target transmission rate and high Signal-to-Noise Ratio (SNR) conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability, thereby providing the highest transmission reliability (Figs. 4 and 5). At a target transmission rate of \begin{document}$ \text{6.5 bit/(s}\cdot \text{Hz)} $\end{document}, the transmission outage probability of the joint relay and transmission mode selection scheme is 6.9% lower than that of the FD transmission scheme (Fig. 4). As the transmit power and the number of relays increase, the covert rate gradually approaches a constant value. Among all transmission schemes, the joint relay and transmission mode selection scheme consistently achieves the highest covert rate (Figs. 6 and 7).  Conclusions  This paper proposes a covert communication system based on relay selection and transmission mode selection over Nakagami-m fading channels. Closed-form expressions for the warden’s detection error probability and the system’s transmission outage probability are derived under different relay selection and transmission mode selection strategies. The asymptotic transmission outage probability and covert rate are then analyzed. Simulation results show that increasing the detection threshold or jamming power weakens the warden’s ability to detect covert communication, causing the detection error probability to approach one. Under identical target transmission rates and high SNR conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability. These results indicate that appropriate relay selection and transmission mode selection not only reduce the warden’s detection capability and protect covert communication, but also improve both transmission reliability and covertness. Future work will consider practical factors, including imperfect channel state information, residual self-interference, and incomplete knowledge of the warden’s channel.
Design of Lightweight Gated Recurrent Unit Network Model Based on Memristor
HUA Honghu, XU Jia, ZHANG Bohao, WANG Wei, LI Zhiwei, LIU Haijun
Available online  , doi: 10.11999/JEIT260152
Abstract:
  Objective  With the slowdown of Complementary Metal-Oxide-Semiconductor (CMOS) technology scaling and the inherent memory-computation separation of von Neumann architectures, conventional computing systems face increasing challenges in processing large-scale sequential data. Memristors provide a promising solution because of their high integration density, fast switching speed, and synaptic plasticity. Memristor crossbar arrays naturally support Vector-Matrix Multiplication (VMM) in the analog domain, enabling energy-efficient in-memory computing. As a representative recurrent neural network, the Gated Recurrent Unit (GRU) has achieved excellent performance in sequential tasks such as trajectory prediction and urban sound classification. However, conventional hardware implementations of GRU networks require frequent data transfer between memory and processing units, resulting in high energy consumption and limited throughput. Although memristor-based GRU implementations improve computational efficiency, their large parameter size and high weight precision require substantial hardware resources and reduce deployment reliability on resource-constrained memristor arrays. In addition, device non-idealities, such as conductance fluctuations, further reduce inference accuracy. Existing memristor-based GRU methods generally treat weights and activations using the same quantization strategy without considering their different hardware implementation characteristics, and they provide limited robustness against device variations. This paper addresses these issues through a hardware-algorithm co-design strategy.  Methods  This paper proposes a lightweight memristor-based GRU network model. A 1T1R (one-transistor-one-resistor) memristor crossbar array is adopted for weight mapping and analog Multiply-Accumulate (MAC) operations. Signed weights are represented by differential pairs of positive and negative conductance matrices because memristor conductance values are inherently non-negative. A linear transformation is used to map trained network weights to memristor conductance values. To account for the different hardware implementation paths of weights and activations, a device-aware fusion quantization method based on performance analysis is proposed. Symmetric quantization is applied to weights stored in the memristor array because the zero-centered quantization range eliminates zero-point storage and simplifies write-driver circuit design. In contrast, asymmetric quantization is applied to activation values computed in peripheral circuits, thereby preserving the dynamic range and reducing quantization error. To improve robustness against memristor conductance fluctuations, weight noise training is incorporated into Quantization-Aware Training (QAT). Gaussian noise with an intensity determined by the device variation parameter is injected into quantized weights during each forward pass. This strategy acts as a regularizer that guides the model toward flatter loss minima and improves tolerance to weight perturbations. During backpropagation, the straight-through estimator updates the full-precision floating-point weights, whereas noise is dynamically resampled in every forward pass.  Results and Discussions  On the public UrbanSound8K dataset, the proposed full-precision lightweight memristor-based GRU network model achieves a classification accuracy of 93.94%. After applying the device-aware fusion quantization method, the 6-bit quantized model achieves 92.68% accuracy, corresponding to only a 1.26% decrease while reducing weight precision by 81.25% (Table 1). The proposed model outperforms Dilated Convolution (78.00%), LM-MFCC+GRU (92.00%), TFFS-DNN (88.74%), TFCNN (93.10%), and CL-Transformer (92.95%) under their full-precision settings (Table 2). Under noisy input conditions with Signal-to-Noise Ratios (SNRs) ranging from −10 dB to 10 dB, the 6-bit quantized model exhibits robustness comparable to or better than that of the full-precision model, demonstrating the effectiveness of the proposed device-aware fusion quantization strategy (Table 3). From the perspectives of storage, hardware resources, and device feasibility, 6-bit quantization reduces weight storage from 5.6 MB to 1.05 MB, corresponding to a compression ratio of 81.2%, while requiring only 2.8 million memristor cells under the 1T1R mapping scheme. Weight noise training also substantially improves robustness against device non-idealities. When the conductance variation reaches 14%, the classification accuracy increases from 82.97% to 91.14%. At the maximum simulated variation of 28%, the accuracy increases from 54.23% to 87.01% (Fig. 7), demonstrating improved tolerance to memristor device variations. On a self-constructed true-false trajectory dataset, the lightweight memristor-based GRU network model achieves 97.35% accuracy at full precision and 96.51% after 6-bit quantization, with only a 0.84% decrease, outperforming the Dilated Convolution baseline (Table 4). To further verify its applicability to different sequential tasks, the lightweight memristor-based GRU network model is evaluated on lithium-ion battery State-of-Charge (SOC) estimation using a public dataset. The 6-bit quantized model achieves Root Mean Square Errors (RMSEs) of 1.48%, 0.79%, and 0.74% at 0 °C, 25 °C, and 45 °C, respectively, outperforming the existing memristor-based GRU implementation. The proposed model also achieves lower RMSEs than the comparison method at all evaluated quantization precisions of 6 bits and above (Table 5).  Conclusions  This paper presents a lightweight memristor-based GRU network model for hardware deployment. By combining device-aware fusion quantization with weight noise training integrated into Quantization-Aware Training (QAT), the model achieves substantial memory compression while maintaining high classification accuracy and improving robustness to memristor device non-idealities. Experimental results on multiple datasets and sequential tasks demonstrate that the 6-bit quantized model preserves competitive accuracy and stable performance, providing an effective solution for deploying GRU networks on resource-constrained memristor-based edge computing platforms.
A Survey of Cooperative Mission Planning for Imaging Satellites Observing Moving Targets
XU Zhuo, FAN Shenghua, YUE Haitao, QU Tao, WANG Dingwen, SUN Shilei
Available online  , doi: 10.11999/JEIT260133
Abstract:
  Significance   Cooperative mission planning for imaging satellites observing moving targets is a key technique that supports the transition of space-based Earth observation systems from static regional coverage to a dynamic closed-loop paradigm consisting of wide-area search, dynamic tracking, and feedback-guided supplementary search. It plays an important role in emergency response, maritime monitoring, wide-area situational awareness, and persistent observation of high-value moving targets. The primary challenge arises from the conflict between uncertainty in future target states and the reliance of conventional mission planning models on deterministic inputs. Unlike static targets, moving targets have neither fixed locations nor fixed visibility windows. Their future states are generally represented by probability distributions, confidence regions, or grid-based target existence probabilities. Effective mission planning therefore requires not only accurate target motion prediction but also systematic integration of uncertainty into planning objectives, constraints, and replanning triggers. A comprehensive review from an uncertainty-driven perspective is therefore needed.  Progress   This survey reviews cooperative mission planning for imaging satellites observing moving targets from an uncertainty-driven perspective. Typical moving targets are classified into maritime moving targets, highly time-sensitive aerospace targets, and ground moving targets according to their operating environments, dynamic characteristics, and observation requirements. Although these target categories differ in maneuverability, prior constraints, and observation windows, they share a common planning challenge: coupling uncertain target motion with deterministic satellite observation actions under stringent platform and resource constraints. Methods for target motion prediction and spatiotemporal uncertainty representation are first reviewed. Physics-based methods characterize target state evolution using kinematic constraints, dynamic models, covariance propagation, reachable sets, and Markov state transition models. Data-driven methods learn motion patterns from historical trajectories, Automatic Identification System (AIS) data, remote sensing observations, meteorological information, and geographic constraints. From the perspective of mission planning, the utility of these methods depends on whether outputs such as covariance, target existence probability, confidence regions, and information gain can be directly incorporated into planning models. Observation task modeling, cooperative planning architectures, optimization algorithms, and closed-loop replanning mechanisms are then analyzed. Deterministic task models simplify uncertainty into trajectory points, visibility windows, or fixed geographic regions, while probabilistic task models incorporate target existence probability, belief states, and information gain into objective functions, constraints, and state transition models. Centralized, distributed, and hybrid planning architectures are compared with respect to global optimization capability, onboard autonomy, communication overhead, and response timeliness. Exact optimization methods, heuristic methods, metaheuristic algorithms, Deep Reinforcement Learning (DRL), and Large Language Model (LLM)-assisted solution strategies and algorithm design are also reviewed. Finally, state-triggered replanning, Receding Horizon Optimization (RHO), and Model Predictive Control (MPC) are summarized as representative approaches for closed-loop dynamic scheduling.  Conclusions  The reviewed studies indicate that cooperative mission planning is evolving from open-loop static scheduling to closed-loop dynamic planning. Nevertheless, several challenges remain. First, uncertainty information generated during target prediction is not fully exploited in planning decisions. Rich probabilistic information is frequently reduced to deterministic time windows, discrete trajectory points, or geometric regions, thereby limiting risk-aware task allocation. Second, distributed cooperation lacks reliable belief-state consistency. Differences in local observations may lead satellites to maintain inconsistent estimates of the same target state, resulting in redundant observations, task conflicts, and inefficient resource utilization. Third, dynamic replanning lacks unified benefit-cost criteria for determining replanning triggers. Excessively frequent replanning increases attitude maneuver time, energy consumption, and onboard storage resource use, whereas delayed replanning may fail to respond to actual target maneuvers. Fourth, LLMs have demonstrated potential for task requirement parsing, constraint modeling, heuristic generation, and algorithm design for satellite scheduling, but their application to cooperative mission planning for moving targets remains limited.  Prospects   Future research should focus on developing a more robust closed-loop planning framework. Prediction uncertainty should be incorporated directly into planning models through chance-constrained planning, belief-state planning, or Partially Observable Markov Decision Processes (POMDPs), enabling covariance, target existence probability, and information entropy to be integrated into planning objectives, constraints, and replanning triggers. Bayesian updating or sequential Bayesian filtering should use both successful detections and missed detections to continuously refine the prediction layer. Distributed cooperation requires lightweight state synchronization and belief fusion supported by compact state-sharing descriptors and event-triggered communication. Replanning decisions should be guided by information gain and benefit-cost evaluation. In addition, LLMs should be developed as verifiable auxiliary tools rather than direct replacements for optimization solvers. They can assist with task requirement structuring, constraint modeling, heuristic generation, and algorithm component design, whereas feasibility verification, solution refinement, and performance evaluation should remain the responsibility of formal verification methods, conventional optimization algorithms, and simulation environments. These research directions are expected to improve the robustness and uncertainty awareness of mission planning for satellite observation of moving targets.
Heterogeneous Task Cooperative Scheduling Architecture for Networked Radar in Saturation Attack Air Defense Early Warning
YE Juhang, FANG Yuyuan, WEI Shaopeng, DUAN Jia, ZHANG Lei
Available online  , doi: 10.11999/JEIT260373
Abstract:
  Objective  To address the severe challenges posed by Unmanned Aerial Vehicle (UAV) swarms and intelligent loitering munitions, which generate massive, sudden, and heterogeneous early warning tasks during saturation attacks, a scalable networked-radar cooperative scheduling architecture is proposed. Existing architectures suffer from rigid dynamic coordination and insufficient capability for heterogeneous task scheduling. The non-convex cooperative scheduling problem is therefore decoupled into a multi-stage decision process consisting of multidimensional dynamic resource coordination and adaptive heterogeneous task scheduling. By incorporating a dispatch mechanism and Hierarchical Reinforcement Learning (HRL), a hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling is developed. The architecture is implemented through an execution-redispatch cognitive closed loop, together with a Target Dispatch Algorithm (TDA) and a hierarchical command-and-scheduling method, to address multi-radar, multi-target, and multi-task scheduling in saturation attack scenarios.  Methods  The cooperative scheduling problem is first decoupled into dispatch-oriented network-level multidimensional resource coordination and hierarchical-command-based node-level adaptive heterogeneous task scheduling. An environment perception layer establishes a dynamic uncertainty model based on the Bayesian Cramér-Rao Lower Bound (BCRLB) and radar detection probability to jointly characterize target threat levels and radar operating states. The network coordination layer then adopts an adaptive weighted dispatch model based on comprehensive combat effectiveness to achieve dynamic target allocation while constructing a scalable execution-redispatch cognitive closed loop. To solve the resulting large-scale, nonlinear, multi-constraint generalized bipartite graph matching problem, the proposed TDA, an MMAS-based algorithm incorporating feasibility-rule-based constraint handling, is developed as a constructive solution. At the node level, Hierarchical Q-Learning (HQL) is implemented through a serial dual-Q-table implementation for distributed heterogeneous task scheduling. Using task proportion control as the hierarchical subgoal, the proposed method transforms upper-level operational intent into lower-level beam dwell scheduling, enabling adaptive, long-term, and interpretable execution of heterogeneous tasks with complex dependency relationships.  Results and Discussions  A simulated point-defense scenario against UAV swarm and loitering munition saturation attacks is established using three networked homogeneous S-band medium-range phased-array radars to counter 200 high-speed maneuvering targets. For network-level coordination, the proposed TDA replaces conventional penalty functions with a hierarchical solution framework based on feasibility-rule constraint handling. By exploiting prior model information, TDA achieves higher solution quality and faster convergence than the Max-Min Ant System (MMAS), Artificial Bee Colony (ABC), and Genetic Algorithm (GA) (Fig. 3). Although computational complexity increases, the execution time remains well within the dispatch cycle, improving solution quality with only millisecond-level computational overhead while satisfying real-time operational requirements (Fig. 4). For node-level scheduling, HQL employs hierarchical macro- and micro-level decisions to ensure policy consistency. Supported by an internal dense transfer-reward mechanism, HQL achieves higher learning efficiency and better long-term policy quality than Q-Learning (QL) and the Priority-Based Method (PBM) (Fig. 5). The hierarchical serial dual-Q-table framework maintains balanced task proportions, maximizes comprehensive combat effectiveness, and improves resource utilization (Fig. 6). Furthermore, comparisons of the target track-loss rate and mean tracking error show that HQL achieves the lowest mean tracking error by prioritizing high-quality tracking tasks, despite a moderately higher target track-loss rate, demonstrating superior long-term scheduling capability (Fig. 7).  Conclusions  The proposed hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling effectively addresses the limitations of rigid dynamic coordination and insufficient heterogeneous task scheduling capability in existing networked radar systems. The proposed framework enables real-time cooperative scheduling of search, confirmation, and tracking tasks during large-scale, high-speed saturation attacks, thereby improving the operational capability of the air defense early warning system. Simulation results demonstrate improvements in applicable processing scale, environmental adaptability, long-term scheduling capability, scalability, and interpretability. Future work will explore online learning to reduce the discrepancy between offline training and online deployment and will further extend the architecture by integrating weapon-target assignment to support unified early warning and fire-control systems.
Analysis of Age upon Decisions and Distortion at Decisions in IoT Status Update Systems with Batch Arrivals
LIU Lei, JIN Wenkai, ZHANG Qingqing, LI Yuzhou, JIANG Fan
Available online  , doi: 10.11999/JEIT260359
Abstract:
  Objective  The rapid development of the Internet of Things (IoT) makes the timely transmission and processing of status updates essential for modern systems, where information freshness at decision epochs plays a critical role. In many IoT applications, such as smart grid fault detection and Industrial Internet of Things (IIoT) cluster monitoring, status updates typically arrive in batches rather than individually. However, most existing studies on Age of Information (AoI) assume single-update arrivals and therefore cannot accurately characterize the queueing dynamics caused by batch arrivals. Besides information freshness, distortion at decision epochs is another key factor because it directly affects decision quality. A fundamental tradeoff therefore exists between information freshness and distortion. Waiting for more complete status update information allows more completed status updates to be incorporated into joint estimation, but increases queueing and transmission delays, thereby reducing information freshness. In contrast, triggering decisions earlier reduces delay but increases distortion because fewer completed status updates are available for joint estimation. To address this problem, this paper investigates the tradeoff between information freshness and distortion in an IoT status update system with batch arrivals by adopting Age upon Decisions (AuD) and Distortion at Decisions (DaD) as performance metrics. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. Furthermore, for the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. The results provide theoretical insight and practical guidance for the design of IoT status update systems with batch arrivals.  Methods  Information freshness and distortion at decision epochs are analyzed for an IoT status update system with batch arrivals. AuD and DaD are adopted to quantify information freshness and distortion, respectively. Based on queueing theory, analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. A typical case with geometrically distributed batch sizes is then investigated. An alternating iterative optimization algorithm is further developed to jointly optimize the batch arrival rate, average batch size, and decision threshold to minimize the weighted sum of the average AuD and average DaD.  Results and Discussions  Simulation results validate the theoretical analysis. The average AuD exhibits a nonmonotonic trend as the batch arrival rate increases, first decreasing and then increasing. In addition, the Batch-size Coefficient Of Variation (BCOV) has a significant effect on the average AuD, with a smaller BCOV providing better information freshness performance. Under high-load conditions, queue backlogs become more severe, and stochastic fluctuations in batch arrivals have a greater effect on the queueing process. This increases service-time variability and amplifies the effect of BCOV on the average AuD (Fig. 2). As the average batch size increases, the system queue length and queueing delay increase, leading to a larger average AuD. At the same time, the decision control unit can utilize more completed status updates for joint estimation, thereby reducing the average DaD (Fig. 3). Moreover, the average DaD decreases as the decision threshold increases because more completed status updates are incorporated into the joint estimation process, improving estimation accuracy. A larger BCOV also increases the number of completed status updates available for joint estimation and therefore further reduces the average DaD (Fig. 4). The optimization results show that the solutions obtained by the proposed algorithm lie on the Pareto frontier, demonstrating its effectiveness. By comparison, fixed batch arrival rates and decision thresholds produce performance that is considerably farther from the Pareto frontier, demonstrating the advantage of jointly optimizing system parameters (Fig. 5).  Conclusions  This paper investigates an IoT status update system with batch arrivals by adopting AuD and DaD to quantify information freshness and distortion, respectively. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. For the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. Simulation results verify the theoretical analysis and reveal the effects of the batch arrival rate, average batch size, and decision threshold on the average AuD and average DaD. The results also demonstrate that the proposed low-complexity algorithm effectively identifies Pareto-optimal solutions for the AuD-DaD tradeoff. This study considers only the batch arrival characteristics of status updates. Future work will incorporate batch service mechanisms to further examine their effects on the tradeoff between AuD and DaD. Flexible decision mechanisms can also be developed to achieve adaptive AuD-DaD tradeoffs according to the heterogeneous requirements for information freshness and distortion across applications with different batch characteristics.
Kolmogorov-Arnold Nonlinear Enhancement Method for Aerial-Ground Person Re-IDentification
CHEN Yijun, ZENG Xianxian, LIU Shun, WANG Leijun
Available online  , doi: 10.11999/JEIT260430
Abstract:
  Objective  Aerial-Ground Person Re-IDentification (AG-PReID) aims to match the same person across Unmanned Aerial Vehicle (UAV) and ground-camera views. Compared with conventional same-platform person re-identification, this task faces larger cross-view appearance variation and more severe cross-domain distribution shifts. Under these conditions, identity-consistent cues are often weakened by strong viewpoint asymmetry and cross-domain appearance distortion. Existing methods mainly focus on feature extraction and cross-view representation alignment. However, the classification supervision branch still relies heavily on linear feature transformation, which limits its ability to model complex nonlinear discriminative relationships in high-dimensional feature spaces. A stronger nonlinear supervision mapping is therefore needed to better exploit high-order feature interactions and local discriminative variations. To address this issue, this paper proposes a Kolmogorov-Arnold Nonlinear Enhancement Module (KANEM). KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier. It uses learnable nonlinear mappings to adaptively enhance features for more discriminative cross-view representation learning.  Methods  The backbone follows the View-Decoupled Transformer (VDT), which introduces an additional view token and performs layer-wise view decoupling. This design separates view-related factors from identity features and reduces representation bias between aerial and ground domains. Based on this framework, KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier, thereby providing adaptive nonlinear mappings for feature enhancement. Specifically, KANEM consists of a base activation branch and a spline branch, which are stacked into cascaded function-mapping layers. This design enables more flexible nonlinear modeling than conventional linear or MultiLayer Perceptron (MLP)-based transformations. It allows the model to capture local nonlinear variations and complex correlations among feature dimensions. To improve discriminability and further separate identity and view information, the network is jointly optimized using identity classification loss, view classification loss, triplet loss, and orthogonality loss. KANEM is used only during training and is removed during inference, so no extra inference cost is introduced.  Results and Discussions  Comprehensive evaluations are conducted on the CARGO and AG-ReID datasets. The results show that the proposed method consistently performs better than the baseline model and existing state-of-the-art methods. On CARGO, the proposed method achieves 70.19%/63.16%/51.34% in Rank-1 accuracy, Mean Average Precision (mAP), and Mean Inverse Negative Penalty (mINP), respectively, under the overall ALL retrieval protocol. It also achieves 58.75%/53.27%/41.11% under the most challenging aerial-ground (A↔G) cross-view retrieval protocol (Table 1). On AG-ReID, the proposed method achieves the best performance under both retrieval protocols. It reaches 84.41%/76.21%/53.05% in Rank-1/mAP/mINP for aerial-to-ground (A→G) retrieval and 86.69%/77.99%/52.28% for ground-to-aerial (G→A) retrieval (Table 2). Ablation studies on CARGO further verify the effectiveness of KANEM. They show that KANEM achieves better overall performance than conventional linear transformation and MLP-based alternatives. This result indicates that the proposed nonlinear enhancement strategy is more suitable for supervision mapping in CARGO (Tables 3 and 4). In addition, integrating KANEM into other person re-identification tasks further demonstrates its potential generalization ability across different scenarios (Table 5). Parameter analysis shows that setting λ to 0.001 enables the model to better balance the complexity difference between view classification and identity classification (Fig. 2(a)). When G and P are set to 5 and 3, respectively, the model effectively fits nonlinear variations in the feature space while preserving the smoothness and continuity of spline functions. This setting achieves effective nonlinear feature enhancement (Fig. 2(b)(d)). The two-dimensional t-distributed Stochastic Neighbor Embedding (t-SNE) visualization shows that the enhanced features have higher intra-class compactness and better inter-class separability (Fig. 3). The top-5 retrieval comparisons further provide qualitative evidence that the proposed method improves ranking quality and retrieval robustness under all four retrieval protocols on CARGO. It promotes correct matches to higher positions and returns more relevant samples among the top-ranked results (Fig. 4).  Conclusions  This paper presents KANEM for AG-PReID. The proposed module is motivated by the large discrepancy between UAV and ground-camera views and by the limited capacity of linear feature transformation in the classification branch to capture complex nonlinear discriminative relationships. By replacing the conventional fully connected feature transformation between backbone output features and the linear classifier, KANEM provides a more flexible nonlinear supervision mechanism for cross-view representation learning. Through adaptive nonlinear enhancement, it better models complex feature interactions in high-dimensional spaces and strengthens the representation of cross-view consistency and fine-grained discriminative cues. Experimental results on CARGO and AG-ReID demonstrate the effectiveness of the proposed method, particularly in challenging scenarios with large view discrepancies. Future work will further refine the nonlinear mapping mechanism of KANEM and explore its use in more complex cross-view settings to improve model discriminability and generalization performance.
A Task Prediction-augmented Hierarchical Offloading Method for Space-Air-Ground Integrated Networks
ZHANG Linghao, XU Bo, SUN Jinlong, LAI Haiguang, ZHAO Haitao
Available online  , doi: 10.11999/JEIT260217
Abstract:
  Objective  Space-Air-Ground Integrated Networks (SAGIN) have become key infrastructure for future 6G communications. They support wide-area coverage and flexible deployment through the coordinated operation of Low Earth Orbit (LEO) satellites, Unmanned Aerial Vehicles (UAVs), and Ground Users (GUs). With the rapid growth of Internet of Things (IoT), Internet of Vehicles (IoV), and smart city applications, terminal devices generate increasingly diverse computation-intensive tasks. These tasks impose high requirements on real-time computing and resource scheduling. Mobile Edge Computing (MEC) has been integrated into SAGIN architectures to provide near-user computing services by using UAVs and satellites as edge nodes, thereby reducing task completion latency. However, efficient task offloading remains challenging when average task completion latency and UAV flight energy consumption must be jointly reduced. This difficulty is caused by the strong coupling among UAV trajectory planning, task offloading, and computational resource allocation. It is further intensified by the dynamic and partially observable nature of SAGIN environments. Existing Multi-Agent Reinforcement Learning (MARL) methods mainly rely on reactive decisions based on instantaneous observations. They lack awareness of future task workload changes, which leads to decision lag and limited adaptability under bursty traffic. To address these issues, a task prediction-augmented MARL method is proposed to support forward-looking decisions in dynamic SAGIN environments.  Methods  A three-layer SAGIN-MEC architecture is considered, including one LEO satellite, multiple UAVs, and GUs. Tasks can be processed locally, offloaded to UAVs through Ground-to-Air (G2A) links, or further relayed to the LEO satellite through Air-to-Satellite (A2S) links under a partial offloading mechanism. The joint optimization of UAV trajectory, user association, offloading ratios, and computational resource allocation is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem. The objective is to minimize the weighted sum of average task completion latency and UAV flight energy consumption. Owing to the nonconvexity and high dimensionality of this problem, it is reformulated as a DECentralized Partially Observable Markov Decision Process (DEC-POMDP). A Prediction-Augmented Multi-Agent Proximal Policy Optimization (PA-MAPPO) algorithm is then developed. A lightweight Exponential Smoothing-Autoregressive (ES-AR) prediction module is used to generate multi-step workload forecasts, which are incorporated into the state space of each agent. The algorithm adopts a bilevel structure. In the outer layer, Centralized Training and Decentralized Execution (CTDE)-based PA-MAPPO generates UAV trajectory actions. In the inner layer, Block Coordinate Descent (BCD)-based convex optimization solves the resource allocation and offloading subproblems, and closed-form resource allocation solutions are obtained through Lagrangian analysis. Generalized Advantage Estimation (GAE) and the PPO-Clip objective are used to improve training stability and convergence.  Results and Discussions  Simulations are conducted with one LEO satellite, five UAVs, and 50 GUs in a 1×1 km2 area. PA-MAPPO is compared with MAPPO without prediction and Prediction-Augmented Multi-Agent Deep Deterministic Policy Gradient (PA-MADDPG). The training curves show that PA-MAPPO converges within 500~700 episodes, with the highest average reward and the smallest variance, indicating better stability (Fig. 3). As the number of GUs increases from 20 to 80, PA-MAPPO consistently achieves the lowest system cost. Compared with MAPPO and PA-MADDPG, it reduces the average cost by approximately 12.4% and 18.7%, respectively (Fig. 4). Experiments with different UAV numbers show a U-shaped cost curve for all algorithms. The best configuration is obtained when U=5, where PA-MAPPO achieves the minimum cost (Fig. 5). Sensitivity analysis of the latency-energy tradeoff weight ω confirms that PA-MAPPO remains robust under different optimization preferences (Fig. 6). The prediction horizon H has a nonmonotonic effect on performance. When H=5, PA-MAPPO obtains the best result and reduces the cost by approximately 14.9% compared with the no-prediction case. Longer horizons degrade performance because prediction errors accumulate (Fig. 7).  Conclusions  The PA-MAPPO algorithm is proposed to solve the joint optimization of UAV trajectory planning, user association, task offloading, and computational resource allocation in dynamic SAGIN environments. By integrating a lightweight ES-AR task workload prediction module into the MARL process, PA-MAPPO enables UAV agents to account for future task dynamics. This design reduces the decision lag caused by purely reactive methods. The inner BCD-based convex optimization converges to a Karush-Kuhn-Tucker (KKT)-stationary point, while the outer CTDE-based PPO mechanism improves training stability and scalability. Simulation results show that PA-MAPPO outperforms baseline methods in average task completion latency, UAV flight energy consumption, and overall system cost. It also maintains strong scalability and robustness under different system configurations. Future work will study online prediction and decision co-optimization in multi-satellite cooperative scenarios and examine the effect of dynamic network topology changes on algorithm performance.
A Dual-polarized Magnetoelectric Dipole Antenna Array with Differential Feeding
TANG Li, WANG Zhihui, ZHAO Luyu
Available online  , doi: 10.11999/JEIT260505
Abstract:
  Objective  This study addresses key challenges in Fifth-Generation (5G) millimeter-wave terminal antennas by designing a compact, high-performance dual-polarized array. Existing designs often face trade-offs among bandwidth, beam-scanning range, and integration complexity. To address these limitations, this paper proposes a differentially fed magnetoelectric dipole array. A stacked stripline-slot-stripline balun is used to enable efficient single-ended-to-differential conversion, and the array design is optimized. The objective is to realize an integrated solution with wideband operation, low cross-polarization, wide-angle beam scanning, and high integration density for practical 5G millimeter-wave applications.  Methods  A structured design method is adopted. First, a stacked differential balun based on a stripline-slot-stripline configuration is developed to achieve efficient single-ended-to-differential conversion. A single-polarized magnetoelectric dipole antenna element is then designed and integrated with the balun, and its performance is characterized. The design is further extended by orthogonally integrating two elements to form a dual-polarized unit, which is used to construct a 1×4 linear array. Iterative full-wave electromagnetic simulation and optimization are conducted to balance wideband impedance matching, high port isolation, stable wide-angle beam scanning, grating-lobe suppression, and mutual-coupling reduction.  Results and Discussions  The optimized 1×4 dual-polarized differentially fed magnetoelectric dipole antenna array uses an element spacing of 4.6 mm, corresponding to 0.4 free-space wavelength at 26 GHz. This spacing achieves a favorable balance between grating-lobe suppression and inter-element mutual-coupling reduction. The measured –10 dB reflection coefficient bandwidths are 25~29.4 GHz for the +45° polarization port and 25~27.7 GHz for the –45° polarization port (Fig. 20). The slight matching difference is attributed to the incomplete structural symmetry of the baluns under the two polarization modes (Fig. 13). At 26 GHz, both polarization modes provide a peak gain of 10.7~11 dBi and support ±60° wide-angle beam scanning, with main-lobe gain attenuation no greater than 3 dB (Fig. 21). The measured radiation performance agrees well with the simulated results. Minor deviations are mainly caused by the high dimensional sensitivity of millimeter-wave structures and small errors in fabrication and test assembly. The array also maintains stable low cross-polarization and high port isolation across the operating band. These results are achieved through equal-length feed lines, symmetric layout, ground-pad shielding, and metallized-via electromagnetic isolation (Fig. 16), which suppress mutual coupling and parasitic radiation and ensure consistent dual-polarized radiation performance.  Conclusions  This paper presents a dual-polarized magnetoelectric dipole antenna array with differential feeding for 5G millimeter-wave applications. By using a stacked stripline-slot-stripline balun and optimizing the radiating structure and array layout, the design achieves wide bandwidth, high gain, low cross-polarization, and wide-angle beam scanning. The differential balun enables efficient single-ended-to-differential conversion with good amplitude and phase balance across the target band. The implemented 1×4 array, with an optimized element spacing of 4.6 mm, achieves a simulated peak gain of 11 dBi at 26 GHz and supports ±60° beam scanning, with gain variation below 3 dB. The overall design verifies the feasibility of a differentially fed magnetoelectric dipole architecture for compact, high-performance 5G millimeter-wave terminal antenna modules. Future work may focus on larger array configurations and further integration with BeamForming Integrated Circuits (BFICs).
An Incremental Density Clustering Method with Time-Difference Prior for Mobile Multi-Station Radar Signal Sorting
CHEN Jinli, FAN Yu, WANG Yanjie, ZHANG Jindong
Available online  , doi: 10.11999/JEIT260151
Abstract:
  Objective  In modern electronic warfare, radar signal sorting is a pivotal technology for electronic reconnaissance. Its primary objective is to deinterleave and categorize pulses from multiple radar emitters within dense and overlapping pulse streams. However, in practical reconnaissance missions, particularly those involving mobile platforms such as aircraft, the observation stations are in constant motion, while the target radar emitters typically remain stationary. This dynamic geometric relationship between the observation stations and the emitters causes the Time Difference of Arrival (TDOA) of intercepted signals to exhibit non-stationary characteristics that evolve over time. Conventional sorting algorithms lack the mechanism to perceive this dynamic feature drift. Consequently, continuous TDOA trajectories generated by the same emitter are prone to being incorrectly partitioned into multiple clusters during data processing, resulting in cluster proliferation and degraded sorting performance. To address these issues, an incremental density-based clustering method incorporating TDOA prior information is proposed. The proposed method effectively exploits the temporal evolution characteristics of TDOA, achieving high stability and high sorting accuracy in mobile multi-station scenarios.  Methods  First, by analyzing the temporal evolution characteristics of TDOA in mobile multi-station scenarios, the TDOA evolution process is approximated as a linear function of time, and a linear prior model is established to characterize its dynamic behavior. An online micro-cluster construction strategy is then adopted for multi-station TDOA samples, and macro-clusters are formed according to the spatiotemporal intersection relationships among micro-clusters. For each macro-cluster, independent Kalman filters are constructed for different TDOA dimensions. The TDOA value and its rate of change are defined as state variables, and recursive state prediction is performed to provide dynamic feature references for newly arriving samples. During sample association, a spatiotemporal joint scoring function is designed, in which prediction residuals generated by the Kalman filters are incorporated as dynamic constraints. Consequently, the matching criterion evolves from a conventional static density-based rule into a joint spatiotemporal consistency constraint, enabling more accurate assignment of newly arriving TDOA samples. To suppress cluster proliferation caused by TDOA feature evolution, a concept drift detection mechanism based on macro-cluster center evolution is further introduced. Once concept drift is detected, posterior state vectors and covariance matrices estimated by the Kalman filters are employed to construct a dual Mahalanobis-distance criterion that jointly evaluates state-distribution overlap and predicted-position overlap. Under a 95% confidence threshold, mis-split radar clusters are adaptively merged and assigned to the same radiation source, thereby effectively suppressing cluster proliferation induced by concept drift.  Results and Discussions  In the simulation experiments, radar pulse data distributions are generated according to the radar parameters (Table 1). The multi-station TDOA corresponding to the same radiation source exhibits a near-linear evolution trend with respect to the Time of Arrival (TOA) at the observation station(Fig. 5). To comprehensively evaluate the performance of the proposed method in complex environments, a simulation analysis on radar cluster proliferation suppression is first conducted (Fig. 6). In this scenario, radar emitters E4 and E9 are selected from the nine simulated radiation sources as representative cases for analysis. Compared with DBSCAN, Incremental Density-Based Clustering (ICDC), cloud model-based sorting, and PointNet++ sorting algorithms, the proposed method effectively suppresses radar cluster proliferation and demonstrates superior cluster stability during the sorting process. Furthermore, to further demonstrate the superiority of the proposed approach, its sorting performance was comprehensively compared with several representative algorithms, including the conventional histogram method, grid-based clustering, DBSCAN, ICDC algorithm, cloud model-based sorting, and PointNet++ sorting algorithms. The sorting accuracy of various algorithms under different TOA measurement errors (Fig. 7) indicates that the proposed method achieves a sorting accuracy exceeding 96% when the TOA measurement error ranges from 50 to 300 ns, demonstrating remarkable robustness to measurement noise. Under different pulse interference rates (Fig. 8), the sorting accuracy of the proposed method remains above 96%, exhibiting excellent interference suppression capability. To further assess the stability of different algorithms under mobile observation platforms, simulations were performed at different observation station velocities (Fig. 9). The proposed method consistently maintains a high sorting accuracy across all tested velocities. Even at relatively high observation station speeds, the sorting accuracy remains at approximately 94%, indicating that the proposed approach retains favorable stability and robustness under dynamic observation conditions. In addition, simulations on radar cluster proliferation and missed-cluster probabilities are conducted (Fig. 10), and the computational complexities of different algorithms are analyzed (Table 2). The results demonstrate that the proposed method provides a favorable balance among sorting accuracy, cluster proliferation suppression, missed-cluster control, and computational cost.  Conclusions  To address the sorting performance degradation caused by the temporal evolution of TDOA features in mobile multi-station scenarios, an incremental density-based clustering method incorporating TDOA prior information is proposed. By incorporating observation station motion information to construct a dynamic TDOA state model and integrating Kalman Filters to achieve recursive prediction of TDOA evolution, the proposed method effectively suppresses radar cluster proliferation and mis-splitting caused by concept drift. In addition, the spatiotemporal joint criterion and the adaptive cluster merging mechanism further enhance the robustness of the algorithm under complex and dynamic environments. Simulation results verify the effectiveness and stability of the proposed method in mobile multi-station cooperative reconnaissance scenarios. Future work will focus on real-time multi-parameter fusion-based sorting methods in complex electromagnetic environments. Furthermore, adaptive adjustment strategies for the micro-cluster spatial intersection threshold, process noise covariance matrix, and measurement noise variance will also be investigated.
Resilience-Aware Cooperative Mission Planning Algorithm for Multi-UAV Systems in Complex Dynamic Environments
ZHAO xuejian, XIE lulu, WANG enliang
Available online  , doi: 10.11999/JEIT260138
Abstract:
  Objective  This paper addresses the strongly coupled problem of task allocation and route planning in cooperative mission planning for multiple UAV systems under complex dynamic environments, where dynamic task arrivals, UAV failures, no fly zone constraints, and link quality degradation coexist.   Methods  A resilience aware hybrid swarm optimization algorithm, termed RAHSO, is proposed. The method first builds an integrated task and route planning model that incorporates task value, route cost, energy consumption, interference penalty, time window constraints, capability constraints, conflict resolution, and link quality. It then combines clustering and genetic initialization, DBO and PSO hybrid search, GA and VNS local refinement, and Tarjan based deadlock detection and repair to obtain high quality baseline solutions. Furthermore, an event driven PPO online replanning module is introduced to rapidly adjust affected task subsets when emergent tasks, UAV failures, or topology changes occur.   Results and Discussions  Comparative and ablation experiments are conducted under static, scale-expansion, dynamic-event, and interruption scenarios. The results demonstrate that the proposed method achieves consistent advantages in task completion rate, total utility, average energy consumption, recovery time, and resilience index over representative baselines while maintaining acceptable replanning latency.   Conclusions  The proposed method provides an effective solution for resilient cooperative mission and route planning of multiple UAV systems in complex dynamic environments.
Survey of Satellite Covert Communications: Status, Key Technologies, and Future Challenges
DENG Hao, SUN Weiyuan, ZHU Zhengyu, PAN Gaofeng, SUN Gangcan
Available online  , doi: 10.11999/JEIT260177
Abstract:
  Objective  This survey comprehensively integrates the theoretical foundations, key technologies, and future challenges in the field of satellite covert communications. Based on the classic Alice-Bob-Willie model, it analyzes the impact of satellite channel characteristics on covert communication capacity, laying the theoretical groundwork. It summarizes the covert communication network model under the space-based, air-based, and ground-based three-layer architecture (Fig. 1), and systematically reviews core technologies and their optimization methods, including signal camouflage coding, beamforming, spectrum diversity, quantum encryption, and AI-assisted techniques. The main security threats faced by satellite covert communications are outlined, and multi-layered defense strategies, such as physical layer security and intelligent collaborative protection, are summarized. This provides theoretical and technical support for promoting highly secure and intelligent development in this field.  Significance   The research significance of this survey lies in its systematic integration of the theoretical framework and technological systems in the field of satellite covert communications. Addressing the threats of detection, interference, and eavesdropping in the vast, open satellite environment, it summarizes representative space-air-ground architectures reported in the literature. These architectures overcome the limitation of traditional encryption technologies that only protect information content, supporting low probability of detection by reducing statistical distinguishability at the physical layer. By elucidating the constraints of satellite channels on covert capacity through the refined square root law, and reviewing enhancement strategies reported in prior work centered on core technologies such as dynamic encoding, beamforming, and spectrum diversity, it provides theoretical and technical foundations for constructing highly survivable space-air-ground integrated secure communication networks. This holds significant strategic value for national defense, emergency communications, and the security assurance of 6G integrated space-terrestrial networks.  Progress   Existing studies reveal the dual impact of Doppler spread on satellite covert capacity: while increasing the missed detection probability, it simultaneously causes signal distortion, necessitating reliance on adaptive coding for compensation. The research further quantifies the detection characteristic differences among terrestrial, aerial, and orbital wardens (Willie) (Table 1), providing a theoretical basis for hierarchical defense design. In terms of covertness enhancement techniques, existing schemes propose multi-level strategies. AI-driven dynamic camouflage combined with sparse coding integrates background noise, inter-satellite links, and dynamic beamforming, improving covert throughput. At the network level, cooperative UAV-assisted transmission and dynamic spectrum coordination are summarized as typical network-level enhancement schemes (Fig. 3). Based on representative studies, a hierarchical defense framework for satellite covert communications is summarized. This combines reconfigurable intelligent surface control, Stackelberg game-theoretic incentives for jamming cooperation, and XOR network coding, and utilizes federated learning to achieve cross-domain threat signature sharing. These systematic advances provide innovative solutions for the covertness and security of satellite communications.  Conclusions  This paper systematically investigates the foundations and advancements of satellite covert communication, highlighting the integration of multi-layer satellite constellations, dynamic aerial relays, and quantum-encrypted, software-defined networks to establish resilient global covert channels. By adapting the Alice-Bob-Willie model to real-world satellite channel imperfections, it guides covert throughput and security optimization. The infusion of AI into coding and waveform design enables adaptive, environment-aware concealment strategies. Future research should focus on robust covert links in dynamic LEO environments, scalable constellation management, and the deep integration of AI and quantum technologies for 6G NTN systems, as the convergence of programmable satellites, intelligent surfaces, and advanced machine learning is expected to further influence secure space communications.  Prospects   Future research challenges and development trajectories focus on four critical domains: robust transmission under non-ideal channels, AI-enabled intelligent decision-making, 6G NTN integrated networking, and quantum-communication integration (Fig. 4). High-precision Doppler compensation models must be developed to mitigate rapid channel variations in LEO satellites. Concurrently, robust transmission mechanisms should be developed under non-ideal CSI conditions, potentially leveraging deep reinforcement learning for real-time resource optimization. Research should prioritize synergistic advancement of AI and quantum technologies. This entails integrating cross-layer Quantum Key Distribution (QKD) designs with covert transmission protocols, while utilizing Software-Defined Satellite (SDS) capabilities for dynamic strategy deployment. Key opportunities include exploiting Reconfigurable Intelligent Surfaces (RIS) for enhanced spatial-domain signal control and implementing blockchain solutions to address trust constraints in multi-node cooperative networks. Essential objectives encompass deploying efficient lightweight onboard algorithms and establishing optimized international coordination frameworks for spectrum and orbital resource allocation.
Cross-Domain Collaborative Enhancement for Tiny Object Detection in Remote Sensing Images
ZHANG Tianyang, ZHANG Xiangrong, WANG Guanchun, TANG Xu
Available online  , doi: 10.11999/JEIT260317
Abstract:
  Objective  The rapid development of deep learning has significantly advanced object detection in remote sensing images (RSIs).Due to constraints imposed by imaging conditions and object size, a large number of tiny objects (with a pixel area of less than 16×16) are widely distributed in RSIs.However, compared with satisfactory detection performance on normal-scaled objects, current object detection methods exhibit a significant performance gap for tiny objects.This is primarily attributed to the two critical limitations: insufficient positive label assignment and weak feature representations.To address these issues, a novel Cross-Domain Collaborative Enhancement Detector (CDCEDet) is proposed.The CDCEDet jointly optimizes the label assignment in the spatial domain and strengthens the feature representations in the frequency domain for tiny objects, thereby establishing an accurate and robust framework for remote sensing tiny object detection.  Methods  The overall framework of the proposed CDCEDet is depicted in Fig.2 and comprises three key components.First, a Scale-Adaptive Anchor Generator (SAAG) is devised to dynamically generate anchors matched with the scales of ground-truth (GT) objects, effectively mitigating the scale mismatch issue neglected in previous works.Compared to conventional uniform-scale anchor generators, the proposed SAAG significantly increases positive samples for tiny objects, even under the IoU-threshold-based label assignment.Second, a Quantile-based Adaptive Label Assignment (QALA) mechanism is proposed to replace the fixed IoU threshold label assignment.The QALA exploits the IoU distribution between each GT object and its matched anchors to dynamically generate an adaptive threshold tailored to each GT object, thereby further boosting positive samples for tiny objects.Third, a Frequency-Adaptive Fusion (FAF) module is constructed to enhance feature representations from a frequency-domain perspective.On the one hand, adaptive high-pass filters are employed to highlight high-frequency information, which mitigates the information loss during channel compression.On the other hand, adaptive low-pass filters are applied to smooth and up-sample high-level features, thereby alleviating the feature inconsistency within the up-sampled objects.  Results and Discussions  Extensive experiments are conducted on two publicly available remote sensing tiny object detection datasets, AI-TODv2 and AI-TOD-R, by comparing the proposed method with several latest approaches, including RFLA, DCNet, and DFCL.On the AI-TODv2 dataset (Table 1), the proposed CDCEDet achieves performance gains of 1.8% and 0.7% in terms of AP50 and AP50-95, respectively, compared to the latest method.On the AI-TOD-R dataset (Table 2), AP50 and AP50-95 are improved by 2.6% and 0.7%, respectively.These results demonstrate that the proposed CDCEDet possesses superior performance and robust generalization capabilities for tiny object detection in RSIs.Ablation studies and configuration analyses of the CDCEDet core modules (Table 3, Table 4, Table 5, and Table 6) further verify the effectiveness of each proposed component and their complementary nature.Qualitative evaluations on both datasets (Fig.3) confirm that the proposed method accurately detects tiny objects in both sparse and dense distributions scenarios.Fig.4 confirms that the proposed SAAG generates scale-matched anchors for each object and assigns more positive samples to tiny objects than widely used uniformly distributed anchor generator.Furthermore, in comparative visualizations with RFLA and DCNet (Fig.5), CDCEDet exhibits superior detection accuracy and demonstrates a more effective reduction in missed detections.  Conclusions  This paper proposes a novel CDCEDet to address the insufficient label assignment and weak feature representations in remote sensing tiny object detection.Specifically, a SAAG is proposed to generate scale-matched anchors tailored to each GT object, significantly increasing the number of positive samples for tiny objects.Additionally, a QALA mechanism is devised that exploits the IoU distribution between GT objects and their matched anchors to dynamically adjust the assignment threshold, thereby effectively mitigating the scale bias induced by the fixed threshold.Finally, a FAF module is constructed from a frequency-domain perspective, which adopts the adaptive high-pass filters and low-pass filters to strengthen the feature representations for tiny objects.Extensive experiments on two tiny object detection benchmarks demonstrate the superior performance and robust generalization of the proposed CDCEDet.Future research will focus on enhancing model efficiency and real-time performance to facilitate the practical deployment of the proposed approach in remote sensing scenarios.
Infrared Small Target Detection Enhanced by Multi-dimensional Fusion Attention
LI Weixing, WANG Shuai, CHEN Huaiyu, SHENG Weidong
Available online  , doi: 10.11999/JEIT260040
Abstract:
  Objective  Infrared imaging boasts advantages such as long operating range, wide coverage, strong concealment, and all-weather visibility, making it widely applicable in aerospace target surveillance, maritime emergency rescue, forest fire monitoring, and earth remote sensing. Infrared payloads are typically mounted on platforms like satellites and aircraft, capturing images over long distances where targets appear small in the imagery and lack distinct texture and morphological features. Due to weak thermal radiation signals, targets are prone to being submerged in strong clutter. Current infrared small target detection networks face several key challenges. First, target segmentation networks expand the receptive field through continuous down-sampling, which can cause the features of infrared small targets to be easily lost in the deeper layers of neural networks. Second, features in the deeper neural network layers tend to diffuse. Therefore, under the constraints of long-range imaging and complex backgrounds, achieving high-precision and highly reliable infrared small target detection remains a research hotspot in the field of infrared imaging.  Methods  A channel-spatial Multi-Dimensional Fusion Attention Mechanism (MFAM) is proposed to address the challenges of small target size and feature diffusion in deep networks. Multi-dimensional features across channels, height, and width are captured, as well as the dependencies across these dimensions. By integrating feature fusion and cross-dimensional synergistic interaction, feature diffusion of small targets in deep networks is mitigated and the robustness of infrared small-target detection is enhanced. The channel attention and spatial attention modules are applied in parallel directly to the input feature maps, performing feature extraction and local information interaction along the channel and spatial dimensions, respectively. In the channel domain, features are compressed and processed through a Multi-layer Perceptron (MLP) to capture inter-channel relationships. In the spatial domain, Global Average Pooling (GAP) and Global Maximum Pooling (GMP) are employed to encode width and height dimensions. Finally, feature fusion is achieved via a Sigmoid activation function. Compared with traditional hybrid attention mechanisms, the channel attention module and spatial attention module are applied in parallel directly to the input feature maps. This approach enables refined encoding of the channel, height, and width dimensions under a global receptive field. Meanwhile, shallow features with spatial details and deep semantic features rich in contextual information are captured. The proposed MFAM features a plug-and-play design, allowing flexible integration into various baseline models such as ResNet and DNA-Net without introducing complex additional structures, which demonstrates excellent compatibility and generalization capability.  Results and Discussions  The publicly available dataset NUDT-SIRST is utilized to conduct a performance analysis of the proposed algorithm. By incorporating the proposed MFAM into the baseline DNA-Net, the Intersection over Union (IoU), detection rate (Pd), and false alarm rate (Fa) are 87.34%, 98.72% and 3.22×10–6, respectively. Compared to CBAM, BAM, GAM, CA, and TA, MFAM improves IoU by 0.4%, 1.68%, 2.46%, 2.11%, and 1.61%, respectively(Table.1). CBAM places the spatial attention module in series of the channel attention, which may lead to information loss during feature propagation. CA and BAM only employ GAP for information encoding, overlooking the role of max pooling in deep feature extraction. TA realizes cross-dimensional dependencies through rotation operations but fails to achieve simultaneous three-dimensional interaction. MFAM simultaneously integrates channel and spatial information, enabling sufficient cross-domain interaction, and combines global average and max pooling to enhance context and detailed feature extraction capabilities, thereby achieving more refined infrared small target detection. MFAM is also embedded into ALC-Net and AMFU-Net, resulting in IoU improvements of 0.43% and 0.28% compared with CBAM(Table.2). Through ablation experiments, detection performance is compared under different attention combination strategies. Compared to the serial fusion approach, the proposed parallel structure achieves improvements of 0.37% and 0.39% in IoU and Pd, respectively, while reducing Fa by 1.41×10–6. The advantage stems from the parallel structure applying channel and spatial attention directly to the input features, mitigating the diffusion of target features in deep networks. To validate the inference performance of the proposed algorithm on edge-side processors, a verification system is designed by using FPGA and NVIDIA Jetson AGX Xavier. Practical testing confirms that the proposed algorithm can be deployed and inferred on edge-side GPUs, with an average single-frame inference latency of 46.7 ms (Fig. 8) for 256×256 input images.  Conclusions  A Multi-dimensional Fusion Attention Module (MFAM) is constructed in this paper. The MFAM module achieves effective aggregation of channel and spatial salient features and enables cross-dimensional adaptive interaction, enhancing the preservation of target features in deep networks and ensuring robust output. The MFAM module exhibits favorable plug-and-play characteristics. Experiments demonstrate that the proposed algorithm performs better in metrics of IoU, Pd, and Fa. A lightweight intelligent processing unit based on FPGA+GPU is developed, which successfully achieves deployment of the algorithm on edge devices and enables high real-time inference, demonstrating promising engineering applicability and future application prospects.
Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism
ZHANG Xin-di, CHEN Hui, ZHANG Hong-yun, LIAN Feng, ZHANG Guang-hua, YIN Zhi-peng
Available online  , doi: 10.11999/JEIT260045
Abstract:
  Objective  Multi-target tracking requires sensor management to adapt the observation process to clutter, missed detections, target-number variations, and changes in target motion. Conventional methods often search over a finite set of sensor actions, resulting in increasing computational cost and limited control resolution. Moreover, rewards formed by combining several single-target quantities may not adequately represent the joint multi-target posterior. A continuous-action sensor management method integrating the Twin-Delayed Deep Deterministic Policy Gradient algorithm with the Labeled Multi-Bernoulli filter is therefore developed to optimize mobile-sensor heading according to the multi-target belief state.  Methods  The LMB posterior, including target existence probabilities and state densities, is used to construct the belief state. At each filtering step, the mobile sensor selects a continuous heading angle that affects the sensor-target geometry, detection probability, and LMB update. Predicted target states and candidate headings are used to generate ideal measurements and obtain pseudo-updated LMB densities. The Cauchy-Schwarz divergence between the predicted and pseudo-updated densities is used to construct the information-gain reward. TD3 employs two critics, target policy smoothing, and delayed actor updates to reduce value-estimation bias. Random control, policy-gradient control, an information-driven discrete method, and DDPG-based continuous control are used for comparison.  Results and Discussions  DDPG-LMB and TD3-LMB produce more continuous steering changes than the discrete-action methods (Fig. 2). TD3-LMB obtains the highest or near-highest detection probabilities for most targets (Fig. 3) and gives relatively large Cauchy-Schwarz divergence values during most time steps, while random control remains relatively low (Fig. 4). TD3-LMB also achieves the lowest overall OSPA in the tested scenario, with DDPG-LMB generally outperforming the discrete baselines (Fig. 5). These results show that continuous heading control improves observation quality and overall LMB tracking performance.  Conclusions  A TD3-based continuous-action sensor management framework for the LMB filter is presented. Candidate heading actions are evaluated through pseudo-updated LMB densities and Cauchy-Schwarz divergence, linking action selection directly to the joint multi-target posterior. The simulation results show more continuous sensor motion, higher detection probabilities for most targets, larger information gain, and lower OSPA in the tested scenario. Future work will consider higher-dimensional actions and cooperative multi-sensor management.
A Behavioral Economics-Based Game Model for Side-Channel Security Attack and Defense Strategies
CAI Juesong, YAN Yingjian, WANG Jindong
Available online  , doi: 10.11999/JEIT260121
Abstract:
  Objective  The field of side-channel security currently lacks a systematic, quantifiable, and reproducible model for guiding the selection of attack and defense strategies, particularly in real-world engineering contexts where resource constraints necessitate informed cost-benefit trade-offs. The absence of such a framework impedes the practical realization of the “appropriate security” principle, often leading to either over-protection or under-protection of cryptographic modules. Traditional approaches to evaluating attack and defense costs rely heavily on subjective expert judgments, which are inherently arbitrary, difficult to replicate, and lack a structured multi-dimensional assessment. To bridge this critical gap, this research proposes an interdisciplinary model that integrates game theory, behavioral economics, and the Analytic Hierarchy Process (AHP). The primary objective is to establish a holistic decision-support system that not only quantifies the multi-faceted costs of various side-channel strategies but also incorporates the psychological dimensions of decision-making under risk, thereby enabling dynamic and economically rational security strategy selection tailored to specific asset values and security levels.  Methods  This study constructs a multi-layered modeling framework based on a static non-cooperative game with incomplete information. First, the side-channel analyst and the defense designer are formally defined as rational players, each possessing a finite set of strategies: the attacker may choose from non-modeling attacks such as DPA/CPA, modeling-based attacks like template attacks, or emerging deep learning-based side-channel analysis; the defender may adopt countermeasures including time hiding, amplitude hiding, or masking/blinding techniques. To systematically quantify the often-overlooked cost dimension, an AHP-based structured cost model is introduced. Through pairwise comparison matrices, the model decomposes costs into multiple criteria—such as time, data storage, computational resources, expertise, and hardware overhead—and assigns objective weights to each criterion, thereby replacing subjective cost estimates with a reproducible, hierarchical evaluation system. Furthermore, to reflect real-world decision-making behavior, key concepts from behavioral economics are integrated: Prospect Theory models how gains and losses are perceived relative to a reference point, while risk aversion coefficients capture players’ tolerance for uncertainty. These behavioral parameters are explicitly linked to the security level of the cryptographic module, allowing the model to adapt to different operational contexts. The resulting behavioral-augmented Bayesian game is then solved using the concept of Bayes-Nash Equilibrium, wherein each player’s optimal mixed strategy is derived based on their private type (behavioral profile) and beliefs about the opponent. To ensure engineering relevance, the As Low As Reasonably Practicable principle is incorporated as a constraint, enforcing that any selected defense strategy must be justifiable in terms of risk reduction versus cost incurred. Numerical solutions are obtained via a customized sequential quadratic programming algorithm implemented in Python.  Results and Discussions  A comprehensive experimental evaluation was conducted to validate the proposed model’s consistency, sensitivity, and practical utility. The AHP-based cost quantification demonstrated strong internal consistency, with all consistency ratios below the 0.1 threshold, confirming the reliability of the judgment matrices. The derived weight distributions revealed intuitive priorities: attackers placed greater emphasis on technical barriers and computational cost, whereas defenders prioritized design complexity and performance overhead. The behavioral adjustment layer successfully modulated perceived costs according to security levels: under low-security conditions (high risk aversion), costs were perceptually inflated, leading to conservative strategy choices; under high-security conditions, decision-making aligned more closely with objectively quantified costs. Equilibrium analysis across varying asset values and security levels yielded interpretable and rational strategy profiles. For low-value assets, both players exhibited a strong tendency toward low-cost or “no action” strategies, adhering to the lower bound of the ALARP region. As asset value increased, a clear threshold effect was observed, triggering a shift toward high-cost, high-efficacy strategies such as deep learning-based attacks and masking-based defenses. Sensitivity analysis further confirmed that defense strategy probabilities increased monotonically with asset value, validating the model’s ability to capture the non-linear relationship between protection intensity and asset criticality. These findings underscore the model’s capacity to support context-aware, adaptive security decision-making that balances risk, cost, and psychological factors.  Conclusions  This research presents a novel, behaviorally informed game-theoretic model for side-channel security strategy selection, addressing a significant void in existing literature regarding structured cost-benefit assessment. By integrating AHP-based objective cost quantification, behaviorally adjusted subjective valuations, and ALARP-driven engineering constraints, the proposed framework offers a multi-dimensional, reproducible, and context-sensitive tool for analyzing attack-defense interactions. The model advances the field by explicitly linking security levels to behavioral parameters, enabling dynamic strategy adaptation in response to both asset value and decision-makers’ risk perceptions. Although the current implementation relies partially on expert-defined parameters and operates within a static game setting, it establishes a critical foundation for transitioning from heuristic-based security decisions to quantitatively grounded, interdisciplinary analysis. Future work will focus on parameter calibration using real-world attack/defense datasets, extension to multi-stage dynamic games to capture strategic evolution over time, and empirical validation in industrial cryptographic evaluation scenarios. This study contributes the first systematic methodology for cost-aware, behaviorally realistic strategy optimization in side-channel security, offering both theoretical insights and practical guidance toward achieving “appropriate security” in cryptographic engineering.
Research Status and Prospects of Mid-Wavelength Infrared Superlattice Detector Technology
LIU Ming, ZHAO Yaqi, GUAN Xiaoning, ZHANG Fan, LU Pengfei
Available online  , doi: 10.11999/JEIT260083
Abstract:
  Significance   Mid-Wavelength Infrared (MWIR) detectors are widely used in civilian and military applications because of their high sensitivity and excellent temperature discrimination. Type-II SuperLattice (T2SL) materials, especially the InAs/GaSb and InAs/InAsSb systems, have become promising candidates for third-generation infrared photodetectors. This review systematically analyzes the research status and future trends of MWIR T2SL detector technology. It focuses on key photoelectric parameters, including Quantum Efficiency (QE), dark current density, and Specific Detectivity (D*). This work provides a reference for material selection and performance optimization in this rapidly developing field.  Progress   Considerable progress has been made in dark current suppression and photoresponse enhancement for MWIR T2SL detectors. For dark current suppression, advanced barrier structures, such as nBn, XBn, and M-structures, are designed through band-structure engineering. These structures effectively block majority-carrier transport while allowing efficient collection of photogenerated carriers. For instance, an nBn device with an AlAsSb/InAsSb superlattice barrier shows a dark current density of 2.01×10–5 A/cm2 at 150 K (Fig. 2(c)). Strain compensation and optimized epitaxial growth further reduce bulk dark current. One device achieves a dark current density of 4.5×10–7 A/cm2 at 140 K (Fig. 4(f)). Device process optimization, including two-step etching and Zn-diffusion-based planar junction formation, also reduces surface leakage current (Fig. 5, Fig. 6). For photoresponse enhancement, the main strategies include micro/nano-optical structure integration, epitaxial growth optimization, and device process improvement. Monolithically integrated metalenses increase th,e peak responsivity to 9.01 A/W at 300 K (Fig. 7(d)). Guided-mode resonance architectures enable a room-temperature External Quantum Efficiency (EQE) of approximately 60% (Fig. 8(c)). Epitaxial optimization, including stepped absorption layers and interfacial graded doping, increases the QE to 59.4% at 150 K (Fig. 10(c)). Device process optimization, such as substrate removal and Anti-Reflection (AR) coating deposition, also improves QE. An average QE of 63.7% is reported in the 3.7~4.8 μm range (Fig.13(c)). Comparative analysis shows that InAs/GaSb detectors are mainly reported at 77~150 K, whereas InAs/InAsSb detectors show stronger potential for higher-temperature operation, especially near 150 K (Fig. 15, Fig. 16). Overall, at 150K, dark current densities are generally suppressed below 10–4 A/cm2, and peak QEs approach 70%.  Conclusions  T2SL materials, with tunable band structures and low Auger recombination rates, have become a core material platform for high-performance MWIR detection. Current studies have addressed key challenges in dark current suppression and photoresponse enhancement. Through advanced barrier design and device process optimization, dark current densities have been suppressed to the 10–6 A/cm2 level at approximately 150 K. Through optical and epitaxial engineering, QEs have been increased to approximately 60% or higher. The InAs/InAsSb material system is particularly promising for High-Operating-Temperature (HOT) applications.  Prospects  Future development will focus on four main directions. First, the HOT limit should be further increased, with the goal of maintaining diffusion-limited performance at 180 K or higher. Second, large-format Focal Plane Arrays (FPAs) should be developed based on highly uniform material growth through mature Molecular Beam Epitaxy (MBE), aiming for pixel operability higher than 99%. Third, multicolor and multispectral detection should be expanded by precisely tuning superlattice periods, enabling integrated dual-band or multiband MWIR detection with reduced crosstalk. Fourth, new device architectures and coupled physical mechanisms should be explored to extend detector performance and application boundaries.
A Review of Advances and Challenges in Intelligent Disaster Assessment of High-Value Objects in Remote Sensing Imagery
ZHANG Yidan, FENG Yingchao, WANG Tianqi, LIU Yu, WANG Mengyu, HOU Zhongyan
Available online  , doi: 10.11999/JEIT251297
Abstract:
  Significance   The rapid and accurate disaster assessment of objects in the aftermath of disasters is critical for enabling effective emergency response and post-disaster recovery operations. Intelligent remote sensing technologies, empowered by deep learning, provide scalable, objective, and efficient means of assessing disaster across large and complex environments, such as densely populated urban centers, transportation hubs, and critical energy infrastructure. By leveraging high-resolution satellite and aerial imagery, these technologies can provide timely situational awareness for decision-makers, supporting prioritization in rescue and recovery tasks. Despite significant advancements in methods and applications, there remains a lack of comprehensive synthesis in the field, leading to fragmented practices and inconsistent benchmarks across studies. This study addresses this critical gap by providing a structured and systematic review that consolidates technical foundations, commonly used datasets, evaluation metrics, and methodological advances in deep learning-based remote sensing disaster assessment. The overarching goal is to accelerate the adoption and operationalization of these technologies in real-world disaster scenarios, ultimately contributing to the construction of resilient cities and infrastructures under increasing environmental and geopolitical risks.  Progress   In recent years, deep learning-based remote sensing disaster assessment methods have developed rapidly, demonstrating notable improvements in classification accuracy, processing automation, and scalability. Significant advances include bi-temporal change detection methods that can precisely localize disaster by capturing differences before and after disaster events, and multi-temporal sequence modeling approaches that extract evolving disaster patterns and degradation trends across time-series data. Furthermore, multi-modal data fusion strategies that combine optical, Synthetic Aperture Radar(SAR), and Light Detection and Ranging (LiDAR)data have enhanced the ability to analyze complex disaster characteristics under varying observation conditions. Advanced techniques to address data scarcity, such as transfer learning and self-supervised learning, have further extended the applicability of disaster assessment methods in data-constrained environments. These advances collectively contribute to improving the responsiveness and effectiveness of disaster assessment systems in supporting emergency decision-making and resource allocation during critical disaster events.  Conclusions  This study systematically categorizes and evaluates the technical landscape of deep learning-based remote sensing disaster assessment for objects, highlighting the strengths and limitations of bi-temporal, multi-temporal, multi-modal, and data-scarce scenario methods. While existing approaches demonstrate considerable potential in addressing various challenges in post-disaster assessment, challenges remain in ensuring robustness across diverse environments and operational conditions. The review underscores the need for standardized disaster classification criteria and comprehensive evaluation frameworks that consider both physical disaster and functional impacts to facilitate practical and consistent disaster assessments in real-world deployments.  Prospects   Future research in remote sensing-based disaster assessment should focus on developing multi-level collaborative frameworks for evaluating diverse objects across spatial and functional scales, enabling holistic disaster impact assessment that captures both direct and cascading effects. Complex scenarios such as airports, industrial zones, and ports contain static structures, moving objects, and interdependent functional units, thus requiring hierarchical modeling and multi-object reasoning. Physics-driven and hybrid approaches integrating structural mechanics, material degradation, and expert knowledge can further improve interpretability and generalization. Meanwhile, lightweight model design and edge deployment are important for real-time assessment on drones and satellites in emergency situations. Standardized evaluation metrics that combine physical disaster mechanisms with functional impact analysis will also be essential for practical deployment. Together, these directions will help transform intelligent remote sensing technologies into actionable tools for disaster response and recovery.
An Overview of Key Technologies for 6G-Enabled Communication-Computing Integration and Energy-Efficiency Optimization
LIU Guangyi, CAI Qing, WANG Xinyao, CHEN Tianjiao, JIN Jing, XUE Yahui, WANG Ailing, WANG Hanning
Available online  , doi: 10.11999/JEIT260399
Abstract:
  Significance   Constrained by size, power consumption, and cost, emerging intelligent terminals often face excessive energy consumption and limited battery life. These limitations have become major bottlenecks to large-scale deployment. Compared with Fifth-Generation (5G) wireless networks, Sixth-Generation (6G) wireless networks are expected to enhance the Radio Access Network (RAN) architecture and move computing capability toward the RAN side. High-energy and compute-intensive Artificial Intelligence (AI) tasks that are originally executed by end devices can therefore be processed by the network. Through End-Edge Collaboration, emerging intelligent terminals can be upgraded toward lightweight design, low cost, and long battery life, thereby supporting the large-scale deployment of ubiquitous intelligence in 6G networks.  Progress   Current progress in Terminal Energy Consumption Optimization through 6G End-Edge Collaboration is reviewed, with emphasis on local execution, full offloading, and partial offloading. In local execution, User Equipment (UE) processes all tasks locally, which leads to high computing energy consumption. In full offloading, all tasks are transferred to the RAN. This reduces terminal-side computing energy consumption but can increase transmission energy consumption, especially under poor channel conditions. Partial offloading combines the benefits of both modes and optimizes energy consumption according to real-time network conditions. For partial offloading, four representative optimization techniques are summarized. (1) Feature Extraction and Filtering. Semantic encoding and information extraction are performed at the UE, and only task-relevant data are transmitted to the RAN. This reduces redundant data transmission and lowers transmission energy consumption. (2) Split Offloading. A large Deep Neural Network (DNN) is divided into layers according to its structure. Simpler shallow layers are processed at the UE, whereas more complex deep layers are offloaded to the RAN. This method balances terminal-side and RAN-side computational loads through End-Edge Collaborative Inference. (3) Model Lightweighting. Model complexity is reduced through pruning, quantization, and knowledge distillation, which lowers computational overhead while maintaining task performance. (4) Incremental Inference. Only changed data or updated features are processed, while historical computations are reused. This reduces redundant computation. Together, these techniques improve terminal performance and energy efficiency within the 6G End-Edge Collaboration framework.  Conclusions  This paper systematically reviews Terminal Energy Consumption Optimization for 6G End-Edge Collaboration. It summarizes the functional evolution of enhanced RAN, constructs an End-Edge Collaborative service framework for Communication-Computing Integration, and establishes a theoretical model of terminal computing energy consumption and transmission energy consumption. The composition and influencing factors of energy consumption under different offloading modes are clarified. Key energy optimization technologies, including Feature Extraction and Filtering, Split Offloading, Model Lightweighting, and Incremental Inference, are then discussed. To address energy consumption fluctuations caused by dynamic wireless channels, the paper proposes energy optimization mechanisms based on Adaptive Semantic Compression, Dynamic Split Offloading, Adaptive Model Pruning, and Incremental Inference. These mechanisms maintain a dynamic balance between energy optimization and task performance. Using embodied intelligent robot video understanding as a typical application scenario, a test platform is developed to verify the effectiveness of the proposed mechanisms. Current challenges and future research directions are also analyzed.  Prospects   Although End-Edge Collaborative energy-saving technologies have achieved initial progress, practical deployment still faces challenges in real network environments, dynamic wireless channels, and large-scale user access. Future research should examine the trade-off between optimization overhead and system robustness. It should also study dynamic communication-computing resource substitution modeling in stochastic resource environments, multi-user collaboration strategies, and global energy-efficiency optimization. As the technology matures, standardization and engineering implementation of End-Edge Collaborative energy-saving frameworks will become critical to the large-scale adoption of 6G applications. Future studies should further integrate algorithm design with network architecture, enabling practical deployment of low-power and high-efficiency intelligent communication systems.
Gating Adaptive Repeat Query Framework for Reliable Collaborative Inference with Edge Heterogeneous LLMs
WANG Tengsheng, YU Tao, LI Jihong, ZHENG Guhan, ZHANG Shunqing
Available online  , doi: 10.11999/JEIT260218
Abstract:
  Objective  Reliable inference at the network edge is indispensable for 6G-enabled ubiquitous AI, yet the deployment of large language models (LLMs) in such environments remains a cornerstone challenge. In resource-constrained edge settings, single-LLM inference often proves unreliable due to knowledge limitations and inherent biases, severely hampering real-world deployment. Collaborative inference leveraging multiple heterogeneous LLMs emerges as a promising remedy to boost robustness, but it introduces nontrivial hurdles under stringent latency and energy budgets, especially when wireless channel conditions and query content vary unpredictably. These challenges include the need for dynamic sequential decision-making for LLM selection and resource allocation, the fundamental paradigm mismatch between bit-level reliability protocols and semantic-level error correction, and the lack of adaptive mechanisms to align and fuse disparate LLM outputs effectively. To fill these critical gaps, this paper presents a novel framework that fundamentally reinterprets collaborative inference as a semantic-driven, closed-loop process, thereby transitioning from conventional bit-retransmission to semantic-retransmission and offering a practical path toward reliable 6G edge intelligence.  Methods  In response to these critical challenges, we propose the Gating Adaptive Repeat Query (G-ARQ) framework. Its core innovation is the Semantic-Space Alignment and Error-Guided Retransmission (SEMAR) mechanism. SEMAR first aligns the token-level probability distributions from heterogeneous LLMs into a unified semantic space using relative representation, enabling comparable outputs. It then models the collaborative process probabilistically, explicitly capturing error dependencies among models, and uses an Expectation-Maximization (EM) algorithm to infer a latent error direction, which guides the selection of the next LLM for query retransmission, steering it towards outputs orthogonal to previous errors. To jointly optimize the LLM gating and uplink power allocation under communication constraints without requiring explicit system dynamics—often unavailable in practice—we design a black-box trajectory optimizer. This optimizer formulates the sequential decision problem as sampling from a target distribution that encodes dynamic feasibility, optimality, and constraints. It employs a diffusion-based sampling process with a model-guided prior and Monte Carlo estimation to generate near-optimal policy trajectories that satisfy hard latency and energy limits.  Results and Discussions  To evaluate the practical viability of G-ARQ under realistic edge conditions, simulations are conducted in a scenario with five base stations hosting five heterogeneous 7B-parameter LLMs (Mistral-7B, Vicuna-7B, Nous-Capybala-7B, Gemma-7B, and Llama-2-7B). The user equipment (UE) performs a question-answering task evaluated on a mixed SQuAD and TriviaQA dataset. Component-level evaluations, each designed to isolate the contribution of a single innovation, validate the effectiveness of every key element. The error-guided gating of SEMAR, compared to a Top-k gating baseline, improves accuracy by 0.23 % on average, and its dynamic weight ensemble contributes an additional 0.7 % gain (Fig. 3). The black-box trajectory optimizer, which operates without any explicit channel model, achieves accuracy close to that of the unconstrained model-greedy strategy while ensuring strict latency constraints (Fig. 4). The convergence of the optimizer is verified by tracking the evolution of \begin{document}$ J(\boldsymbol{S}) $\end{document} and selection probability over diffusion steps (Fig. 5). System-level performance under varying latency and energy constraints demonstrates that G-ARQ consistently surpasses two baselines: one combining model-greedy selection with Proximal Policy Optimization (PPO) for power optimization, and another combining model-greedy selection with Simulated Annealing for power optimization, both using average output weights. The accuracy improvement is most significant under the most stringent resource limits, reaching up to 2.2 % for \begin{document}$ {K}_{\max }=1 $\end{document} and 1.9 % for \begin{document}$ {K}_{\max }=2 $\end{document} (Fig. 6, Fig. 7). The framework successfully establishes a Pareto boundary that characterizes the inherent trade-off between inference accuracy and communication latency, providing a valuable design guideline for resource-constrained edge systems and offering actionable insights for real-world deployment. The GARQ-S variant is noted to outperform GARQ-E by avoiding the integration of outputs from previously erroneous models.  Conclusions  This paper proposed G-ARQ framework, an innovative closed-loop framework that transforms collaborative edge inference into a semantics-guided retransmission process. By introducing SEMAR for error-based alignment and selection of heterogeneous LLMs, and employing a black-box trajectory optimizer for joint model selection and power allocation, the framework achieves up to a 2.2% accuracy improvement under strict resource constraints. The results validate G-ARQ as an effective and practical approach toward reliable and efficient 6G edge intelligence.
A Two-layer Closed-loop Cooperative Resource Allocation Framework for Improving QoS of MEC Network Slicing
XU Juntao, FAN Xinggang, XU Changfu, SHEN Minyang, LIANG Yuzhu, WANG Tian
Available online  , doi: 10.11999/JEIT260156
Abstract:
  Objective  Driven by 5G/6G networks, Multi-access Edge Computing (MEC) environments face major challenges in ensuring Quality of Service (QoS) for heterogeneous network slices while improving resource utilization. Existing resource allocation methods often lack dynamic adaptability in heterogeneous settings. They also fail to jointly optimize caching, bandwidth, and computing resources, which reduces resource utilization and service success rates. This paper addresses these limitations by proposing a robust framework for joint multi-dimensional resource optimization under strict QoS constraints. The framework provides a tailored solution for heterogeneous MEC environments.  Methods  The network slicing resource allocation problem is first formally defined. Its NP-hardness is proved by reducing the NP-hard multidimensional 0-1 knapsack problem to this problem. To solve this complex optimization problem, a cache-aware two-layer closed-loop cooperative framework, termed QCache, is proposed. The framework uses a synergistic “generate-evaluate-feedback” loop. In the upper Global Exploration Layer, a hybrid heuristic algorithm is designed by combining a population evolution strategy, including selection, crossover, and mutation, with a particle update mechanism guided by historical individual and global best positions. This layer broadly explores the solution space and generates high-quality candidate resource allocation schemes under complex constraints. Adaptive parameter adjustment and elite retention strategies are used to avoid local optima. In the lower Multi-Dimensional Weight Evaluation Layer, a quantitative assessment model is constructed. This model converts low-latency and high-bandwidth service demands into explicit QoS constraints by normalizing key performance indicators, including delay and rate, and by dynamically assigning weights through the entropy weight method. The weighted score reflects different slice priorities. The evaluated score (SCORE) from this layer is fed back to the upper layer as the fitness value, guiding the iterative evolution of candidate solutions until convergence.  Results and Discussions  Extensive simulations are conducted to validate the effectiveness of QCache against several baseline methods, including Genetic Algorithm (GA), PSO-Leader, GraphSAGE, and the No-Consideration-of-Service-Quality (NCSQ) scheme. Under identical resource and user demand scenarios, the overall comparison (Fig. 3) shows that QCache achieves the highest resource utilization rate of 80.30% and the best average service score of 1.109. Compared with the baseline methods, QCache improves resource utilization by 2.29% to 24.50% and increases user service scores by 4.13% to 59.34%. Experiments with varying total cache resources (Fig. 4) show that QCache maintains superior performance across different cache states. It improves resource utilization by 2.43% to 27.53% and service scores by 10.81% to 119.56%, confirming its cache-aware adaptability. Tests with increasing user numbers (Fig. 5) show that QCache scales effectively, achieving up to 85.83% resource utilization and a service score of 3.06. These results demonstrate its ability to handle dense access scenarios. Experiments with time-varying user demands (Fig. 6) further confirm the dynamic robustness of the framework. In these tests, QCache achieves average improvements of 9.20% in resource utilization and 23.45% in service score over the baselines.  Conclusions  This paper studies the NP-hard resource allocation problem in dynamic MEC environments with heterogeneous network slices. The proposed cache-aware two-layer closed-loop cooperative framework, QCache, jointly optimizes caching, bandwidth, and computing resources under explicit QoS constraints. The upper-layer hybrid heuristic provides strong global search capability. The lower-layer multi-dimensional weight model supports accurate QoS quantification and dynamic feedback. Comprehensive experimental results show that QCache outperforms existing methods in both resource utilization efficiency and user QoS satisfaction. Future work will explore reinforcement learning and traffic prediction mechanisms to further improve the response of the framework to bursty traffic and anomalous demands. This may support more intelligent and autonomous MEC network slice resource management.
Queue Stability Constrained Robust Secure Beamforming for Low-Altitude UAV-ISAC Systems
OUYANG Jian, REN Wei, XU Ba, LIU Xiaoyu, JIANG Wanmu
Available online  , doi: 10.11999/JEIT260275
Abstract:
  Objective  To address the challenges of antenna array angle errors caused by UAV jitter, transmission instability induced by random data arrivals, and secure transmission guarantee in multi-eavesdropper scenarios for low-altitude UAV-ISAC systems, this paper proposes a robust secure beamforming algorithm based on queue stability constraints. The proposed algorithms aims to minimize long-term transmit power consumption while maintaining data queue stability while enhancing beamforming robustness against UAV jitter.  Methods  To guarantee the stability of the wireless transmission in UAV-ISAC systems, this paper formulates an optimization problem aimed at minimizing the long-term average transmit power, subject to constraints on data queue stability, secure communication rate, sensing performance, and maximum transmit power. Since this long-term optimization problem is intractable, the Lyapunov optimization framework is employed to transform it into a sequence of short-term subproblems. To handle the antenna array angle errors caused by UAV jitter within each short slot, we jointly adopt the second-order Taylor series expansion and the S-Procedure method to approximate the short-term subproblem into a convex form. Consequently, a robust secure beamforming algorithm based on penalty successive convex approximation optimization is proposed.  Results and Discussions  Simulation results demonstrate the impact of the number of antennas, secrecy rate threshold, beam gain threshold, UAV jitter error, and Lyapunov weight factor on the system transmit power. As illustrated by the beam gain pattern in Fig. 3, the communication beamformer facilitates cooperative sensing toward the sensing area, while the sensing beamformer enhances system security by directing interference toward potential eavesdroppers. This validates the effectiveness of the proposed algorithm in simultaneously improving sensing and secure communication performance. Furthermore, leveraging the dual-function characteristics of the communication and sensing beamformers, the proposed integrated sensing and communication scheme achieves significantly higher resource utilization efficiency than the communication-only and sensing-only schemes, as shown in Fig. 5. Additionally, Fig. 6 indicates that, compared with the non-robust scheme, the proposed robust scheme strictly satisfies security requirements under various angle errors. Finally, Fig. 8 shows that the proposed queue-aware scheme can effectively suppress the transmit power fluctuations caused by random data arrivals, exhibiting superior stability compared to the queue-free baseline.  Conclusions  This paper investigates a robust secure beamforming method for UAV-ISAC systems subject to queue stability constraints. First, based on the Lyapunov optimization framework, the long-term stochastic optimization problem is transformed into a sequence of short-term subproblems. Second, to address the issue of UAV jitter, the second-order Taylor series expansion and the S-Procedure method are jointly employed to approximate the non-convex constraints into tractable convex forms. Finally, a robust secure BF optimization algorithm based on penalty successive convex approximation is proposed to efficiently solve the deterministic short-term subproblems. Simulation results demonstrate that the proposed scheme can effectively tackle the challenges posed by random data arrivals, UAV jitter, and eavesdropping threats, thereby ensuring the stability and security of downlink data transmission in low-altitude UAV-ISAC systems.
VT2R: Video and Text-driven Method for Generating Large-scale Millimeter-wave Radar Data
DENG Kaikai, LING Yue, XING Ling, WU Honghai, ZHAO Dong, MA Huahong
Available online  , doi: 10.11999/JEIT260240
Abstract:
  Objective  The lack of large-scale training data impedes progress in developing robust and generalized deep learning models. However, existing millimeter-wave radar data generation methods are ineffective due to a lack of sufficient data sources. To address this gap, this paper proposes a video and text-driven radar data generation method, VT2R, which utilizes video or text data to generate large-scale, realistic radar data, solving the key problem of constructing the mapping relationship between video and text and radar data.  Methods  The proposed method consists of three main components: video feature encoding network, text feature encoding network, radar feature encoding network and data fitting and decoding network. Video feature encoding networks and text feature encoding networks extract temporally consistent visual representations and alignable semantic features, respectively, while the radar encoding network learns the structure and dynamic information of point clouds through hierarchical spatiotemporal modeling. In the data fitting and decoding network based on Variational AutoEncoder (VAE), multi-modal features are mapped to a unified latent distribution space and decoded into radar data through reparameterized sampling. During training, reconstruction loss, Kullback-Leibler (KL) divergence loss, and cross-modal similarity loss are jointly optimized.  Results and Discussions  This paper constructs the first radar point cloud dataset for reclining gesture recognition (Figs. 6 and 7), covering 5 gesture categories, 32 participants, and a total of 14,400 samples. Experimental results based on this dataset show that VT2R achieves a recognition accuracy of 89.2% when trained using only generated radar data, a 33.88% improvement over the representative RFGen (Figs. 9 and 10). When combined with a small amount of real radar data for joint training, the accuracy further improves to 97.62%, a 21.48% improvement over RFGen (Figs. 9 and 11). Furthermore, VT2R still achieves average recognition accuracies of 89.35% and 97.21% under different scenarios and factors (Figs. 16-18). In addition, this paper also verifies the accuracy of VT2R under different postures, achieving average accuracies of 89.98% and 97.55% in the first and third settings, respectively (Fig. 19), which is basically consistent with the result obtained when lying down, demonstrating its robustness under cross-posture conditions.  Conclusions  This paper proposes a radar data generation system, VT2R, which addresses the severe lack of realistic radar training data when users are performing gestures in a lying position. Through a video feature encoding network built on a vision-language pre-trained model, a text encoding network incorporating cue templates, a hierarchical radar encoding network for sparse point clouds, and a VAE-based data fitting and decoding network, these components collaboratively generate large-scale, realistic radar data. It also supports augmented reconstruction based on limited real radar data, providing rich data support for radar perception tasks. Future work will focus on solving multi-modal data generation for more complex gesture scenes, providing better data support for emerging large-scale models.
Multi-task Lightning Nowcasting with Spatio-temporal Focal Perception and Synergistic Weighted Loss
TANG Zhihao, HAN Yuanpeng, ZHANG Hui, SONG Lin, ZHANG Qilin, LIU Yi
Available online  , doi: 10.11999/JEIT260234
Abstract:
  Objective  Lightning nowcasting is essential for early warning systems and for protecting critical infrastructure, including aviation, power grids, and transportation systems. Traditional numerical weather prediction models depend strongly on parameterization schemes and require high computational costs, which limits their use in rapid-update nowcasting. Although deep learning methods have advanced, they still have difficulty handling extreme data sparsity, suffer from serial-computation bottlenecks in recurrent architectures, and mainly focus on binary occurrence prediction rather than the joint optimization of lightning-frequency prediction and regional localization. Moreover, conventional loss functions are easily dominated by extensive non-lightning areas, which biases predictions toward zero or causes excessive false alarms. To address these limitations, a Spatio-Temporal Focal perception and synergistic weighted loss Network (STF-Net) is proposed as a multi-task lightning nowcasting model that jointly predicts lightning frequency and occurrence regions. It integrates three key components: a Lightning Adaptive Attention Module (LAAM) for explicit spatio-temporal dependency modeling, a Spatio-Temporally Weighted Hybrid Loss for data sparsity and imbalance, and a spatio-temporal dual-branch Generative Adversarial Network (GAN) to improve prediction fidelity and temporal coherence.  Methods  STF-Net is built on the SimVP video prediction architecture and adopts an encoder-translator-decoder paradigm. LAAM uses a three-dimensional decoupled attention mechanism along the height, width, and channel dimensions, enabling adaptive focus on convectively sensitive regions while maintaining computational efficiency. The Spatio-Temporally Weighted Hybrid Loss combines Temporally Weighted Mean Squared Error (TW-MSE) for frequency regression and Dual-Weighted Cross-Entropy loss (DWCE) for regional localization. Time-increasing weights are incorporated to improve medium- to long-term forecast robustness (Fig. 5). DWCE integrates static class weights with dynamic grid weights, thereby balancing global class proportions and local lightning-frequency heterogeneity. A spatio-temporal dual-branch GAN, consisting of a spatial PatchGAN discriminator and a temporal three-dimensional convolutional discriminator, is used to improve the textural fidelity and temporal coherence of predicted lightning-frequency fields. The model uses six consecutive historical lightning-frequency frames at 256×256 resolution and 10-min intervals to predict the next six frames, corresponding to a 1-h forecast window. Experiments are conducted on a high-resolution Very Low Frequency Lightning Location Network (VLF-LLN) dataset containing 11,748 images that cover different seasonal and weather conditions. The dataset is split at a ratio of 7:3 for training and testing.  Results and Discussions  Comprehensive evaluation metrics are used, including Mean Squared Error (MSE) and Mean Absolute Error (MAE), which are computed only on lightning pixels; Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) for image fidelity; and Probability Of Detection (POD), False Alarm Rate (FAR), and Critical Success Index (CSI) for regional detection. STF-Net achieves a CSI of 0.663 within the 1-h forecast window, representing a 14.5% improvement over the SimVP baseline (0.579). It also reduces FAR from 0.351 to 0.216, corresponding to a relative reduction of 38.5% (Table 1). Ablation studies validate each component. Adding GAN improves CSI to 0.624 and reduces MSE to 0.109. Further incorporation of LAAM increases CSI to 0.629 and yields the highest POD of 0.894. The complete STF-Net with the hybrid loss achieves the best performance, with a CSI of 0.663 and an MSE of 0.105 (Table 1, Fig. 5). The joint prediction of frequency and region is supported by simultaneous improvements in regression metrics (MSE and MAE) and detection metrics (CSI and FAR). Time-step analysis shows that LAAM reduces long-term performance degradation, with STF-Net maintaining the highest CSI compared with SimVP+GAN and SimVP (Fig. 6). Comparative experiments with ConvLSTM and PredRNN further demonstrate the superiority of STF-Net across all lead times. STF-Net consistently achieves higher CSI and lower FAR, and its advantage becomes more evident as the forecast horizon increases (Fig. 7). Its consistent gains in PSNR and SSIM further indicate that the spatio-temporal GAN helps generate coherent, detail-rich predictions. Visualization results show that STF-Net produces structurally clear and continuous lightning-activity regions centered on high-frequency areas. It accurately tracks dynamic evolution patterns, including movement, merging, and splitting, while generating minimal noise in non-lightning regions. These results demonstrate effective collaborative prediction of both frequency magnitude and spatial distribution.  Conclusions  STF-Net is presented as a deep learning model for joint lightning-frequency prediction and regional localization. It explicitly models long-range spatio-temporal dependencies, focuses on critical convective zones, addresses extreme data sparsity and class imbalance, and jointly optimizes frequency regression and regional localization while suppressing false alarms. Its spatio-temporal dual-branch GAN further improves the spatial structural consistency and temporal coherence of predictions. Experimental results show that STF-Net outperforms baseline and state-of-the-art models, achieving a CSI of 0.663, an FAR of 0.216, and the best MSE and MAE values within a 1-h forecast window. The model effectively reduces long-term performance degradation, captures the evolution trends of lightning regions, and generates physically plausible predictions with minimal background noise. This study provides an efficient end-to-end solution for operational lightning nowcasting systems and offers guidance for model design in sparse meteorological spatio-temporal sequence prediction.
Intelligent Privacy-Aware Computation Offloading Method against Multi-server Joint Inference Attacks
MIN Minghui, LIU Mingcheng, ZHANG Peng, DUAN Jincheng, LI Shiyin, ZHANG Hongliang
Available online  , doi: 10.11999/JEIT260249
Abstract:
  Objective  With the rapid development of the low-altitude economy, services such as intelligent transportation, smart healthcare, and low-altitude logistics have become increasingly common. Their efficient operation depends on the real-time processing of massive sensing data. Mobile Edge Computing (MEC) improves task execution efficiency and reduces device computational burdens by offloading tasks to nearby servers. However, user privacy and security risks have become increasingly severe. In dynamic scenarios where multiple MEC servers jointly process tasks, information sharing can enable multi-server joint inference attacks and greatly increase the risk of user location privacy leakage. Although existing studies have used Differential Privacy (DP) to protect user location privacy, current DP-based solutions remain limited. These methods inject noise into offloading decisions, but unconstrained noise may reduce task allocation accuracy. In addition, user mobility causes continuous changes in channel states during dynamic computation offloading. Privacy leakage risks and attacker behaviors are also uncertain. Traditional optimization methods based on static system models are therefore unsuitable for such dynamic environments. To address these challenges, this paper proposes an Asynchronous Advantage Actor-Critic (A3C)-based Intelligent Privacy-Aware Computation Offloading (AIPCO) scheme against multi-server joint inference attacks. The proposed scheme protects user location privacy while maximizing the overall utility of the MEC system.  Methods  This paper proposes a DP-based task offloading rate perturbation mechanism. By adding controlled noise, the mechanism increases the randomness of user task offloading toward multiple MEC servers. A truncated Laplace mechanism is used to constrain the perturbed offloading rates within valid boundaries. This design satisfies the mathematical guarantees of DP and reduces the accuracy of multi-server joint inference attacks on sensitive user locations. Privacy entropy is then introduced to dynamically evaluate the real-time privacy protection level. Finally, the AIPCO scheme is constructed. Through a multi-threaded asynchronous training mechanism, the scheme interacts with the environment through iterative trial and error and efficiently learns the optimal real-time offloading policy online. The proposed scheme dynamically protects user privacy, reduces computational cost, and maximizes comprehensive system utility.  Results and Discussions  The AIPCO scheme jointly optimizes user privacy and task offloading cost by incorporating multidimensional performance variables into the reinforcement learning reward function. A comprehensive performance analysis (Fig. 4) shows that, when the number of continuous learning iterations reaches 200, the privacy protection level of AIPCO increases by 2.52%, 3.56%, and 22.90% compared with RCLM, JODRL, and DODA-DT, respectively. This advantage is mainly attributed to the DP-based task offloading rate perturbation method, which uses the truncated Laplace mechanism to increase data randomness while strictly constraining the perturbation range. By contrast, RCLM perturbs the task offloading rate through range-limited DP without using the truncated Laplace mechanism. JODRL increases randomness only through network policy optimization, resulting in a lower privacy protection level. DODA-DT focuses on balancing energy consumption and system latency without optimizing user privacy. For the privacy weight parameter (Fig. 5), increasing $\omega$ improves privacy protection. For example, the privacy protection level increases by 5.64% when $\omega$ rises from 0.2 to 0.7, with a clear performance gain at 0.7. As the system agent reduces its focus on computational cost, user utility remains optimal despite increased cost. When the physical distance between users and the MEC server is adjusted (Fig. 6), AIPCO shows stronger privacy protection in long-distance scenarios. A greater distance reduces the number of tasks offloaded to the server. Therefore, attackers obtain less information, and privacy protection improves. Although computational cost increases with distance, AIPCO consistently outperforms competing schemes. These results confirm that AIPCO achieves optimal MEC system utility while protecting user privacy.  Conclusions  To mitigate multi-server joint inference attacks caused by information sharing among collaborative MEC servers, this paper proposes an AIPCO method. A DP-based task offloading rate perturbation scheme is designed to increase randomness, and a truncated Laplace mechanism is used to constrain perturbed rates within reasonable boundaries. The scheme is proven to satisfy strict DP mathematical guarantees, and privacy entropy is introduced to quantitatively evaluate the privacy protection level. In addition, the AIPCO scheme uses a multi-threaded asynchronous training mode, enabling the agent to efficiently learn the optimal perturbed offloading policy in a continuous space and maximize overall system utility. Simulation results show that the proposed scheme outperforms the baselines in both dynamic and average performance. It achieves optimal system utility while protecting user privacy.
Bearing Fault Diagnosis of Roadheader via Cross-modal Kernel Fusion-sphere Space Learning
SU Shuzhi, GUI Yang, MA Tianbing, ZHU Yanmin, WU Kanghui
Available online  , doi: 10.11999/JEIT260494
Abstract:
  Objective  Traditional roadheader bearing fault diagnosis methods often struggle with high-dimensional and nonlinear multi-sensor data. They also fail to effectively perceive cross-modal, multi-scale fault information or integrate local and global structural features. To address these limitations, this paper proposes a Cross-modal Kernel Fusion-sphere Space Learning (CKFSL) method. By perceiving cross-modal multi-scale fault information, CKFSL extracts highly discriminative features from roadheader bearing cross-modal fault samples and improves diagnostic accuracy.  Methods  CKFSL first maps roadheader bearing cross-modal fault samples into a high-dimensional kernel space through implicit transformation. Dual extremal point anchoring and polar neighbor allocation mechanisms are then used to capture fault sample clusters with similar isomorphic information, forming kernel fusion-spheres. An adaptive binary partitioning strategy is designed according to the geometric span of internal fault samples. This strategy tightens isomorphic boundaries, constructs micro-neighbor kernel fusion-spheres, and achieves highly isomorphic manifold aggregation at the microscopic scale. A micro-neighbor kernel fusion-sphere space is further formed to re-evaluate local isomorphism (Fig. 1). To characterize wide-area topological correlations, a wide-area topological isomorphism constraint is proposed. This constraint constructs a wide-area dynamic isomorphism graph among micro-neighbor kernel fusion-spheres (Fig. 1). Finally, an objective optimization function is formulated within the space learning framework. It integrates local manifold isomorphism and wide-area topological correlations of roadheader bearing cross-modal fault samples, as shown in the CKFSL diagnostic flowchart (Fig. 2). The analytical solution for spatial projection is theoretically derived to obtain discriminative cross-modal kernel fusion-sphere space isomorphic features from roadheader bearing cross-modal fault samples.  Results and Discussions  CKFSL is first validated on the self-built AUST roadheader bearing cross-modal fault dataset, with the experimental platform shown in Fig. 3. The average recognition rates obtained with increasing numbers of training fault samples are shown in Fig. 4. On the AUST dataset, CKFSL achieves a recognition rate of 99.49% with only 70 training fault samples and reaches 100% as the number of training fault samples increases. Table 1 summarizes the standard deviations under different training fault sample sizes. The results show that CKFSL has the lowest standard deviation and stronger robustness than the other seven comparison algorithms. Three-dimensional fault feature distributions are shown in Fig. 5. The results confirm that CKFSL effectively separates highly overlapping fault samples into different clusters and reduces the boundary confusion observed in the comparison algorithms. To verify generalization capability, CKFSL is further evaluated on the public Paderborn dataset, with the experimental setup shown in Fig. 6. As shown in Fig. 7 and Fig. 8, CKFSL achieves a 100% average recognition rate across four complex fault categories. It also outperforms the comparison algorithms, which have difficulty exceeding an 85% recognition rate for the F4 fault category.  Conclusions  CKFSL effectively addresses the inability of traditional roadheader bearing fault diagnosis methods to perceive complex multi-scale fault information. By using the wide-area dynamic isomorphism graph learned in the micro-neighbor kernel fusion-sphere space, CKFSL integrates local manifold isomorphism with wide-area topological correlations of roadheader bearing cross-modal fault samples. This process enables CKFSL to extract highly discriminative cross-modal kernel fusion-sphere space isomorphic features. It improves the accuracy of roadheader bearing fault diagnosis and supports the reliability and continuous operation of roadheaders.
A General Evaluation Framework for Mission Planning Algorithms for Remote Sensing Satellite Constellations
LI Jinfei, YU Xiaogang, TIAN Jing, HE Haochen, XING Xiangwei, ZHANG Xiaohan
Available online  , doi: 10.11999/JEIT260335
Abstract:
  Objective  The rapid growth in remote sensing satellite constellations has shifted mission planning from single-satellite static scheduling to large-scale dynamic coordination across heterogeneous constellations. However, evaluation methods have not kept pace with algorithm development. Existing studies often rely on private datasets, simplified metrics centered on Completion Rate, and idealized simulations that ignore realistic constraints, such as attitude maneuvers, illumination conditions, and dynamic task insertion. These limitations prevent fair cross-paper comparison and slow engineering application. To address this gap, this paper proposes the Remote Sensing Constellation Mission Planning Benchmark (RSCMP-Bench), a general, open, and reproducible evaluation framework. It is designed as a unified benchmark for the community, similar to ImageNet in computer vision and General Language Understanding Evaluation (GLUE) in Natural Language Processing (NLP).  Methods  RSCMP-Bench consists of three components. First, the multi-scenario standard task library contains 300 standardized scenarios at three difficulty levels: Low, Medium, and High, with 100 scenarios per level. Satellite numbers range from 30 to 200, and task demands range from 56 to 560. All scenarios are generated from public Two-Line Element (TLE) data and explicitly model realistic constraints. Optical satellites require a minimum solar elevation angle, and Synthetic Aperture Radar (SAR) satellites require incidence angles within specified ranges. General constraints, such as per-orbit maximum on-time, minimum single-operation on-time, attitude maneuver time, and valid execution windows, are also modeled. The scenarios include point tasks, area tasks, static tasks, and dynamically inserted tasks. Second, the multi-dimensional effectiveness evaluation system includes a Basic Performance layer and a Dynamic Adaptability layer. The Basic Performance layer uses Completion Rate, Weighted Completion Rate, Average Response Delay, and Time Utilization. The Dynamic Adaptability layer uses multi-stage rolling evaluation with random dynamic task insertion. The Dynamic Adaptability Score measures the post-insertion Completion Rate relative to the baseline, and Dynamic Response Efficiency measures the performance gain per unit replanning time. A composite RSCMP-Bench Score is also provided. Third, the simulation and evaluation platform uses a client-server architecture. It integrates a Simplified General Perturbations 4 (SGP4) propagator, algorithm adapters, two-stage constraint verification, an intelligent scenario generator, and visualization tools. The platform has been deployed at https://www.tianzhibei.com and has supported a national competition with more than 80 research teams.  Results and Discussions  Baseline experiments comparing Random Scheduler and Priority Greedy validate the feasibility, reproducibility, and discriminative capacity of RSCMP-Bench. Random Scheduler yields very low Completion Rates of 7.3%, 3.8%, and 1.9% on the Low, Medium, and High levels, respectively. These results confirm the extreme sparsity of the feasible solution space. Priority Greedy achieves higher Completion Rates but still degrades as scenario difficulty increases, decreasing from 76.1% at the Low level to 63.7% at the Medium level and 49.2% at the High level. These findings indicate that high-difficulty scenarios remain challenging even for reasonable heuristic methods. They also show considerable room for more advanced algorithms. The dynamic adaptability protocol quantifies algorithm robustness under unexpected dynamic task insertion, which is not captured by static evaluations. The two-stage constraint verification module rejects infeasible plans and generates detailed error reports to support debugging.  Conclusions   RSCMP-Bench provides a unified, fair, and reproducible benchmark for remote sensing constellation mission planning. By combining a public library of 300 standardized scenarios, a multi-dimensional effectiveness evaluation system based on Basic Performance and Dynamic Adaptability, and a simulation and evaluation platform with realistic constraints and automated scenario generation, the framework addresses the long-standing lack of standardized evaluation in this field. Baseline results confirm its discriminative capacity and reveal clear performance bottlenecks in large-scale dynamic scenarios. Inspired by ImageNet and GLUE, RSCMP-Bench can support systematic community evaluation and fair competition. The framework has been deployed at https://www.tianzhibei.com, and its adoption can accelerate progress in intelligent mission planning for next-generation remote sensing constellations.
Dual-MPC-Driven Modeling and Spatiotemporal Evolution of Intelligent Connected Traffic Risk Fields
JIANG Linyuan, DING Fei, FAN Xuan, YANG Xuechao, SONG Aiguo, ZHANG Dengyin
Available online  , doi: 10.11999/JEIT260194
Abstract:
  Objective  With the deployment of intelligent connected vehicle-road-cloud cooperative systems, roadside infrastructure is evolving from traffic-state sensing units into intelligent decision-support platforms for multi-vehicle interaction analysis and dynamic risk inference. In highway and urban freeway scenarios, traffic operation is affected not only by the kinematic responses of individual vehicles but also by lane-changing intentions, car-following competition, and conflict propagation under local interactions. From a roadside perspective, a unified framework is therefore needed to continuously represent traffic risk, reveal its spatiotemporal evolution, and couple risk information with behavior decision-making and trajectory planning. Existing car-following models, such as the Optimal Velocity Model (OVM), Full Velocity Difference (FVD) Model, and Intelligent Driver Model (IDM), can describe speed-spacing evolution. However, these models mainly focus on longitudinal interactions and usually embed risk implicitly in safety-distance or acceleration constraints. They cannot explicitly characterize the coupling between longitudinal following and lateral lane changing, nor can they provide a continuous risk representation suitable for regional traffic assessment. Although Artificial Potential Field (APF) methods and Model Predictive Control (MPC) methods can improve trajectory safety, existing studies still lack a unified mechanism that links risk assessment, behavior decision-making, and motion planning. In addition, discrete behavior choices and continuous control actions are difficult to process efficiently within a single optimization framework.  Methods  An intelligent connected traffic risk-field model oriented toward vehicle-road cooperation is first established. The model integrates vehicle-interaction risk, lane-marking constraint risk, and road-boundary repulsive risk (Fig. 1). In the vehicle-interaction layer, motion-state-induced risk is formulated by considering the relative speed and relative orientation between the ego vehicle and surrounding vehicles. Distance-induced risk is modeled to reflect attenuation as separation distance increases (Fig. 2(a)(b)). To represent the stronger influence of forward hazards than lateral and rear hazards, a directional non-uniformity coefficient is used. This coefficient adjusts the angular attenuation of field strength and enables anisotropic spatial risk representation around the vehicle. In the road-constraint layer, lane markings and road boundaries are modeled separately. The total driving risk field is obtained by weighting and combining the lane-marking field, road-boundary field, and multi-vehicle interaction field (Fig. 2(e)(f)). Based on this representation, a dual-MPC hierarchical decision and motion-planning architecture is designed (Fig. 3). In each control cycle, the upper layer evaluates candidate behavior modes according to the vehicle state, surrounding traffic state, and dynamic risk field, and then outputs a unique behavior mode. The lower layer activates the corresponding control branch. When lane keeping is selected, longitudinal speed-planning MPC is used. When lane changing is selected, lane-change trajectory-planning MPC is activated under road-boundary, lane-marking, and safe-gap constraints.  Results and Discussions  The proposed framework reconstructs microscopic traffic risk evolution under different datasets and time scales. In the HighD highway scenario, when the slicing interval is 0.4 s, the local evolution of following and lane-changing interactions is captured in detail. This includes the process in which the ego vehicle initially follows a preceding vehicle and then starts changing to the adjacent lane (Fig. 4(a)(d)). When the interval is increased to 1.0 s, a wider spatiotemporal interaction range becomes visible, and lane-change completion and the subsequent return maneuver are identified more clearly (Fig. 4(e)(h)). In the NGSIM scenario, a smaller interval provides a finer description of the lane-change disturbance process. By contrast, a larger interval reveals the wider reconstruction of interaction relationships among the original lane, target lane, and surrounding vehicles (Fig. 4(i)(p)). These results indicate that the proposed roadside-oriented risk field can describe both local interaction details and larger-scale evolution trends, depending on the selected reconstruction interval. Sensitivity experiments further confirm the role of the directional non-uniformity coefficient. As this coefficient increases, forward risk concentration becomes stronger, local peak risk increases, and the coverage of high-risk regions decreases (Table 2). This finding shows that the coefficient effectively regulates anisotropic field distribution. Comparative experiments with IDM, OVM, FVD, and APF show that the proposed method performs better in most representative scenarios and error metrics (Fig. 5, Table 3). In the lateral cut-in scenario, its advantage lies in the early representation of lateral intrusion risk, which enables the behavior decision layer to anticipate conflict and the motion-planning layer to generate continuous evasive actions. In congested scenarios, the superposition of forward congestion risk, lateral neighboring-vehicle influence, and road-boundary constraints allows the dual-MPC controller to evaluate safety and feasibility simultaneously in local space.  Conclusions  A unified framework for roadside-oriented traffic-risk modeling and behavior-driven trajectory planning is developed. By integrating multi-vehicle interaction risk, lane-marking constraint risk, and road-boundary repulsive risk into a continuously evolving dynamic risk field, the spatial quantification of multi-vehicle interaction risk is realized. The directional non-uniformity coefficient further enables asymmetric risk perception modeling in forward, lateral, and rear directions. On this basis, a dual-MPC hierarchical architecture is constructed to couple behavior decision-making with motion planning, so that lane-keeping and lane-changing behaviors can be adaptively selected and optimized under a unified risk-driven mechanism. Experiments based on HighD and NGSIM datasets show that the proposed method can effectively characterize the spatiotemporal evolution of traffic-risk fields and outperform representative comparison models in most typical scenarios and error metrics.
MGM-3DUNet: A Multi-scale Edge Semantic-guided GraphConvolutional Sequence Method for Brain Tumor Segmentation
ZHUANG Jianjun, LI Xiang, JING Shenghua, LÜ Zhenglong
Available online  , doi: 10.11999/JEIT260128
Abstract:
  Objective  Feature fusion in U-Net and its 3D variants mainly relies on simple single-scale concatenation, which limits the use of encoder features and weakens fine-grained segmentation of Tumor Core (TC) and Enhancing Tumor (ET) regions. Recent methods such as VM-UNet improve sequence modeling efficiency, but they mainly focus on global information modeling. Local detail preservation and edge enhancement remain insufficient. Therefore, current methods still have limitations in segmentation accuracy and clinical utility. To address these problems, this paper proposes MGM-3DUNet for brain tumor segmentation.  Methods  The Multi-Scale Edge semantic Guidance Module (MEGM) is designed to improve tumor boundary segmentation through learnable edge detection. The Graph Convolutional Sequence Module (GCSM) combines the local aggregation ability of graph convolution with efficient long-range modeling based on a Mamba-like structure. This design improves semantic consistency while preserving small tumor structures with fewer parameters. The Multi-scale Context Perception Module (MCPM) is introduced to strengthen feature complementarity across different tumor scales through dual-scale fusion.  Results and Discussions   Experiments show that the proposed method achieves better average Dice similarity coefficient (Dice) and 95th percentile Hausdorff Distance (HD95) than the comparison methods. With only 2.3M parameters, MGM-3DUNet achieves Dice values of 91.2%, 90.4%, and 89.2% for Whole Tumor (WT), TC, and ET, respectively. The visualization results (Fig. 9, Fig. 10) further show that MEGM improves boundary localization. Overall, the proposed method shows improved sensitivity to edge details and contextual correlations while maintaining a low parameter count.  Conclusions   This method improves tumor boundary prediction by introducing shallow-layer edge enhancement to emphasize tumor contours. Local and global semantic information is fused in the bottleneck layer, and multi-scale contextual features are integrated during decoding. The proposed design achieves accurate segmentation with low computational cost and is suitable for deployment on resource-constrained platforms.
A Low-latency Synchronization Header Detection Algorithm and Circuit for the JESD204C Interface
YIN Peng, ZHANG Chao, LEI Changan, HOU Weizhou, SHU Zhou, LIU Shubin, ZHU Zhangming
Available online  , doi: 10.11999/JEIT260163
Abstract:
  Objective  With rapid advances in high-speed electronics, front-end Analog-to-Digital Converters and Digital-to-Analog Converters (ADCs/DACs) continue to increase in sampling rate and resolution. Back-end Field-Programmable Gate Arrays and Application-Specific Integrated Circuits (FPGAs/ASICs) also provide stronger computing capability. These trends impose strict requirements on high-speed data interfaces, including high bandwidth, low latency, low power consumption, and reliable synchronization. As a mainstream high-speed Serializer/Deserializer (SerDes) interface, the JESD204C interface still suffers from long link initialization latency and high synchronization power consumption. These limitations restrict system real-time performance and energy efficiency. To address these issues, this study optimizes the link-layer design of the JESD204C receiver and proposes an efficient Synchronization Header (SH) detection method. The method implements exponential compression of the search set through global observation and iterative convergence. Detection efficiency is improved, fast and accurate SH positioning is achieved, link synchronization latency is reduced, and synchronization stability and energy efficiency are enhanced.  Methods  A typical JESD204C interface uses serial sliding detection for SH detection, which causes high link initialization latency and large delay jitter. To solve these problems, an Iterative Set Screening (ISS)-based SH detection algorithm is proposed. The SH detection task is modeled as the rapid localization of a deterministic pattern in a binary random sequence. A theoretical model based on information theory and stochastic processes is constructed. Expected space utilization and Bit Error Rate (BER) are introduced to support quantitative performance evaluation. In this model, SH candidate positions are defined as a dynamic set. Based on the inherent polarity inversion characteristic of the SH and global observations in each clock cycle, multilevel XOR logic is used to verify all candidate hypotheses in parallel. Non-inverting candidate positions are eliminated, and the search space is dynamically compressed. This design improves synchronization speed and position robustness, providing a low-latency and reliable initialization solution for high-speed SerDes links.  Results and Discussions  The proposed ISS-based SH detection algorithm is validated under harsh conditions, including SH crossing block boundaries and loss of lock caused by burst errors. The results demonstrate strong robustness, with rapid SH locking and link resynchronization under all test conditions (Figures 1116). To evaluate performance, four representative schemes are reproduced: a single-bit serial locking circuit, a 66-bit serial locking architecture, a register-intensive block synchronization method, and a parallel search circuit. A systematic comparison is then conducted between these schemes and the proposed design. The results show that the normalized locking time of the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method varies substantially with SH position (Figure 17(a)), especially at block boundaries (Figure 17(b)). When the SH is located at the Most Significant Bit (MSB), typical sliding detection requires about 1.8 times the time needed at the Least Significant Bit (LSB), indicating strong sensitivity to the starting position and search path. In contrast, the proposed ISS scheme maintains a stable normalized locking time within 1.0 ± 0.05 across all positions, with the standard deviation reduced by more than 70%. By evaluating all candidate positions equally through parallel filtering, the scheme eliminates position dependence. Synchronization can be completed within tens of clock cycles whether the SH is located at the LSB, the MSB, or any other position in the block. The experimental results verify that the ISS algorithm improves synchronization robustness and predictability while accelerating link initialization. Table 3 summarizes the performance metrics. The average locking time is only 24.4 clock cycles, representing an overall improvement of more than 70% compared with the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method. The standard deviation of locking time is only 4.7, indicating a more stable synchronization process. In terms of resource utilization, the design consumes 509 Look-Up Tables (LUTs) and only 2.0 mW, much lower than the 3 503 LUTs and 94.1 mW required by the register-intensive scheme. Its energy efficiency reaches 0.03 mW/bit, which is better than those of the three conventional methods. Compared with the parallel search circuit, the average locking time is reduced by 6.11%, power consumption is reduced by 50.3%, and energy efficiency is improved by 53.8%. Therefore, the proposed JESD204C receiver link shows advantages in SH detection speed, stability, power consumption, and energy efficiency.  Conclusions  An ISS-based SH detection algorithm is proposed for the JESD204C receiver. By screening the data stream in parallel through multilevel XOR logic, dynamically compressing the search space, and efficiently eliminating non-inverting candidate positions, the algorithm converges to the true SH position. This approach improves the conventional serial detection mechanism. The design is verified on the Xilinx KC705 FPGA platform. A Pseudorandom Binary Sequence 31 (PRBS31) is used to emulate the random distribution of polarity transitions, and a high-speed SubMiniature version A (SMA) cable is used for data loopback transmission. The results show that the algorithm achieves an average locking time of only 24.4 clock cycles, with a standard deviation as low as 4.7. High robustness is maintained for the SH at any position within the 66-bit block, and the energy efficiency reaches 0.03 mW/bit. The algorithm is superior to existing typical schemes in locking speed, delay stability, and energy efficiency. It provides a low-latency, reliable, and energy-efficient synchronization initialization approach for high-speed SerDes links.
Survey on Intelligent Semantic Covert Communication
FENG Zhaoxin, XU Yifan, XING Chengwen, XU Yuhua, ZHAO Nan, WANG Jinlong
Available online  , doi: 10.11999/JEIT260184
Abstract:
  Significance   As the Sixth-Generation mobile communication network (6G) evolves from the Internet of Everything to the Intelligent Internet of Everything, the communication paradigm is shifting from reliable bit transmission to effective semantic transmission. Semantic communication extracts and compresses task-related semantics to reduce redundancy and resource use. However, because semantic information is highly structured and task-specific, it is vulnerable to eavesdropping, inference, and attacks. Covert communication addresses this risk by hiding transmission behavior from unauthorized monitoring. With support from Artificial Intelligence (AI), covert communication can use reinforcement learning to adjust power and resource allocation in dynamic environments. Generative models can also conceal transmitted signals by learning and reproducing environmental patterns. However, strict covertness constraints limit the achievable transmission rate and make large-scale information transmission difficult. Intelligent semantic covert communication integrates semantic extraction with covert transmission, providing a reliable approach to secure and efficient 6G communications.  Progress   With the development of AI, especially deep learning for complex feature modeling, semantic communication can support efficient semantic extraction and nonlinear compression of multimodal data. Research on semantic communication has also shifted from Separate Source-Channel Coding (SSCC) to Joint Source-Channel Coding (JSCC), which supports end-to-end training and improved transmission performance. For image transmission, Convolutional Neural Networks (CNNs) use local receptive fields to capture spatial correlations. For sequential data transmission, Long Short-Term Memory (LSTM) networks use gating mechanisms to maintain temporal coherence. In covert communication, Generative Adversarial Networks (GANs) and diffusion models can learn the statistical patterns of environmental noise in the time, frequency, and spatial domains, thereby concealing transmitted signals. These methods reduce the effectiveness of unauthorized monitoring and detection, and improve system adaptability in dynamic environments. AI also improves autonomous decision-making in dynamic covert communication. By modeling covert transmission as a Markov Decision Process (MDP), Deep Reinforcement Learning (DRL) can learn resource allocation strategies through interaction with the environment. This approach reduces computational complexity compared with traditional convex optimization methods. By integrating semantic extraction and covert transmission, intelligent semantic covert communication further supports semantic-driven covert transmission. Large Language Models (LLMs) can evaluate semantic sensitivity and contextual risks, enabling selective covert transmission of sensitive semantic information.  Conclusions  Research on intelligent semantic covert communication shows the advantages of coordinated semantic perception and physical-layer covert mechanisms. AI improves semantic extraction efficiency and strengthens adaptation to dynamic and complex environments. By integrating semantic understanding with covert transmission strategies, intelligent semantic covert communication supports both efficiency and security for ubiquitous 6G services.  Prospects   Future research on intelligent semantic covert communication should address several key challenges, including AI-enabled detection, unified semantic metrics, lightweight model design, multimodal semantic alignment, system interpretability, and semantic hallucination. Active threat detection and adaptive defense strategies are needed to counter AI-driven surveillance. Causal reasoning in Large Multimodal Models (LMMs) can help mitigate semantic hallucination and improve data transmission reliability. Advances in model compression and cloud-edge collaboration are also needed to deploy high-complexity AI models on resource-limited terminals. With the rapid development of AI, intelligent semantic covert communication is expected to provide core support for intelligent connectivity of everything and help build more secure, efficient, and reliable 6G networks.
An Inverse-Hybrid-Modeling Digital Twin System for Natural Gas Energy Metrology
LIU Bin, ZHONG Lu, FENG Quanyuan, CHEN Yihong
Available online  , doi: 10.11999/JEIT260289
Abstract:
  Objective  Global natural gas consumption continues to increase at an average annual rate of 3.2%. A 0.1% reduction in energy measurement error can reduce trade disputes by approximately $750 million per year. Traditional studies mainly use indirect methods for energy measurement. Among these methods, chromatographic analysis and acoustic velocity correlation are the most widely used, but both have clear application limits. Chromatographic analysis has a low interference error, but it shows delayed dynamic response at high flow rates and limited dynamic calibration capability. It also has poor adaptability to multi-gas-source switching, requires manual calibration, and has high operation and maintenance costs. The lack of interoperability standards for energy networks further increases the difficulty of system integration. Acoustic velocity correlation provides a low-latency dynamic response for flow measurement, but it has a high interference error. This error may increase when the content of a single component changes, such as when the hydrogen content increases from 5% to 10%. The method may even fail under complex operating conditions, such as multi-gas-source mixing and dynamic pressure fluctuations. To address these issues, new mechanism-modeling-oriented methods have been developed. The two most representative directions are mechanism-modeling-driven methods and hybrid-modeling methods. Both methods combine multi-source data fusion with virtual-physical interaction to establish mechanism models that link flow rate, other parameters, and energy. These methods provide a new approach for accurate energy measurement, but new challenges remain. Mechanism-modeling-driven methods are usually based on static flow modeling using Computational Fluid Dynamics (CFD). However, their dynamic parameter updates are slow, with delays of more than 30 s. They also have difficulty adapting to real-time operating-condition changes, rely on large labeled datasets, and have limited interpretability. Hybrid-modeling methods still face unresolved problems in collaborative optimization across multiple modules. In addition, existing studies lack support from industrial-grade verification platforms. These limits restrict their ability to solve the dynamic response delay, parameter identification difficulty, excessive physical simplification, and weak interference resistance of traditional natural gas energy metrology methods under complex conditions. Based on recent progress in mechanism-modeling-driven and hybrid-modeling methods, this study proposes an inverse-hybrid-modeling-driven digital twin system. The system introduces a Variational AutoEncoder (VAE)-based operating-condition feature extraction algorithm and a Dynamic Bayesian Network (DBN)-based parameter calibration mechanism. It also uses a Variational Expectation-Maximization (VEM) algorithm for offline calibration. The proposed system aims to improve the accuracy, adaptability, and interference resistance of natural gas energy metrology under complex operating conditions.  Methods   A natural gas energy metrology digital twin system based on inverse hybrid modeling is proposed. The system is built on a three-tier “algorithm-system-scenario” architecture. It integrates calorific value, flow, and energy mechanism models with multi-source real-time data streams. The VAE is used for unsupervised mining of operating-condition features. A parameter self-correction loop is then constructed by combining the DBN with VEM-based system calibration. Industrial-grade devices, including ultrasonic flowmeters and gas chromatographs, are integrated to ensure real-time data transmission and closed-loop control. The system covers key operating conditions, including dynamic pressure fluctuations, hydrogen-blended gas mixtures, and multi-gas-source switching. This design ensures strong adaptability between the model and practical applications. The system was continuously verified for 25 weeks on a full-scale industrial-grade experimental platform. The results show an operational delay of ≤3.8 s, data transmission jitter of ≤0.5 s, average daily energy consumption per device of ≤1.2 kW·h, Mean Time Between Failures (MTBF) of ≥4 100 h, energy measurement error of ≤0.25%, calorific value error of ≤0.12%, and flow indication error of ≤0.2%. The system also meets security requirements through industrial Ethernet encryption and hierarchical access control. It provides engineering support for intelligent pipeline-network optimization and standardized integration.  Results and Discussions  First, a multi-level hybrid modeling framework is established. Modular hybrid modeling is achieved through the algorithm-system-scenario three-tier architecture. Numerical methods combined with data are more flexible than purely analytical models and can represent complex multiphysics systems with fewer lumped physical parameters. These parameters may change during energy measurement under mechanical, energy, and hydrodynamic effects. The VAE and DBN are used to deeply integrate mechanism models with real-time data. This reduces the parameter synchronization delay to 3.8 s and supports fluid-acoustic co-simulation and rapid response under complex operating conditions, such as hydrogen-blended natural gas. Second, an integrated algorithm for inverse hybrid modeling and system calibration is proposed. By incorporating the VAE, DBN, and VEM algorithm, the inverse hybrid modeling algorithm forms a self-supervised, adaptive intelligent system with an internal closed-loop operation. The VAE encoder compresses high-dimensional operating-condition data into low-dimensional feature vectors. This enables unsupervised feature extraction without large labeled datasets. Based on the learned internal data distribution, the VAE can also generate perturbed data similar to the input data. These data are used to simulate abnormal operating conditions and verify interference resistance. The DBN constructs a continuous “prior-evidence-posterior” iterative cycle to support system self-correction and adaptive response to operating-condition changes. The VEM algorithm compensates for systematic errors that are difficult for the DBN to capture, thereby overcoming the limits of traditional static models.  Conclusions  This study describes and validates a hybrid digital twin system that combines experimental data-driven methods with physical models. The system successfully simulates the physical characteristics of natural gas energy metrology. A full-scale test platform was constructed, and the main system parameters were validated using experimental measurement data and compared with industry benchmarks. Each independent module in the algorithm-system-scenario three-tier hybrid modeling architecture, including calorific value measurement, flow calculation, and energy conversion, was continuously verified for 25 weeks. The results confirm strong consistency between model predictions and actual measurements. On the natural gas energy metrology digital twin experimental platform, systematic validation was performed for three core functions: flow measurement under dynamic conditions, multi-component calorific value determination, and energy accumulation. The results show that the output of the digital twin model matches the physical device measurement data with an accuracy of more than 99.5%. Under complex operating conditions, such as pressure pulsations and hydrogen-blended gas mixtures, the system maintains the measurement error within 0.5%. This performance is better than that of traditional methods and meets the Class A accuracy requirements for natural gas measurement. By introducing a multi-tier hybrid modeling framework, this study addresses the parameter identification difficulty and excessive physical simplification of traditional natural gas energy metrology methods. The integration of the VAE, DBN, and VEM algorithm enables unsupervised feature extraction under complex operating conditions and adaptive calibration of model parameters. This reduces dependence on prior physical knowledge and large labeled datasets. The experimental results show that the proposed method maintains high precision and strong stability under complex scenarios, including pressure pulsations and hydrogen-blended gas mixtures, where traditional models have difficulty providing accurate descriptions.
Non-Terrestrial Network Architecture and Key Technologies for Civil Aviation
LIU Xiangnan, QIU Yu, HUANG Zhipeng, ZHANG Haijun
Available online  , doi: 10.11999/JEIT260348
Abstract:
  Significance   Civil aviation communication systems are entering a new stage of development driven by the rapid growth of global air transportation, the increasing demand for intelligent air traffic management, and the continuous expansion of in-flight connectivity services. Traditional civil aviation communication systems mainly rely on high frequency radio, high frequency radio, terrestrial air-to-ground links, and conventional satellite communication systems. These technologies have supported aircraft operation, air traffic control, airline operational communication, and low-rate data transmission for a long time. However, they still face limitations when applied to future civil aviation scenarios characterized by global coverage, high-speed mobility, low latency, high reliability, and service diversification. Particularly, terrestrial networks are difficult to deploy in transoceanic routes, polar regions, deserts, mountains, and remote airspace, while traditional geostationary satellite systems suffer from large propagation delay and limited capacity. Current systems cannot fully meet the requirements of continuous aircraft access, real-time flight monitoring, engine health data transmission, aviation safety communication, and passenger broadband services. Non-Terrestrial Networks (NTNs) provide a promising technical path for overcoming these limitations. By integrating GEOstationary satellites (GEO), Medium Earth Orbit satellites (MEO), low Earth orbit satellites (LEO), Very Low Earth Orbit satellites (VLEO), High-Altitude Platform Stations (HAPS), Unmanned Aerial vehicles (UAV), electric Vertical Take Off and Landing (eVTOL), and terrestrial infrastructures, NTN can construct a multi-layer air-space-ground integrated communication system. Such a system is able to provide continuous coverage, flexible deployment, resilient connectivity, and differentiated service support for civil aviation. NTN is becoming an important enabling technology for future civil aviation communication systems and for the digital and intelligent transformation of the aviation industry.  Progress   This paper reviews the development of NTN technologies for civil aviation and summarizes key research progress from three aspects: network architecture, access and mobility management, and resource management and scheduling. (1) We propose an aviation-oriented NTN networking framework composed of three layers: the satellite edge layer, the airborne core layer, and the terrestrial assistance layer. The satellite edge layer includes GEO, MEO, LEO, and VLEO satellites connected through inter-satellite links. GEO satellites are suitable for wide-area broadcasting and non-real-time services, MEO satellites can support navigation and intermediate-delay services, LEO satellites are suitable for low-latency and high-capacity broadband access, and VLEO satellites can further reduce propagation delay for future near-real-time aviation applications. The airborne core layer includes civil aircraft, HAPS, UAVs, and eVTOL platforms. HAPS can act as a regional relay, edge computing node, or software-defined control carrier, while UAVs and eVTOL platforms can provide flexible low-altitude coverage, emergency communication, and local access support. The terrestrial assistance layer consists of terrestrial base stations and gateway stations, which support air-to-ground communication and satellite-terrestrial interconnection. Civil aviation services can be divided into air traffic control and air traffic management services, airline operational control services, and airline passenger communication or in-flight entertainment services. Through network slicing, these heterogeneous services can be logically isolated and managed over a shared air-space-ground infrastructure. In congestion, rain attenuation, or shortened visibility-window scenarios, safety slices should be protected with the highest priority, while passenger service slices can be rate-limited, buffered, or degraded. (2) We analyze the characteristics of NR-NTN access and air-to-ground direct access in civil aviation. NR-NTN can provide continuous coverage for oceanic, polar, desert, and remote flight routes through satellites or HAPS, while air-to-ground direct access can provide low-latency and high-rate links in areas where terrestrial base stations can be deployed. However, aircraft differ significantly from ordinary terrestrial terminals because their flight trajectory, altitude, speed, and route are highly predictable. Therefore, the key issue in aviation NTN access is not only how to execute random access, but how to predict the access window, timing compensation, frequency offset, and target access node before the aircraft enters the coverage area. By using satellite ephemeris, Global Navigation Satellite System information, aircraft trajectory, and velocity parameters, civil aircraft can predict satellite visibility and pre-compute timing advance, scheduling offset, and Doppler compensation before initiating access. This transforms random access from a passive response process into a proactive and predictive access process, thereby improving access certainty and synchronization stability in highly dynamic aviation scenarios. For mobility management, a signaling interaction process for aircraft handover is designed. Based on trajectory prediction and satellite visibility prediction, the network can select a target satellite or gateway with longer residence time and better service capability. Before the aircraft reaches the handover boundary, the source and target network sides can complete context preparation, user-plane path preparation, radio resource reservation, and protocol data unit session update. When the handover condition is triggered, the aircraft performs random access to the target satellite or beam and then switches the user-plane path. This “prediction–preparation–fast handover” mechanism can reduce service interruption and maintain session continuity. For safety-critical traffic, priority and isolation policies should remain consistent and auditable throughout session preparation, handover execution, and path switching. (3)We discuss computing and caching resource management in civil aviation NTN. As onboard computing capability is limited and aviation applications generate increasing computing demands, NTN can provide mobile edge computing and caching services through LEO satellites, HAPS, UAVs, and inter-satellite cooperation. The paper introduces several computing offloading modes, including on-orbit satellite collaborative offloading, network-level integrated offloading, and cloud-edge-terminal hybrid offloading. These mechanisms can support tasks such as aviation monitoring, trajectory analysis, intelligent inference, and in-flight service optimization. In addition, caching mechanisms such as onboard satellite caching, inter-satellite cooperative caching, and named-data-networking-based content caching can improve content delivery efficiency and service continuity. Cache placement should consider content popularity, regional demand prediction, visibility windows, cache prefetching, and cooperative cache sharing among different satellite layers.  Conclusions   NTN can effectively complement traditional civil aviation communication systems by filling coverage gaps in remote and oceanic airspace, enhancing service continuity, and supporting differentiated aviation services. The proposed aviation-oriented NTN architecture integrates multi-orbit satellites, HAPS, UAVs, civil aircraft, and terrestrial infrastructures into a unified framework. The on-demand isolated slicing mechanism can provide differentiated protection for ATC/ATM, AOC, and APC/IFE services. Ephemeris-map-assisted access and predictive mobility management can improve access reliability and reduce handover interruption in high-speed aviation scenarios. Computing offloading and cooperative caching further enhance the ability of NTN to support intelligent and data-intensive aviation applications.  Prospects   Future civil aviation NTN should evolve toward deeper integration of low-altitude networks, space networks, and terrestrial networks. Cross-domain topology visualization, link-state sharing, policy distribution, and programmable logical networks are essential for improving controllability and scalability. In mobility management, integrated cross-domain handover mechanisms should be developed to cope with satellite beam switching, terrestrial cell handover, and air-to-air relay reconstruction. In resource management, communication, navigation, computing, and caching resources should be jointly scheduled and transformed according to aviation service requirements. With continuous advances in NTN architecture, network slicing, predictive access, mobility management, computing offloading, and caching, NTN is expected to provide more efficient, stable, and intelligent communication support for civil aviation and to promote the digital transformation of future air transportation systems.
Research on Secure and Covert Transmission for UAV-assisted Visible Light Communication Systems
WU Mengru, LIN Jiale, LU Weidang, LI Bo, GUO Lei
Available online  , doi: 10.11999/JEIT260239
Abstract:
  Objective  Unmanned Aerial Vehicles (UAVs) can serve as aerial base stations for Visible Light Communication (VLC) because of their mobility and on-demand coverage capabilities. However, air-ground communication links are exposed to open environments, which makes VLC vulnerable to data eavesdropping and malicious detection. To address this issue, this paper proposes a secure and covert transmission strategy for a UAV-assisted VLC system from the perspectives of Physical Layer Security (PLS) and Covert Communication. The proposed strategy jointly optimizes UAV transmit power and hovering altitude to maximize the system secrecy capacity. The optimization is subject to covert communication requirements, illumination requirements, and operational constraints on UAV transmit power and hovering altitude.  Methods  This paper investigates secure and covert communication in a UAV-assisted VLC system. A UAV-assisted VLC system model is first established. In this model, a mobile UAV equipped with a Light-Emitting Diode (LED) is used to establish a VLC link with a legitimate ground user in the presence of an eavesdropper (Eve) and a warden (Willie). An optimization problem is then formulated to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude. To solve this problem, a Two-Layer OPtimization (TLOP) algorithm is proposed. The transformed problem is decomposed into two subproblems: an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude design problem. A closed-form expression for the optimal transmit power is derived for the inner-layer problem. A Particle Swarm Optimization (PSO) algorithm is then developed to solve the outer-layer problem.  Results and Discussions  In the simulations, the proposed optimization scheme is compared with two baseline schemes. First, the convergence of the proposed TLOP algorithm is verified (Fig. 3). The results show that the algorithm converges rapidly within a limited number of iterations. Second, the optimal UAV hovering altitude with respect to the UAV horizontal coordinates is illustrated under the spatial distribution (Fig. 4). The results indicate that the optimal hovering altitude decreases as the UAV approaches the legitimate ground user. The secrecy capacity with respect to the UAV horizontal coordinates is then presented (Fig. 5). The secrecy capacity increases as the UAV approaches the legitimate ground user. This is because the legitimate VLC channel gain increases when the UAV is closer to the user. In contrast, when the UAV approaches Eve and Willie, the security and covertness constraints become stricter. The UAV is then forced to reduce its transmit power or increase its hovering altitude, which decreases the system secrecy capacity. Furthermore, the secrecy capacity of all schemes increases as ϵ increases (Fig. 6). This is because a larger ϵ relaxes the covertness requirement. The UAV can therefore adjust its hovering altitude and transmit power more flexibly to increase the system secrecy capacity. In addition, the secrecy capacity decreases as the number of symbols increases (Fig. 7). This occurs because more symbols provide Willie with more signal samples for detection, thereby improving Willie’s detection capability. Finally, the secrecy capacity of all schemes decreases as the uncertainty-region radius of illegal nodes increases (Fig. 8). This trend occurs because greater location uncertainty forces the UAV to address potential threats over a wider area. The UAV must therefore adopt a more conservative strategy under worst-case eavesdropping and detection conditions. Overall, the simulation results confirm that the proposed scheme improves the secrecy capacity of the UAV-assisted VLC system.  Conclusions  This paper investigates secure and covert communication in a UAV-assisted VLC system. The objective is to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude under covert communication, illumination, transmit power, and hovering altitude constraints. Because the formulated problem is highly non-convex, a PSO-based TLOP algorithm is designed to solve it. The proposed algorithm decomposes the problem into an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude optimization problem. Simulation results show that the proposed algorithm converges rapidly and improves the system secrecy capacity compared with the baseline schemes.
Energy-Efficient Trajectory Planning and Resource Optimization for UAV Relay Communications over Hybrid RF/FSO Links
LI Baolong, PAN Wenwei, JIANG Hao, FENG Simeng, WU Qihui
Available online  , doi: 10.11999/JEIT260139
Abstract:
  Objective  In low-altitude communication networks, hybrid Radio Frequency/Free-Space Optical (RF/FSO) Unmanned Aerial Vehicle (UAV) relaying can ease RF spectrum congestion and improve uplink data aggregation. However, in obstacle-rich urban environments, FSO backhaul links are vulnerable to blockage and intermittent outages. This creates a severe mismatch between the RF access-link rate and the FSO backhaul-link rate. UAV trajectory planning is also constrained by obstacle avoidance and flight dynamics. To address these coupled issues, this paper investigates an energy-efficiency maximization problem. Multiuser Non-Orthogonal Multiple Access (NOMA)-based RF access and the Three-Dimensional (3D) obstacle-avoiding UAV trajectory are jointly optimized, and buffer-assisted RF/FSO rate decoupling is incorporated.  Methods  A time-slotted UAV relaying model is considered, in which multiple ground users upload data to the UAV through an RF link using NOMA. The UAV decodes superposed signals by Successive Interference Cancellation (SIC), and the decoding order in each slot is determined according to the received-power ranking. The successfully received data are then forwarded to a Base Station (BS) through an FSO backhaul link. Urban blockage is modeled using 3D geometric obstacles. A visibility test is used to determine whether each relevant link is in Line-Of-Sight (LOS) or Non-Line-Of-Sight (NLOS), which captures the spatially correlated and time-varying RF access-link rate and intermittent FSO backhaul capacity. To suppress blockage-induced rate mismatch between the RF access link and the FSO backhaul link, an onboard finite-capacity buffer is deployed at the UAV. In each slot, the forwardable data amount is jointly limited by the instantaneous FSO backhaul capacity and the data available in the buffer, and buffer-capacity constraints are imposed to prevent overflow. System energy efficiency is defined as the ratio of cumulative data successfully delivered to the BS over the mission horizon to UAV propulsion energy consumption. Propulsion power is modeled as a function of UAV velocity and acceleration to reflect the effect of flight dynamics. Under 3D flight-region boundaries, prescribed start and end locations, discrete-time kinematic equations, maximum velocity and acceleration limits, and obstacle collision-avoidance constraints, a non-convex optimization problem is formulated. The decision variables are cross-slot multiuser transmit powers and the 3D UAV trajectory. An alternating optimization framework is then developed. For a fixed trajectory, propulsion energy is fixed, so maximizing energy efficiency is equivalent to increasing end-to-end successfully forwarded data. This yields a power-optimization subproblem. Because of NOMA coupling and logarithmic rate expressions, this subproblem remains non-convex and is solved by Successive Convex Approximation (SCA). For fixed transmit powers, Particle Swarm Optimization (PSO) is used to search candidate 3D trajectories in continuous space. To ensure feasibility under strict dynamics and safety constraints, Quadratic Programming (QP) projection is used to enforce velocity and acceleration constraints. Collision checks are performed for trajectory waypoints and inter-slot line segments to ensure obstacle-free flight. These two optimization procedures are performed alternately. The resulting joint design satisfies flight-dynamics feasibility and collision-avoidance requirements and improves energy efficiency.  Results and Discussion   Simulations are conducted in an urban airspace with multiple users, a BS, and dense 3D obstacles. Blockage causes frequent LOS/NLOS switching as the UAV moves. Fig. 2 and 3 compare the 3D trajectory and its planar projection, respectively. Compared with the initial trajectory, the optimized trajectory shows clear detours and necessary altitude adjustments. It achieves collision-free flight while satisfying velocity and acceleration constraints, thereby verifying the feasibility and safety of the proposed trajectory planning method. Fig. 4 shows the convergence of energy efficiency under different user transmit-power budgets. The proposed alternating optimization generally stabilizes within a small number of outer iterations. The converged energy efficiency increases with the power budget, indicating synergy between power control and trajectory adaptation. Fig. 5 shows buffer evolution over time. The buffer gradually accumulates data when the backhaul is blocked or experiences strong fading. It is quickly drained when the UAV enters regions with LOS backhaul and improved FSO capacity. To quantify buffering gain, Fig. 6 compares system energy efficiency between the proposed buffering mechanism and the no-buffer scheme. The proposed mechanism enables store-and-forward temporal smoothing during backhaul interruptions and improves system energy efficiency. Fig. 7 shows energy-efficiency convergence under different buffer capacities. As buffer capacity increases, the converged energy-efficiency level improves. A larger buffer enhances the UAV’s ability to temporarily store incoming data and reduces data accumulation and transmission blockage when RF access-link and FSO backhaul-link rates are mismatched or the backhaul link is constrained. Figure 8 compares four benchmark schemes, namely a non-optimized baseline, a power-optimization scheme, a trajectory-optimization scheme, and the proposed joint power-and-trajectory optimization scheme. The coordinated design of power allocation and obstacle-avoiding trajectory improves end-to-end energy efficiency. Trajectory optimization also plays a more dominant role under blockage-limited conditions.  Conclusion  This paper investigates a hybrid RF/FSO UAV relaying scheme with NOMA and an onboard buffering mechanism for low-altitude urban communication. Given dense obstacles, frequent blockage, FSO-link susceptibility, and strict flight-dynamics constraints, an energy-efficiency maximization problem is formulated for the joint optimization of multiuser NOMA power allocation and UAV trajectory. An SCA-based power-allocation method and an obstacle-avoiding trajectory design that combines PSO with QP projection are developed. The obtained trajectory satisfies flight-dynamics feasibility and collision-avoidance requirements and improves throughput per unit propulsion energy. Simulation results show that the planned trajectory can avoid obstacles, and that the onboard buffer provides an effective cushion between RF access and FSO backhaul to mitigate rate mismatch. The proposed method consistently outperforms benchmark schemes in energy efficiency. Trajectory optimization is also shown to be generally more effective than power allocation in improving overall system performance.
Joint Optimization Method for Pairwise Constrained Projection Clustering Integrating a Two-row Simultaneous Update Strategy
ZHU Jianyong, CHEN Kun, YANG Hui, NIE Feiping
Available online  , doi: 10.11999/JEIT260111
Abstract:
  Objective  As data structures become increasingly complex, conventional unsupervised clustering methods often fail to achieve satisfactory performance. Semi-supervised clustering has therefore attracted growing attention because it uses limited prior information to improve clustering quality. However, existing methods have two major limitations. First, traditional constrained projection clustering algorithms usually use a two-step independent strategy, in which the projection matrix is learned before k-means clustering is performed. This separation allows projection errors to be propagated directly to the clustering stage, causing accumulated learning errors. In addition, applying pairwise constraints only during projection deviates from the goal of using prior information to guide clustering. Second, many existing methods, including spectral clustering-based approaches, handle pairwise constraints implicitly, for example through eigen-decomposition of a modified similarity matrix. Such implicit processing may not strictly satisfy the constraints, especially Cannot-Link (CL) constraints, which are non-transitive, resulting in high constraint violation rates. To address these issues, this paper proposes a joint optimization method for pairwise constrained Projection Clustering Integrating a Two-row simultaneous Update Strategy (PCITUS). The objective is to unify dimensionality reduction and clustering within a single framework to reduce information loss, while designing an explicit optimization strategy that lowers constraint violations and improves computational efficiency.  Methods  The proposed PCITUS model integrates constrained projection and clustering into a unified objective function for collaborative optimization, with pairwise constraints optimized directly. First, the algorithm uses the transitive property of Must-Link (ML) constraints. Samples belonging to the same ML connected component are merged into a single hyper-point in the feature space. This preprocessing step ensures that all ML constraints are naturally satisfied. A trade-off parameter is then introduced to incorporate projection learning into the clustering framework as a regularization term, allowing both components to be jointly optimized under one objective. Prior information is further embedded into the clustering process by transforming pairwise constraints into row-wise constraints on the indicator matrix. An improved coordinate descent method is then used to optimize the discrete indicator matrix directly, which improves computational efficiency and produces better clustering results. A key feature of PCITUS is the two-row simultaneous update strategy for CL constraints. PCITUS explicitly checks CL conflicts by simultaneously evaluating objective function values obtained by moving conflicting rows to suboptimal classes and then selects the case with the higher value.  Results and Discussions  Extensive experiments are conducted on eight benchmark datasets and compared with nine state-of-the-art semi-supervised clustering algorithms. Quantitative results based on ACCuracy (ACC) and Normalized Mutual Information (NMI) demonstrate the superiority of PCITUS (Table 4 and Table 5). PCITUS achieves the best performance on most datasets. In particular, on the Mushroom dataset, NMI is improved by 7.29% compared with the second-best algorithm. The comparison with CNP, a two-step projection method, confirms that the unified framework effectively reduces error propagation and information loss. This effect is also supported by the mutual reinforcement between projection and clustering: a better projection space produces a clearer clustering structure, while a more reasonable clustering structure guides the formation of a more discriminative projection space. The effectiveness of explicit constraint handling is further illustrated (Fig. 1). PCITUS produces no ML constraint violations because of the hyper-point merging strategy. For CL constraints, the two-row simultaneous update strategy enables PCITUS to maintain an extremely low violation rate, such as 0.57% on Mushroom and 0.41% on Satimage, greatly outperforming methods that handle constraints implicitly. Additionally, the parameter sensitivity analysis (Fig. 2) shows that PCITUS remains stable across a wide range of trade-off parameter values. The noise sensitivity experiments (Fig. 3a and Fig. 3b) confirm its robustness. The convergence curves (Fig. 3c and Fig. 3d) and runtime comparisons (Table 7) further verify its computational efficiency, showing rapid convergence and a stable objective function value within approximately 10 iterations in most cases.  Conclusions  This paper presents PCITUS, a semi-supervised clustering framework that jointly optimizes pairwise constrained projection and clustering structures. The method addresses the difficulty of optimizing CL constraints and overcomes the limitations of traditional constrained projection clustering frameworks based on a two-step separation scheme. By integrating the projection objective into the clustering framework as a regularizer, the proposed method enables subspace learning and data partitioning to reinforce each other and jointly approach the global optimum. Pairwise constraints are used throughout the learning process, allowing prior knowledge to guide optimization more fully. The coordinate descent method with the two-row simultaneous update strategy directly and accurately allocates samples under CL constraints, significantly reducing constraint violations. Experimental results show that PCITUS outperforms existing algorithms in clustering performance.
Robust Optimization of Low-altitude Communication and Computation Resources in Uncertain Environments
GONG Yucheng, LI Bin, WANG Xinyi, FEI Zesong
Available online  , doi: 10.11999/JEIT260090
Abstract:
  Objective  Low-altitude edge computing networks provide flexible computing services and extended coverage for user equipment. However, quality of service is often degraded by uncertainty in task data size and by Unmanned Aerial Vehicle (UAV) position jitter caused by environmental disturbances. Existing robust methods commonly rely on deterministic uncertainty sets, which tend to be conservative and cannot accurately describe the stochastic distribution of task demands. To address these challenges, a robust energy minimization framework is proposed for multi-UAV-assisted Mobile Edge Computing (MEC) networks. The objective is to minimize the weighted sum of system energy consumption. This is achieved by developing a joint optimization model that coordinates UAV flight trajectories, task splitting decisions, and computation and communication resource allocation. The model explicitly accounts for the dual uncertainties of task data size and UAV trajectory.  Methods  To handle the nonconvexity and strong coupling among optimization variables, the problem is first modeled as a Markov Decision Process (MDP). A comprehensive state space is defined to characterize real-time system dynamics, and a continuous action space is designed for trajectory control and resource management. A Distributionally Robust Optimization Soft Actor-Critic (DRO-SAC) algorithm is then developed to solve the MDP. In this framework, an ambiguity set based on the L1-norm distance is constructed to characterize the distributional uncertainty of the task demand distribution. A maximum-entropy reinforcement learning mechanism is used to learn an optimal policy under the worst-case distribution within the ambiguity set. In this way, UAV trajectories, task splitting, and computation and communication resource allocation are jointly optimized to improve system robustness under dynamic environmental fluctuations.  Results and Discussions  The performance of the proposed DRO-SAC algorithm is evaluated through simulations. DRO-SAC achieves faster convergence and higher rewards than Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO) algorithms (Fig. 3). For energy consumption, the proposed method consistently achieves higher efficiency under different user densities (Fig. 4). The robustness of the system against position errors is also verified, with energy fluctuations kept at a low level (Fig. 5). Dynamic trajectory adjustment further confirms that the proposed method can provide effective user coverage while reducing system energy consumption (Fig. 6).  Conclusions  A DRO-SAC-based joint optimization framework is proposed to address uncertainty in task data size and UAV position jitter in multi-UAV-assisted MEC networks. By constructing an ambiguity set for the task demand distribution and optimizing the worst-case expected objective, the proposed method mitigates the limitations of traditional deterministic models in dynamic environments. Weighted system energy consumption is minimized while latency and safety constraints are satisfied. Simulation results demonstrate that the proposed scheme achieves stable convergence and high energy efficiency, even when communication and computation resources are limited and environmental parameters fluctuate strongly.
A Joint Source-Channel Coding Modulation Scheme for the Transmission of Gaussian Sources
LV Yaping, MA Xiao
Available online  , doi: 10.11999/JEIT251224
Abstract:
  Objective  The Separated Source-Channel Coding (SSCC) scheme has been proven, which will not incur performance loss as long as the source block length goes to infinity. However, the SSCC scheme usually leads to a large buffer and long delay, and may cause error propagation in case of a single symbol error incurred in the communication channel. In order to alleviate these issues, Joint Source Channel Coding (JSCC) schemes have been investigated to transmit Gaussian sources. In this paper, a JSCC modulation scheme for the transmission of Gaussian sources is proposed, and a Gaussian source reconstruction scheme and its reconstruction expression are provided.  Methods  In this paper, the Gaussian source sequence is quantified as a sequence of M-ary symbols by a Lloyd-Max quantizer. For the M-ary quantization symbol sequence, the matching M-ary Fourier Transform Pair (FTP) code is constructed, and the modulation mode adopts the corresponding M-ary Pulse Amplitude Modulation (M-PAM). In particular, the modulated M-ary symbol sequences are transmitted in a block Markov superposition way, which constructs the Block Markov Superposition Transmission FTP (BMST-FTP) code. In addition, in order to obtain the shaping gain, the constellation Geometry Shaping (GS) scheme is also proposed. For the proposed source reconstruction scheme, the system output is the weighted average of the representative elements of the Lloyd-Max quantizer, which replaces the representative elements.  Results and Discussions  The simulations are conducted over M-PAM modulated AWGN channels using GF(3) and GF(5) BMST-FTP codes. For FTP codes employing random mapping, the WER approaches the Union Bound (UB) at high SNR. Similarly, the FTP codes with m repeated transmissions exhibit WER performances that approach the corresponding UBs. Furthermore, the WER performance of BMST-FTP codes with memory m matches UBs in the high SNR region (Fig. 6). For Symbol Error Rate (SER), the GF(3) BMST-FTP code outperforms the GF(5) BMST-FTP code (Fig. 7(a)). For the GF(5) BMST-FTP code, the GS can achieve an SER performance gain of approximately 0.3 dB (Fig. 8(a)). In terms of distortion performance, the GF(3) BMST-FTP code outperforms the BMST-FTP code in the low SNR region, whereas in the high SNR region, the GF(5) BMST-FTP code performs better (Fig. 7(b)). Furthermore, compared with other work, the GF(3) BMST-FTP code with m=1 has a similar performance, and the GF(5) BMST-FTP code with m=1 performs better (Fig. 7(b)).  Conclusions  This work has proposed a joint source-channel coding modulation scheme for the transmission of Gaussian sources. In the proposed scheme, two types of BMST-FTP codes were constructed, each matched with a corresponding Lloyd-Max quantizer and M-PAM modulator. Additionally, a Gaussian source reconstruction scheme and its reconstruction expression were provided. Simulation results demonstrate that the appropriate transmission scheme can be selected according to the aim performance. The proposed GS scheme can obtain a gain of SER of about 0.3dB, which can improve the distortion performance of the waterfall area.
A Lightweight True Random Number Generator Based on Chain-Coupled Oscillation Rings
ZHANG Yuan, YING Haixuan, GAO Kai, YE Jin, WANG Shuang, ZHANG Jiliang
Available online  , doi: 10.11999/JEIT260377
Abstract:
  Objective  With the rapid growth of the Internet of Things, 5G/6G, and satellite Internet, resource-constrained devices increasingly require high-quality random numbers for key generation, authentication, masking, and other security functions. Although pseudo-random number generators are efficient, their outputs may be predictable once the seed or internal state is compromised. True random number generators (TRNGs) offer a hardware root of trust by extracting entropy from physical randomness, but many existing designs rely on multiple entropy sources or complex post-processing, leading to increased area and power consumption. To address this issue, this paper proposes a lightweight TRNG based on chain-coupled oscillation rings for high-quality randomness with very low FPGA overhead.  Methods  Starting from the state evolution of a Galois oscillation ring (GARO), this work demonstrates that ideal matched-delay conditions can result in periodic and predictable oscillation. However, in practical circuits, delay mismatch, jitter, and process variation disturb the ideal evolution and can be exploited as entropy sources. On this basis, a compact delay-feedback XOR ring is proposed to enhance state uncertainty, introduce feedback competition, and improve randomness through inter-stage delay differences. In addition, a second-order oscillation ring is incorporated to eliminate the all-zero stop state and provide continuous excitation. Multiple rings are then chain-coupled, enabling adjacent rings to mutually interfere with one another and thereby generate stronger irregular oscillations. The proposed design is modeled in MATLAB and implemented on a Xilinx Artix-7 FPGA. Finally, we evaluate its performance by NIST SP 800-22, NIST SP 800-90B, bias, autocorrelation, and voltage-temperature robustness tests.  Results and Discussions  Simulation confirms that the proposed structure avoids stable periodic locking and produces sustained irregular oscillation. Experimental results show that the TRNG passes all NIST SP 800-22 tests and achieves an average minimum entropy of 0.9936 in NIST SP 800-90B test, outperforming conventional RO and GARO-based TRNGs under similar conditions. The measured bias is only 0.0228%, and the autocorrelation remains well below the threshold, indicating excellent statistical independence. The design also maintains high entropy over temperatures from 0 °C to 80 °C and supply voltages from 0.9 V to 1.1 V. Implemented on Artix-7, our proposed TRNG achieves 200 Mbps throughput using only 11 LUTs and 4 DFFs, with 0.108 W power consumption.  Conclusions  This paper presents a lightweight chain-coupled oscillation-ring TRNG that exploits delay mismatch, phase disturbance, and feedback competition to generate high-quality physical randomness. The theoretical analysis clarifies how practical nonidealities transform ideal periodic oscillation into irregular oscillation, providing a design basis for compact oscillator-based entropy sources. By combining delay-feedback XOR rings with chain-coupled mutual disturbance and continuous excitation, the proposed design enhances entropy while avoiding excessive hardware overhead and complex post-processing. FPGA implementation and statistical evaluations verify high entropy, low bias, and high randomness under voltage and temperature variations. Therefore, the proposed TRNG achieves high randomness quality and high throughput while effectively reducing hardware overhead, making it suitable for resource-constrained security applications such as IoT terminals, lightweight cryptographic modules, and embedded authentication systems.
From Touch to Semantics: A Cross-Modal Framework for Zero-Shot Spiking Tactile Object Recognition
CHI Wei, XU Jin
Available online  , doi: 10.11999/JEIT260158
Abstract:
  Objective  Tactile perception enables robots to understand object properties and perform dexterous interactions. However, tactile data are costly to collect and difficult to scale, which limits conventional supervised learning in open-world scenarios. Zero-Shot Learning (ZSL) provides a promising solution by transferring knowledge from seen to unseen categories through semantic representations. Existing tactile ZSL methods either rely on auxiliary visual information or use manually designed attributes, which are often subjective and limited in generalization. Event-based spiking tactile signals are sparse and asynchronous, with rich spatiotemporal dynamics. These properties make semantic modeling more challenging. Systematic studies on zero-shot recognition for such data remain limited. To address these issues, this paper proposes a zero-shot object recognition framework for spiking tactile perception. The framework aims to bridge low-level tactile dynamics and high-level semantics in a scalable manner.  Methods  The proposed framework consists of three components (Fig. 1): spiking tactile feature extraction, semantic prototype construction, and cross-modal tactile-semantic alignment. First, a biomimetic Spiking Graph Neural Network (SGNN) is used to model raw event-based spiking tactile signals. By integrating Leaky Integrate-and-Fire (LIF) neurons with graph-based message passing, the SGNN captures temporal firing dynamics and spatial relationships among tactile sensing units. It then generates discriminative and biologically interpretable high-level tactile embeddings. Second, instead of using manually annotated attributes, a Large Language Model (LLM) is used to generate structured, fine-grained, and extensible tactile attribute descriptions for each object category. These textual descriptions are encoded as continuous semantic vectors to form class-level semantic prototypes with consistent dimensionality across categories. This strategy supports flexible semantic expansion and avoids labor-intensive attribute engineering. Third, a bidirectional tactile-semantic alignment mechanism is designed to improve generalization to unseen categories. A forward mapping projects tactile embeddings into the semantic space for classification, whereas a reverse mapping reconstructs tactile features from semantic representations. A cycle-consistency constraint is imposed between the two mappings to preserve structural coherence and semantic stability across modalities. The overall framework is trained only on seen categories. During zero-shot inference, tactile embeddings of unseen samples are matched with their corresponding semantic prototypes in the shared embedding space.  Results and Discussions  The proposed method is evaluated on the Ev-Object event-based tactile dataset under a strict zero-shot setting, with disjoint seen and unseen category sets. Performance is assessed using Mean Class Accuracy (MCA), Top-k accuracy, and the Semantic Alignment Score (SAS). The proposed framework consistently outperforms representative tactile ZSL baselines across all metrics. It achieves an MCA of 73.48%, a Top-1 accuracy of 62.68%, and a Top-2 accuracy of 88.75%. Ablation studies show that removing the LLM semantic module, bidirectional mapping, or cycle-consistency constraint reduces recognition performance and semantic alignment quality. Removing the LLM semantic module causes a substantial decrease in MCA, which confirms the role of structured LLM-generated tactile semantics in knowledge transfer. Removing the bidirectional mapping or the cycle-consistency constraint also reduces performance, indicating that both components help maintain stable cross-modal alignment. The t-SNE visualization further shows that cycle-consistent alignment yields more compact intra-class clusters and clearer inter-class separation for unseen categories. Semantic prototypes are also better located near the centers of tactile feature clusters. These results indicate that combining biologically inspired spiking models with LLM-generated tactile semantics provides an effective solution for open-world tactile perception.  Conclusions  This paper presents a zero-shot object recognition framework for spiking tactile perception by integrating SGNN-based tactile representation with semantic prototypes. The proposed method addresses key limitations of existing tactile ZSL approaches by avoiding visual data and manual attribute design while effectively modeling the spatiotemporal dynamics of event-based spiking tactile signals. Experimental results under strict zero-shot settings confirm the effectiveness and robustness of the proposed framework. This work provides a strong baseline for zero-shot spiking tactile recognition and offers a principled path toward open-world tactile cognition in robotic systems. Future work will explore generalized zero-shot tactile perception, multimodal extensions, and real-world robotic deployment under noisy and dynamic sensing conditions.
A Noise Reduction Strategy via Coprime-Spacing Subarrays for Biodiversity Acoustic Indices
CHEN Lei, XU Zhiyong, ZHAO Zhao
Available online  , doi: 10.11999/JEIT260237
Abstract:
  Objective  As a popular tool for rapid biodiversity assessment, acoustic indices have attracted increasing attention in the field of soundscape ecology in recent years. Nevertheless, most commonly used acoustic indices are susceptible to background noise. Traditional single-channel noise reduction strategies, including spectral subtraction, high-pass filtering, and threshold detection, have been widely adopted as preprocessing approaches to optimize the calculation of acoustic indices. However, when dealing with anthropogenic interference that overlaps with biotic signals in both time and frequency domains, the denoising capability of single-channel methods degrades severely. Although spatio-temporal adaptive whitening filtering based on microphone arrays provides a feasible approach for suppressing directional interference, it suffers from a non-uniform two-dimensional spatio-temporal amplitude response and the self-cancellation of target signal in the unconstrained interference cancellation. These disadvantages lead to distortion in the time-frequency distribution of target signals, causing acoustic index calculations to deviate from the ground truth. Therefore, this study aims to propose a noise reduction strategy via coprime-spacing subarrays for biodiversity acoustic indices. This method effectively suppresses directional interference while maximally preserving the time-frequency distribution structure of biotic signals.  Methods  The noise reduction strategy based on microphone array spatio-temporal adaptive whitening filtering is proposed, incorporating the Frequency-dependent Acoustic Diversity Index (FADI), which is insensitive to fluctuations in the array's two-dimensional spatio-temporal amplitude response. A noise-robust acoustic index method, termed Adaptive Interference Cancellation–Frequency-dependent Acoustic Diversity Index (AIC-FADI), is subsequently developed. Specifically, a non-uniform linear array is first constructed using three microphones to form two dual-element subarrays with coprime spacing. This design fully exploits the high spatial resolution of wide-spacing arrays to narrow the null width in the direction of interference. Meanwhile, it avoids the physical implementation difficulties and mutual coupling effects associated with small-spacing array designs caused by the ultra-wideband characteristics of target signals. The spatio-temporal adaptive whitening filtering is then performed on each coprime-spacing subarray separately, adaptively forming two-dimensional nulls within the interference support region, thereby suppressing directional anthropogenic interference in analytical data before index calculation. Next, a frequency-dependent threshold scheme is utilized to obtain the binary spectrogram for each coprime-spacing subarray output, abating the influence from gain differences along the frequency axis for a certain direction. Afterwards, by leveraging the high spatial resolution of wide-spacing arrays and the interleaved characteristics of spatial aliasing null positions between the spatio-temporal frequency responses of the two subarrays with coprime spacing, a pointwise maximum fusion is applied to the above two binary spectrograms. This process reconstructs the binary time-frequency distribution structure of target signals outside the interference support region, leading to a single binary spectrogram where biological sound components are preserved to a great extent and anthropogenic interference is considerably suppressed. Ultimately, from this single binary spectrogram, the proportions of non-zero time-frequency bins within each frequency band are calculated and forwarded to the entropy function, resulting in the final AIC-FADI result.  Results and Discussions  The simulation result indicates that the proposed AIC-FADI maintains numerical robustness across an SINR range down to –15 dB (the yellow line in Fig. 5), substantially outperforming the classical ADI version based on single-channel noise reduction algorithm (FADI) and other ADI versions based on single-array interference suppression processing mentioned in this paper (AIC-FADI-s, AIC-FADI1, and AIC-FADI2). The real-world experiment confirms that the proposed spatio-temporal adaptive whitening filtering effectively suppresses wideband interference signals in complex scenarios, thereby improving the SINR of the analyzed recording. This enables some weaker biotic signals to exceed their corresponding frequency-dependent adaptive thresholds, greatly reducing missed detection of the target signal. In addition, by performing pointwise maximum fusion of the binary spectrograms from the two coprime-spacing subarray outputs, AIC-FADI further alleviates the extent of target signal missed detection (Fig. 8). Nevertheless, the real-world experiments also reveal that the interference suppression performance of AIC-FADI degrades for highly time-varying interference components.  Conclusions  This paper addresses the challenge of calculating acoustic indices reliably in complex soundscapes where directional anthropogenic interference overlaps with biotic signals in both time and frequency domains. A noise reduction strategy using coprime-spacing subarrays is proposed, and a new noise-robust acoustic index (AIC-FADI) is then developed. The method is evaluated through simulations and real-world recordings, and the results show that: (1) By applying spatio-temporal adaptive whitening filtering on each coprime-spacing subarray followed by pointwise maximum fusion, the proposed method achieves both wideband interference suppression capability and target information fidelity in complex soundscapes containing strong interference. (2) As a result, the proposed AIC-FADI maintains numerical robustness down to –15 dB SINR, substantially outperforming the classical FADI algorithm and other ADI versions based on single-array interference suppression methods. (3) The proposed method provides a feasible technical solution for extending the practical application scenarios and spatio-temporal coverage of biodiversity acoustic indices in human-dominated areas. However, this study only considers directional interference that is relatively stable or slowly time-varying. Hence, the interference suppression performance degrades for highly time-varying or uncorrelated noise components. These challenges should be addressed in future work through more advanced signal processing techniques to further improve the robustness of acoustic indices in highly complex acoustic environments.
A Survey of Quantum Covert Communication Integration Schemes and Application Scenarios
SUN Yiheng, XU Yongjun, ZHANG Haibo, HUANG Zishan
Available online  , doi: 10.11999/JEIT260282
Abstract:
  Significance   With the growing demand for network communication security, research and development in covert communication and quantum communication have continued to evolve. However, current covert communication suffers from inherent security vulnerabilities; the transmission reliability of quantum communication has been limited by information eavesdropping and harmful interference. Therefore, quantum covert communication has become a research hotspot, integrating the advantages of both covert and quantum communication while addressing their respective security limitations. To this end, this paper provides a comprehensive survey of quantum covert communication integration schemes and application scenarios, including the principles of covert communication and typical enabling techniques; protocols for quantum communication and important quantum techniques; and three types of quantum covert communication integration schemes summarized by different application scenarios. This paper contributes to the design of advanced secure communication networks while offering guidance for the development of future quantum covert communication systems.  Progress   This paper presents a comprehensive survey of recent advances in quantum covert communication integration schemes and application scenarios, with an in-depth discussion of the principles of covert communication and key enabling techniques, such as Fluid Antenna (FA), Reconfigurable Intelligent Surface (RIS), and Unmanned Aerial Vehicle (UAV). FA actively reshapes wireless channel characteristics, particularly the spatial correlation of multipath components, by dynamically adjusting the transmitter physical configuration, thereby reducing information leakage. In Non-Line-of-Sight (NLoS) scenarios, RIS can dynamically alter the direction of reflected transmission of the incident signal, not only enhancing the Channel State Information (CSI) quality of the covert signal but also reducing signal leakage. In flexible or temporary communication networks, UAVs can increase CSI uncertainty, preventing unauthorized users from establishing a stable monitoring model and thereby complicating eavesdropping. Then, key protocols and significant techniques of quantum communication are introduced, including BB84, B92, and E91 for Quantum Key Distribution (QKD), and BF02, Two-Step for Quantum Secure Direct Communication (QSDC). Additionally, the quantum repeaters and Quantum Random Number Generator (QRNG) are reviewed. Based on different application scenarios, quantum covert communication integration schemes can be categorized into enabling, covert, and symbiotic integration schemes, depending on the integration mechanisms. To be specific, the enabling integration scheme leverages the unconditional security of quantum communication to address the security vulnerabilities in covert communication, the covert integration scheme utilizes enabling techniques in covert communication to reduce the detection probability of quantum communication, and the symbiotic integration scheme combines both advantages of covert communication and quantum communication to achieve mutual empowerment and deep symbiosis. Finally, critical challenges are highlighted, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles in large-scale applications. Promising directions for future research are also identified, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development.  Prospects   Despite remarkable progress in preliminary applications and specific scenarios, research on quantum covert communication remains in its infancy. As quantum covert communication scenarios become increasingly diverse and complex, future studies should prioritize challenges that restrict further development and large-scale application of quantum covert communication. The stringent hardware precision requirements are the primary challenge, limiting reliable transmission distance and stability. Low resource allocation efficiency is another challenge, as the quantum covert communication system that generates quantum entanglement over lossy channels remains subject to the Square Root Law (SRL) constraints, while signal transmission exhibits burstiness and dynamics. Additionally, high deployment costs and the lack of standardization present significant hurdles. To address the challenges mentioned, future directions should include R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development to facilitate the development of high-performance, large-scale, and multi-scenario quantum covert communication.  Conclusions  This paper provides a comprehensive survey of quantum covert communication with particular emphasis on integration schemes and application scenarios. The fundamentals and typical enabling techniques of covert communication are first reviewed, highlighting its Low Probability of Detection (LPD) secure paradigm and unique channel characteristics. The typical protocols and important techniques of quantum communication are then examined, including QKD, QSDC, quantum repeaters, and QRNG. Three types of quantum covert communication integration schemes have been further classified by different integration mechanisms and corresponding application scenarios. Finally, several existing challenges are identified, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles to large-scale applications. Relevant research directions are also outlined, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development. These directions are expected to serve as a valuable reference for advancing and standardizing quantum covert communication in future secure networks.
Full-Space Covert Integrated Sensing and Communications Assisted by Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface
XIE Wenwu, ZHANG Qinke, YANG Liang, WANG Ji, YU Chao, LIU Xinzhong, CUI Yaru
Available online  , doi: 10.11999/JEIT260145
Abstract:
  Objective  The evolution of Sixth Generation (6G) mobile communications toward higher frequencies and larger antenna arrays has made Integrated Sensing And Communication (ISAC) a key enabling technology. However, ISAC systems still face limited communication covertness and resource competition between sensing and communication. Covert communication and Reconfigurable Intelligent Surface (RIS) techniques provide promising solutions. However, most existing studies use reflective RISs with half-space coverage and assume far-field propagation. These assumptions limit deployment flexibility and fail to capture near-field spherical-wave characteristics. To address these issues, this paper proposes a near-field full-space ISAC framework assisted by an Extremely Large-Scale Simultaneously Transmitting And Reflecting Reconfigurable Intelligent Surface (XL-STAR-RIS). The objective is to jointly optimize active transmit beamforming and passive XL-STAR-RIS coefficient design to improve the covert communication rate while satisfying sensing performance and covertness requirements.  Methods  The detection capability of warden Willie is first analyzed, and a closed-form lower-bound expression for the minimum Detection Error Probability (DEP) is derived. A non-convex optimization problem is then formulated to maximize the covert communication rate under sensing Signal-to-Noise Ratio (SNR), covertness, and total transmit power constraints. Direct solution is difficult because the active transmit beamforming vectors and passive XL-STAR-RIS coefficients are strongly coupled. An Alternating Optimization (AO) framework is therefore adopted to decompose the original problem into two tractable subproblems. The active transmit beamforming subproblem is solved using SemiDefinite Relaxation (SDR) combined with a penalty-based successive convex approximation method. The passive XL-STAR-RIS coefficient design subproblem is solved using the Dinkelbach algorithm and a rank-one penalty method. The two subproblems are solved alternately until convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed framework. The algorithm converges within approximately 10 iterations and achieves a covert communication rate of about 11.5 bit/(s·Hz). This rate is higher than those of the passive-RIS scheme (9.8 bit/(s·Hz)) and the non-RIS scheme (8.0 bit/(s·Hz)). The performance gain becomes more evident as the transmit power increases, which indicates strong power adaptability. The proposed framework also maintains robust performance under strict operational constraints. When the sensing SNR threshold increases, it achieves a higher covert communication rate than the benchmark schemes. Under a stricter covertness requirement, it also preserves a higher communication rate. These results show that joint active transmit beamforming and passive XL-STAR-RIS coefficient design can effectively balance communication, sensing, and covertness in near-field ISAC systems.  Conclusions  This paper presents an XL-STAR-RIS-assisted covert communication framework for near-field ISAC systems. By jointly designing active transmit beamforming and passive XL-STAR-RIS coefficients through an efficient AO algorithm, the proposed framework balances communication rate, sensing performance, and communication covertness. Simulation results confirm its advantages over conventional passive-RIS and non-RIS schemes, especially under strict sensing and covertness constraints. The results also indicate the potential of XL-STAR-RIS for secure full-space 6G applications. Future work will consider imperfect Channel State Information (CSI), dynamic propagation environments, and multi-RIS collaboration to improve practical robustness.
Millimeter-Wave Air-to-Ground Channel Prediction Assisted by Visual Information of the Propagation Environment
CHENG Yuanxun, HU Qingsong, ZHANG Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260274
Abstract:
  Objective  Accurate prediction of air-to-ground (A2G) channel states is essential for adaptive transmission and resource optimization in unmanned aerial vehicle (UAV) communications. In urban millimeter-wave scenarios, however, A2G links are highly sensitive to blockage, reflection, scattering, and the rapidly changing geometric relationship among the transmitter, the receiver, and surrounding buildings. As a result, the channel exhibits strong spatial and temporal nonstationarity, and conventional pilot- or feedback-based acquisition methods may become ineffective because the obtained channel state information is easily outdated. Recent data-driven approaches have shown potential, but many of them rely heavily on historical channel observations or directly use raw images as network inputs, which may introduce redundant visual information and weaken physical interpretability. To address these limitations, this paper proposes a vision-assisted millimeter-wave A2G channel prediction method that extracts low-dimensional geometric features from the propagation environment instead of using raw visual data directly. The objective is to preserve the key structural information governing channel evolution while reducing irrelevant redundancy, thereby improving the prediction of channel.  Methods  A communication-and-sensing integrated dataset with strict spatial and temporal alignment is established for millimeter-wave UAV A2G channel prediction. On the sensing side, a high-fidelity three-dimensional urban scenario containing 23 buildings, roads, and intersections is constructed in Unreal Engine 4.27, where synchronized RGB and depth images are collected through AirSim using a multirotor UAV equipped with RGB and depth cameras. The UAV flies along 10 preset trajectories at a height of 55 m with a spatial sampling interval of 1 m, yielding 2160 valid visual samples (Fig. 1, Fig. 2). On the communication side, the same scene is reconstructed in Wireless InSite, and the transmitter-receiver positions are synchronously updated along the same trajectories to ensure frame-level alignment between visual and channel data (Fig. 3). To obtain compact and physically meaningful environmental representations, a cross-modal spatial feature extraction scheme is developed. Buildings are first detected from RGB images using YOLO-V8 (Fig. 4), and the detected regions are then registered with depth images to reconstruct three-dimensional point clouds. After Euclidean clustering and axis-aligned bounding-box fitting, key geometric attributes, including planar position, height, and volume, are extracted. These features are combined with the transmitter-receiver distance to form the spatial feature vector of each frame, and their relevance to path loss, received power, and RMS delay spread is evaluated through cosine-similarity-based correlation analysis (Fig. 6). Based on the extracted features, a hybrid Transformer-MLP network is designed for channel prediction (Fig. 5). Building features are first projected into a latent space, and a stacked Transformer encoder is employed to capture global interactions among buildings through masked multi-head self-attention. Masked average pooling is then used to aggregate building-level representations into a scene-level environmental descriptor, which is concatenated with the link distance feature and fed into a multilayer perceptron regressor to predict the three target channel parameters.  Results and Discussions  The results confirm the effectiveness of the proposed spatial feature representation. Correlation analysis shows that the extracted geometric features are consistently related to path loss, received power, and RMS delay spread under different aggregation strategies (Fig. 6), indicating that compact building descriptors can effectively characterize the propagation environment. Among them, building height exhibits the strongest correlation with all three channel parameters, highlighting its important role in blockage, attenuation, and multipath propagation in urban millimeter-wave A2G channels. In prediction experiments, the proposed method accurately tracks the variation trends of all three targets. It remains effective in deep-fading and sharp-fluctuation regions for path loss prediction (Fig. 7), achieves high consistency with the ground truth for RMS delay spread (Fig. 8), and follows rapid local fluctuations of received power with good fidelity (Fig. 9). In contrast, the benchmark model only captures the general trend and shows larger deviations in peaks, valleys, and abrupt-changing intervals. Residual analysis further demonstrates the superiority of the proposed method. Its errors are more concentrated around zero and fluctuate within narrower ranges than those of the benchmark model across all three tasks (Fig. 10). Quantitatively, both the mean absolute error and the root mean squared error are reduced (Fig. 11). In addition, the model maintains acceptable complexity, with about 5.5 M parameters and a single-frame inference delay of about 3.4 ms, indicating good potential for real-time deployment.  Conclusions  A vision-assisted millimeter-wave A2G channel prediction method for UAV communications is proposed. By constructing a strictly aligned communication-and-sensing dataset and extracting low-dimensional spatial features with clear physical meaning, the method establishes an effective mapping from environmental geometry to channel parameters. The proposed Transformer-MLP framework achieves accurate prediction of path loss, received power, and RMS delay spread, while offering better interpretability, robustness, and efficiency than the benchmark model.
Transfer Learning Aided CNN for Efficient Data Detection in ReRAM with Sneak-Path Interference
DAI Bin, WU Anni
Available online  , doi: 10.11999/JEIT260354
Abstract:
  Objective  Sneak path interference (SPI) in resistive random-access memory (ReRAM) introduces unpredictable inter-cell correlations, significantly increasing the complexity of signal detection. Traditional detection methods typically rely on assumptions about known channel noise states, resulting in limited generalization capability in practical applications. To address this issue, three data detection methods based on convolutional neural networks (CNNs) are proposed, which can effectively model and mitigate interference without relying on prior channel information: first, a method combining constrained coding with a multi-layer CNN, which uses constrained coding to determine the sneak path interference state and recover data; second, a dual-CNN framework that first employs a lightweight CNN for sneak path interference identification, followed by a multi-layer CNN for refined detection; third, an approach incorporating transfer learning, which maintains detection accuracy while reducing the required training sample size to one-thousandth of that of traditional methods. Simulation results demonstrate that the proposed method achieves superior bit error rate (BER) performance under unknown channel conditions, with a BER reduction of at least half relative to existing algorithms, approaching the theoretical performance limit. Moreover, the integration of transfer learning reduces the required training samples from \begin{document}$ {10}^{6} $\end{document} to \begin{document}$ 1000 $\end{document}, corresponding to a reduction of three orders of magnitude.  Methods  To address distinct challenges in sneak path interference detection, this paper proposes three methods sequentially:1. The integrated constrained coding aided convolutional neural network (CC-CNN) detection framework effectively addresses the complex inter-cell correlations introduced by sneak path interference. This approach first employs constrained coding to detect the presence of interference and subsequently utilizes a CNN to learn and capture the random correlations under the influence of interference, thereby achieving accurate signal recovery.2. The dual-CNN-based detection method resolves the code rate loss associated with traditional constrained coding. By directly leveraging a CNN to learn and identify sneak path interference patterns from raw data, this method eliminates the need for redundant coding or additional overhead. It ensures high-precision interference detection while preserving the overall code rate performance of the system.3. The transfer learning-based CNN (TL-CNN) detection method overcomes the dependence of high-performance CNNs on large-scale training datasets. By reusing knowledge from pre-trained models, this method enables rapid adaptation to ReRAM signal detection tasks. It significantly reduces the required number of training samples while maintaining high detection accuracy and resource efficiency, thereby enhancing the feasibility of the solution in practical scenarios.  Results and Discussions  Simulation results demonstrate that the performance of the three proposed methods consistently approaches the theoretical lower bound (Fig.6), outperforming baseline methods such as the Belief Propagation (BP) detector, Deep Neural Network (DNN) detector, and Elementary Signal Estimator (ESE) detector. The two-step network achieves performance comparable to that of the single-step network while successfully avoiding code rate loss. Notably, the transfer learning-aided CNN attains near-optimal BER with only 1000 target domain samples, and its performance stabilizes when the sample size exceeds 1000 (Fig.7), fully validating its data efficiency. The integration of SK modules enables the models to effectively capture SPI-induced spatial correlations, while the transfer learning strategy ensures the models’ robust performance under different noise conditions.  Conclusions  The crossbar array architecture of ReRAM is susceptible to sneak-path interference during storage operations, leading to reduced data reliability. To address this issue, this paper proposes three deep learning-based detection methods. Type-I integrates constrained coding with a CNN to achieve efficient and fast interference detection. Type-II adopts a two-stage processing approach: it first classifies interference patterns in the memory array and then performs detection specifically on affected units, thereby ensuring high detection accuracy while minimizing coding rate loss. Type-III introduces a transfer learning framework that leverages a pre-trained model from the source domain, significantly reducing the number of training samples required in the target domain and effectively lowering training overhead. Experimental results show that under different noise conditions, all three proposed methods achieve performance close to the theoretical lower bound, providing an effective solution for enhancing the reliability of ReRAM storage systems.
Physical-layer Security in Visible Light Communications: Fundamental Theories, Key Techniques, and Future Challenges
WANG Jinyuan, YAN Xinrun, LIN Zihan, LI Yuanyuan, LI Zheng, ZHANG Xin
Available online  , doi: 10.11999/JEIT260338
Abstract:
  Significance   Due to the broadcast nature of optical signals, information security represents a critical research direction in visible light communication (VLC). Conventional encryption techniques address network security issues at the upper layers of the protocol stack through access control, cryptographic protection, and end-to-end encryption. However, their security relies on the assumption that eavesdroppers possess limited computational capabilities, an assumption that currently faces significant challenges. In recent years, physical layer security (PLS) has emerged as a novel information security paradigm and has attracted considerable attention from researchers worldwide. PLS exploits the randomness, heterogeneity, and distinctiveness between the main channel and the eavesdropping channel to achieve secure information transmission at the physical layer. To date, extensive research achievements have been made regarding PLS techniques in conventional radio frequency wireless communications (RFWC). Nevertheless, due to substantial differences in frequency bands, transmitted signals, power representations, and channel characteristics, PLS research results from RFWC systems cannot be directly applied to VLC. Although scholars worldwide have conducted research on VLC PLS technology, the foundational theories, key techniques, and future challenges involved in VLC PLS still lack a systematic review. To bridge this gap, this paper presents a comprehensive survey of VLC PLS technology.  Progress   To evaluate and enhance system performance, a classic VLC PLS system model—comprising the received signal model, the input constraint model, and the channel gain model—is initially established. A comprehensive theoretical framework for performance evaluation is then developed, encompassing instantaneous performance metrics, statistical performance metrics, and asymptotic performance metrics. Specifically, to characterize instantaneous performance, existing works on instantaneous secrecy capacity and instantaneous secrecy rate across different scenarios are summarized. As statistical performance metrics, average secrecy capacity, average secrecy rate, secrecy outage probability, probability of strictly positive secrecy capacity, and interception probability are analyzed. To demonstrate asymptotic performance, secrecy diversity order and secrecy degrees of freedom are derived. Furthermore, to enhance the PLS performance, advanced technologies, including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity, are summarized.  Prospects   Despite existing research achievements, numerous challenges remain in VLC PLS. This paper identifies four critical challenges: (i) Accurate PLS performance limit: Deriving exact expression of secrecy capacity under VLC's unique physical constraints remains challenging. (ii) Incomplete evaluation framework: Some key metrics widely used in RFWC have not been investigated in VLC, and the construction of a comprehensive VLC PLS performance evaluation framework remains unresolved. (iii) Limitations of existing methods: Conventional PLS performance enhancement methods typically adopt a “modeling-optimization-verification” separated research paradigm, often falling into a vicious cycle of “inaccurate modeling-suboptimal solutions-limited performance gains”. Therefore, it is imperative to integrate novel technologies (such as deep learning, reinforcement learning, and digital twins) to construct a data-model dual-driven framework for VLC PLS performance enhancement. (iv) Hardware platform gap: The absence of dedicated hardware platforms featuring adversarial topologies and real-time processing capabilities significantly impedes the practical deployment of VLC PLS technologies. Therefore, addressing these challenges is essential for transitioning VLC PLS from theoretical advances to commercial applications.  Conclusions  The broadcast nature of optical signals renders VLC systems vulnerable to eavesdropping attacks. This paper presents a comprehensive survey of PLS in VLC, covering system models, performance metrics (instantaneous, statistical, and asymptotic), and key performance enhancement technologies including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity. Despite significant progress, challenges remain in establishing accurate performance bounds, complete evaluation frameworks, novel enhancement techniques, and practical hardware implementations. By exploiting channel disparities at the physical layer without relying on complex encryption, PLS represents a paradigm shift in security assurance, paving the way for next-generation secure and reliable VLC networks.
Performance Optimization and Gate Oxide Electric Field Analysis of 1200V Trench SiC MOSFET Based on PCL-CSL Collaborative Design
FANG Shaoming, LI Hongda, GAO Yuan
Available online  , doi: 10.11999/JEIT260164
Abstract:
  Objective  1 200 V Silicon Carbide (SiC) trench Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) are key devices in medium- and high-voltage power conversion systems. They feature high switching performance, low conduction loss, and high-temperature stability. However, conventional trench structures suffer from electric-field concentration at the trench corner and bottom gate oxide. This effect can cause the peak gate oxide electric field to exceed the industrial reliability criterion of 3 MV/cm, reducing long-term reliability. In addition, strong trade-offs exist among breakdown voltage, specific on-resistance, threshold voltage, and peak gate oxide electric field. These trade-offs make it difficult to achieve high efficiency and high reliability at the same time. To address these issues, this work studies a synergistic structure that combines deep P-type Column (PCL), Carrier Storage Layer (CSL), and locally thickened gate oxide. The aim is to regulate the electric-field distribution, suppress electric-field concentration, improve carrier transport, and achieve balanced device performance. This study provides a systematic design method for high-reliability and high-performance 1 200 V Trench SiC MOSFETs for industrial applications.  Methods  Numerical device simulations were performed using a Technology Computer-Aided Design (TCAD) platform to analyze and optimize the electrical performance of 1 200 V Trench SiC MOSFETs. To ensure reliable simulations, physical models were used for bandgap narrowing, Shockley-Read-Hall (SRH) recombination, Auger recombination, avalanche breakdown, incomplete dopant ionization, doping- and temperature-dependent mobility, and high-field mobility saturation. A device structure with deep PCL, CSL, and locally thickened bottom gate oxide is constructed to reduce the peak gate oxide electric field and improve device reliability. Key structural and process parameters were swept and quantitatively analyzed. These parameters included epitaxial layer thickness (TEpi), epitaxial layer doping concentration (NEpi), trench width, trench depth, P-Well (PW) implantation dose, PCL spacing, and CSL implantation dose. Static electrical characteristics, including threshold voltage (Vth), specific on-resistance (Ron,sp), Breakdown Voltage (BV), and peak gate oxide electric field (Eox,max) are extracted and evaluated. The final parameter combination is finally determined through a trade-off analysis between conduction performance and long-term device reliability.  Results and Discussions  The simulation results show that the deep PCL structure redirects electric-field lines away from the trench bottom gate oxide and reduces electric-field concentration. When this structure is combined with the locally thickened bottom gate oxide, Eox-max is reduced below 3 MV/cm, meeting the industrial reliability criterion. The CSL broadens the vertical conduction path, reduces current crowding, and decreases Ron,sp. Parameter optimization shows that TEpi, NEpi, trench dimensions, PW implantation dose, and CSL implantation dose determine the trade-off between BV and conduction performance (Fig. 5, Fig. 6, Fig. 9, Fig. 10, and Fig. 19). PCL spacing has a strong effect on electric-field shielding and gate oxide protection (Fig. 16 and Fig. 17). After multi-parameter optimization, the device achieves VTH=4.7 V, BV=1 708 V, Ron,sp=1.57 mΩ·cm2, and Eox-max=2.5 MV/cm (Table 2). These results indicate balanced performance for high-voltage power applications.  Conclusions  A synergistic PCL-CSL structural design for 1 200 V Trench SiC MOSFETs is studied and validated through TCAD simulation. The design addresses key limitations of conventional Trench SiC MOSFETs, including high peak gate oxide electric field, limited breakdown capability, and the trade-off between conduction performance and reliability. The effects of TEpi, NEpi, trench dimensions, PW implantation dose, PCL spacing, and CSL implantation dose on device performance and gate oxide reliability are clarified through parameter sweeping and comparative analysis. With coordinated structural optimization, the optimized device achieves low Ron,sp, high BV, suitable VTH, and suppressed electric-field concentration near the trench bottom oxide. Eox-max is controlled below the 3 MV/cm industrial reliability criterion, which reduces the risk of oxide degradation under high-bias operation. The proposed structural strategy and optimization method provide guidance for the design, simulation, and process development of high-voltage, high-reliability SiC power devices.
Research on Energy Efficiency Optimization of Rotatable Hybrid Intelligent Reflecting Surface Communication
ZHANG Guangchi, GUO Xuan, WANG Luyao, CUI Miao, FU Hao
Available online  , doi: 10.11999/JEIT260119
Abstract:
  Objective  With the evolution of 6G communication networks, reconfigurable intelligent surfaces (RIS) have emerged as a pivotal technology for reshaping wireless environments and enhancing spectral efficiency. However, conventional fixed RIS architectures face two critical challenges in practical deployment: the “angle mismatch” loss, where the effective aperture significantly diminishes when users are located at large angles from the RIS normal, and the “energy consumption bottleneck,” caused by the high cumulative power consumption of radio frequency (RF) circuits and static control elements in large-scale arrays. Existing research often treats mechanical rotation and element switching in isolation, lacking a unified framework to balance the trade-off between mechanical/circuit energy consumption and communication gain. To address these limitations, this paper investigates a rotatable and switchable hybrid RIS (H-RIS) assisted downlink communication system. The primary objective is to maximize the system’s energy efficiency (EE) by jointly optimizing the base station transmit power, subarray activation states, physical rotation angles, and electronic phase shifts. This approach aims to introduce mechanical rotation degrees of freedom to compensate for path loss and employ dynamic switching mechanisms to reduce redundant power consumption, thereby achieving sustainable green communication.  Methods  A joint optimization framework is established for the H-RIS aided single-user multiple-input single-output (MISO) system. The system model explicitly accounts for the dynamic power consumption induced by mechanical rotation and the static power consumption of active subarrays. The resulting optimization problem is formulated as a non-convex mixed-Integer non-linear programming (MINLP) problem, involving coupled binary variables (activation status) and continuous variables (power, angles, phases). To solve this challenging problem, a block coordinate descent (BCD)-based alternating optimization (AO) algorithm is proposed to decouple the variables into three sub-problems.Firstly, to tackle the exponential complexity caused by binary switching variables, a channel contribution-based ranking strategy is developed. By performing eigenvalue decomposition on the cascaded channel correlation matrix, the priority of each subarray is quantified, reducing the search space from exponential to linear.Secondly, for the power allocation sub-problem, the non-convex fractional objective function is transformed into a parametric subtractive form using the Dinkelbach algorithm, which is then solved via the interior-point method.Thirdly, for the physical rotation and electronic phase optimization, the problem is decomposed into single-variable sub-problems. A Golden Section Search algorithm is employed to iteratively find the optimal rotation angle and phase shift for each subarray within bounded constraints, ensuring the monotonic convergence of the objective function.  Results and Discussions  Extensive simulations are conducted to evaluate the performance of the proposed H-RIS scheme compared with benchmark schemes, including “Only-Rotation” (always on), “Only-Switching” (fixed angle), and “Conventional” (fixed and always on).The simulation results regarding the maximum transmit power Pmax(Fig. 2 and Fig. 3) demonstrate that the proposed method achieves the highest energy efficiency across the entire power range. Specifically, in the low power regime, the proposed algorithm intelligently turns off redundant subarrays where the rate gain cannot offset the circuit power cost, thereby significantly outperforming the “Only-Rotation” scheme which suffers from high static power consumption.The impact of user distance is also analyzed (Fig. 4 and Fig. 5). Results indicate that the proposed scheme maintains high spectral efficiency comparable to the “Only-Rotation” scheme by dynamically adjusting the rotation angles to align with the Line-of-Sight (LoS) path, effectively compensating for the angle mismatch loss observed in the “Only-Switching” and “Conventional” schemes.Furthermore, the activation pattern of the subarray varies in a “U” shape with distance (Table 1), which allows for flexible adjustment of array size and orientation according to user-RIS geometry.  Conclusions  This paper proposes an energy-efficient transmission scheme for H-RIS aided communication systems by integrating mechanical rotation and dynamic switching capabilities. A low-complexity BCD-based algorithm is developed to jointly optimize the transceiver design. The results confirm that introducing mechanical rotation significantly mitigates the angle mismatch loss, while the proposed channel contribution-based switching strategy effectively eliminates redundant energy consumption. The proposed H-RIS architecture offers a superior trade-off between spectral efficiency and energy efficiency compared to traditional fixed RIS architectures, providing a viable solution for future green 6G networks.
CRLB Optimization for O-RIS-Assisted VLP Systems
ZHANG Zengjie, WU Qi, ZHANG Jian, DUAN Ruijie, FENG Yunhan
Available online  , doi: 10.11999/JEIT260120
Abstract:
  Objective  With the rapid development of indoor location-based services, Visible Light Positioning (VLP) has emerged as a promising high-accuracy positioning technology. The integration of Optical Reconfigurable Intelligent Surfaces (O-RIS) into VLP systems can effectively enhance signal coverage and improve positioning performance. However, optimizing the positioning accuracy and fairness across different user areas in RIS-assisted VLP systems remains a challenging issue. This study focuses on optimizing the Cramer-Rao Lower Bound (CRLB) of the system under both near-field and far-field channel models, aiming to enhance overall positioning precision and fairness through RIS configuration.  Methods  Under the far-field channel model assumption, the RIS orientation optimization problem is formulated as a received power maximization problem. A positioning algorithm combining Particle Swarm Optimization (PSO) and N-step iteration is proposed to dynamically adjust the RIS orientation optimally without prior knowledge of the receiver’s position. Under the near-field channel model assumption, the allocation problem between RIS elements and LEDs is constructed as a Markov Decision Process (MDP). A reinforcement learning method based on experience replay and knowledge utilization is designed to solve this problem, aiming to minimize the CRLB while ensuring positioning fairness for users in different regions.  Results and Discussions  Simulation results demonstrate that the proposed algorithms effectively enhance system positioning performance under both models. In the far-field model, the PSO-based iterative algorithm achieves dynamic optimization of RIS orientation, significantly improving positioning accuracy (Fig. 3). Under the near-field model, the reinforcement learning approach not only minimizes the CRLB but also considerably improves positioning fairness across the entire area, with a noticeable reduction in performance disparity among users in different zones (Fig. 5, Fig. 6). Comparative experiments show that the proposed methods outperform conventional RIS configuration strategies in terms of both average positioning error and fairness index (Table 1).  Conclusions  This paper investigates CRLB optimization methods for O-RIS-assisted VLP systems under near-field and far-field channel models. In the far-field scenario, a PSO-based iterative algorithm is proposed to optimize RIS orientation, enhancing positioning accuracy without requiring prior receiver location information. In the near-field scenario, a reinforcement learning-based approach is designed to optimize RIS element–LED allocation, which effectively minimizes the CRLB and improves positioning fairness across the whole area. Simulation results validate the effectiveness of the proposed algorithms in both models. Future work may consider more practical channel impairments and multi-user scenarios to further improve the robustness and scalability of the system.
A Radio Frequency Fingerprint Open-set Identification MethodCombining Multi-scale Wavelet Front-end and Hyperspherical Metric Learning
TIAN Xinyu, LI Zirui, ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260214
Abstract:
  Objective  Open-set Radio Frequency Fingerprint (RFF) identification under low Signal-to-Noise Ratio (SNR) conditions is challenging because fingerprint features are easily masked by noise, multipath effects induce nonlinear distortions, and existing methods struggle with feature extraction and unknown device detection. This study proposes a deep learning framework that integrates a multi-scale wavelet front-end with hyperspherical metric learning to achieve robust open-set RFF identification.  Methods  The proposed method, MS-RANet, comprises three key components. First, a multi-scale wavelet front-end based on one-dimensional stationary wavelet transform performs full-resolution, multi-scale decomposition of I/Q signals, preserving discriminative fingerprint information while suppressing noise. Second, a multi-scale residual attention network incorporates deep residual learning, global self-attention, and Bidirectional LSTM (BiLSTM) to enhance sensitivity to subtle fingerprint features and capture long-range temporal dependencies. Third, hyperspherical metric learning constrains the feature space onto a unit hypersphere, optimizing angular margins to produce compact intra-class and separable inter-class feature distributions. Unknown devices are subsequently detected using cosine similarity.  Results and Discussions  Experiments on a high-fidelity IEEE 802.11 simulation dataset demonstrate the effectiveness of MS-RANet. The method achieves an average classification accuracy of 65.34% across SNR levels from –5 dB to 20 dB, and an Area Under the Curve (AUC) of 0.81 at –5 dB SNR, outperforming DNN, GRU, CNN-LSTM, ResNet50, and DRSN-CA. Confusion matrices and Receiver Operating Characteristic (ROC) curves confirm robustness under extreme channel conditions. t-SNE visualization shows well-separated, compact clusters for known devices, while unknown samples are effectively isolated from known class regions. Ablation studies verify the contributions of the multi-scale wavelet front-end, global attention, BiLSTM, and hyperspherical metric learning modules.  Conclusions  This study presents a robust open-set RFF identification method combining a multi-scale wavelet front-end with hyperspherical metric learning. The framework exhibits strong noise resilience, enhanced feature discrimination, and reliable detection of unknown devices under low-SNR and multipath fading conditions. Future work will focus on reducing computational complexity, improving inference speed, evaluating generalization across diverse scenarios and protocols, and integrating the method with complementary physical-layer security mechanisms for collaborative authentication.
A Cross-Precision Motion Compensation Technique for Security Surveillance Video Coding
JIANG Wei, MA Wei, LU Jinghui, ZHANG Yue, ZHANG Yundong
Available online  , doi: 10.11999/JEIT251301
Abstract:
  Objective  In the field of modern security surveillance, high-altitude dome cameras are often deployed at critical locations such as bridges and tower tops that are susceptible to external interference, resulting in problems such as jitter and blurring in captured videos, which pose great challenges to video coding. In video compression coding, high-precision motion compensation is the key to improving coding efficiency. The existing Ultimate Motion Vector Expression (UMVE) technique suffers from insufficient precision and lack of flexibility in adaptive adjustment. Although high-precision coding tools such as Registration-Based Coding Mode (RCM) and Affine Motion Compensation Prediction (AFFINE) can improve compensation accuracy, they have disadvantages of high computational complexity and hardware cost, making it difficult to meet the multiple requirements of coding efficiency, power consumption and real-time performance in high-altitude surveillance scenarios. Therefore, aiming at the core pain points of video coding for high-altitude dome cameras, it is of important academic value and practical application significance to design an optimized UMVE scheme that combines high-precision motion compensation, low computational complexity and scene adaptability, so as to improve coding efficiency and balance resource consumption.  Methods  This study proposes an Ultimate Motion Vector Expression technique supporting Cross-Precision Motion Compensation (UMVE_CPMC). Its core is to improve motion compensation accuracy by constructing an extended Up-Precision Motion Vector (UPMV), whose mathematical expression is UPMV = BaseMV + MMV(p, angle), where BaseMV is the basic motion vector obtained by the existing UMVE method, and MMV is the refined fine-tuning motion vector based on specific precision p and angle, with incremental candidates only provided at the 1/8 precision level to balance computational complexity and compression efficiency. For step-size adaptive adjustment, an improved scheme with six modes is proposed, covering enhanced UMVE, conventional UMVE and four precision-improved modes, allowing the encoder to switch flexibly according to scene characteristics. The average image gradient is adopted as an objective evaluation index; test scenes are divided into Class A (high-definition motion scenes) and Class B (low-definition scenes), and different coding configurations, sequences and parameters are set to compare coding gains and computational efficiency under different modes.  Results and Discussions  Experiments show that UMVE_CPMC achieves effective performance improvement in various scenes and modes. In Class A high-definition motion scenes, with the adaptive strategy disabled and RCM disabled, the average gains of Y, U and V components in Fusion Mode 1 reach -2.912%, -1.656% and -1.654% respectively, and the average coding time is reduced to 94.55% of the baseline; the average gain of the Y component in Independent Mode 1 reaches -2.925%, with coding time reduced to 91.91% of the baseline. Compared with traditional UMVE, when CPMC Independent Mode 1 is enabled under the scenario where RCM is enabled and other tools work collaboratively, the gain is improved from -0.276% to -1.310%, showing significantly higher cost performance. In Class B low-definition scenes, after enabling adaptive adjustment, the gain losses of Fusion Mode 1 and Mode 0 are significantly reduced, with average gain losses controlled at 0.071% and 0.108% respectively, successfully maintaining the original coding gain. In multi-scene comprehensive tests, when RCM and AFFINE are disabled, 9 out of 10 test sequences in adaptive Fusion Mode 1 show positive gains, including a Y-component gain of -10.691% for the yuxuedaolu sequence and -11.400% for the BQTerrace sequence. When all existing coding tools are enabled, the Y-component gains of dianjing, yuxuedaolu and BQTerrace sequences reach -1.29%, -2.05% and -1.21% respectively, with coding time reduced to 94%–96% of the baseline. In addition, correlation analysis between average image gradient and gain reveals a significant positive correlation: images with high average gradient (high definition) achieve greater gains from UMVE_CPMC, while those with low average gradient (low definition) hardly benefit. Principle analysis indicates that pixel changes in low-definition images are gentle, and high-precision interpolation fails to generate effective pixel values, resulting in insignificant compensation effects. Performance differences among modes match computational complexity: the fusion mode balances gain and stability, while the independent mode further reduces computation. The six step-size adaptive modes can meet real-time and precision requirements of different scenes.  Conclusions  The proposed UMVE_CPMC technique, by integrating cross-precision motion compensation with the UMVE algorithm, effectively solves the core problems of insufficient precision in traditional UMVE and high computational complexity of high-precision coding tools, achieving a favorable balance among coding efficiency, computational complexity and scene adaptability. This technique delivers remarkable coding gains in Class A high-definition motion scenes, with gains exceeding 10% for some sequences without other high-precision compensation tools and 1%–2% when cooperating with other tools. In Class B low-definition scenes, the original coding gain can be maintained through frame-level adaptive adjustment interfaces. Meanwhile, the fusion mode does not increase hardware complexity, and the independent mode significantly reduces coding time, suitable for encoder designs with limited resources or simplified requirements. UMVE_CPMC provides a new effective approach to solving the low coding efficiency caused by jitter and blurring in high-altitude dome camera video coding, enriches the video coding toolset, and offers important practical guidance for the optimization of video coding technologies in the security surveillance field. Future work can further optimize the adaptive strategy, explore integration with other advanced coding technologies, develop personalized coding schemes, and improve performance in complex scenarios.
Semantic Relation-enhanced Adaptive Graph Representation Learning for Next POI Recommendation
WANG Zhuolu, XU Shenghua, WANG Yong, JIANG Shunshun
Available online  , doi: 10.11999/JEIT251357
Abstract:
  Objective  In recent years, next Point Of Interest (POI) recommendation has played an increasingly important role in Location-Based Social Networks (LBSNs). However, existing Graph Representation Learning (GRL)-based recommendation methods have struggled to balance node distributions across different domains (i.e., node types) effectively and have often overlooked feature differences among heterogeneous relations. Thus, complex semantic dependencies in contextual information cannot be fully captured when users’ temporal preference patterns are modeled.  Methods  To address these issues, a next POI recommendation method based on Semantic Relation-enhanced adaptive Graph Representation Learning (SR-GRL) is proposed. A heterogeneous transition graph is constructed to integrate three entity types, namely POIs, POI categories, and regions, and their complex interrelationships. An adaptive balanced random walk sampling strategy is designed to balance node distributions across different domains dynamically and to reduce information redundancy. A type-aware attention mechanism is then used to learn semantic associations among nodes through relation-specific transformation matrices, so that feature differences across node types can be identified effectively. The obtained disentangled POI representations are then used for spatiotemporal encoding of user check-in sequences, and a self-attention mechanism is applied to aggregate users, temporal preference features. Finally, next POI recommendation is generated through a Softmax function.  Results and Discussions  Experiments on the Foursquare datasets from Tokyo and New York and the Sina Weibo dataset from Shanghai show that, compared with state-of-the-art baselines, the SR-GRL method achieves Recall@10 improvements of 2.22%\begin{document}$ \sim $\end{document}24.16%, F1@10 improvements of 1.16%\begin{document}$ \sim $\end{document}10.48%, and NDCG@10 improvements of 3.01%\begin{document}$ \sim $\end{document}17.37%, indicating better recommendation performance.  Conclusions  Overall, the SR-GRL approach can balance the distributions of different node types dynamically and strengthen the modeling of complex semantic dependencies in heterogeneous contextual information.
Multi-Agent Deep Reinforcement Learning Strategy for Multi-Spacecraft Long-Distance Orbital Game
DI Peng, YIN Zengshan, LIN Zheng, YAO Ye
Available online  , doi: 10.11999/JEIT251384
Abstract:
This paper introduces a novel research scenario for multi-spacecraft Orbital Pursuit-Evasion Game (OPEG), which has not yet been systematically studied. To enhance the decision-making capabilities of spacecraft and enable them to formulate more robust policies in complex multi-agent games, this paper proposes a multi-agent deep reinforcement learning algorithm based on a progressive adversarial training framework to solve the game policies of each spacecraft. Two sets of examples with different orbital characteristics and various simulation conditions were set up for simulation verification, and behavioral deviation analysis is conducted to verify the robustness of the policy. The impact of different orbital characteristics, simulation conditions, and behavioral deviations on the game policy was analyzed. Simulation results show that the proposed method enables each spacecraft to formulate an effective game policy that satisfies all set constraints and has good robustness.  Objective  As the space environment becomes increasingly complex, space security has become a hot research area. The existence of a large amount of space debris and failed spacecraft poses a serious threat to high-value spacecraft in orbit. Therefore, the study of Orbital Pursuit-Evasion Game (OPEG) for non-cooperative target spacecraft has attracted widespread attention. Existing research focuses on OPEG for two spacecraft, but less on OPEG for multiple spacecraft. When there are more than two players in the game, zero-sum game design is not feasible, and it is difficult to solve using traditional methods. Furthermore, existing research ignores engineering dynamic constraints and simplifies or defines the dynamics as a two-dimensional scene when modeling the problem, which can cause considerable errors. To overcome the limitations of existing spacecraft game scenarios, this paper proposes a novel multi-spacecraft OPEG research scenario. The aim is to investigate the application of the MADRL algorithm in solving the approximate steady-state policies of each spacecraft in long-distance multi-spacecraft OPEG, highlighting the significant advantages of the MADRL algorithm in solving multi-spacecraft OPEG, and providing a feasible solution for truly realizing autonomous multi-spacecraft game play in the future.  Methods  The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm based on the Progressive Adversarial Training Framework (PATF) is used to solve the optimal game policy for each spacecraft in the Multi-Spacecraft OPEG. First, a multi-constrained multi-spacecraft OPEG model is established based on actual engineering constraints, and the problem is transformed into a Decentralized Partially Observable Markov Decision Process (Dec-POMDPs). Secondly, in order to improve the decision-making ability of agents in complex multi-agent game environments and formulate more robust game policies, a novel PATF is introduced, with different reward functions designed for the specific missions of each spacecraft. Finally, two sets of simulation examples with different orbital characteristics were set up, and four different simulation conditions were set up for simulation and behavioral deviation analysis was performed.  Results and Discussions  The MAPPO algorithm based on the PATF proposed in this paper is compared with the original MAPPO (Fig. 3). The results show that the proposed method can learn effective policies more quickly, reduce ineffective exploration, and achieve a higher final convergence reward value with less fluctuation in the reward curve. This also demonstrates that the PATF can significantly enhance the decision-making ability of agents, enabling them to formulate robust policies more effectively. Simulation verification was performed using two sets of examples in four different settings (Figs. 4, 5, 6, and 7). Simulation results (Tables 3 and 4) show that the proposed method performs well in both sets of examples. Furthermore, it was verified that when the pursuer and the interceptor are on the same orbital plane, the pursuer is more likely to be intercepted. When the interceptor and the target are not on the same orbital plane, the interceptor has a relatively easier time carrying out the interception mission. This paper also analyzes the situation where both sides of the game have behavioral biases, and models this by adding control noise. Simulation results (Tables 5 and 6) show that both sides adopt relatively conservative policies to counter the control noise. The game policy formulated by the method in this paper is an approximate steady-state policy. Behavioral deviations will lead to a decrease in one’s own payoff and an increase in the opponent's payoff, and the game policy has good robustness.  Conclusions  The method proposed in this paper can be well applied to solving the long-distance OPEG problem involving multiple spacecraft in non-coplanar elliptical orbits, enabling each spacecraft to formulate excellent game policies. The PATF facilitates better decision-making by the spacecraft in complex multi-spacecraft dynamic systems, with robust control policies developed by the pursuer and interceptors. The results also demonstrate the accuracy and effectiveness of the reward function design. Through two sets of examples and simulation results with different settings, the impact on the policies of both parties when the pursuer and interceptor have different orbital characteristics is analyzed. When interceptors have different maximum thrusts, the decision-making of each spacecraft changes accordingly. The behavior deviation analysis proves that the game policies of each spacecraft have good robustness. When one party’s behavior deviates, the approximate steady-state policy balance will change, resulting in a decrease in its own benefits and an increase in the other party’s benefits. The research scenario formulated in this paper expands the scope of existing research on multi-spacecraft game problems.
Communication, Computation, and Caching Resource Collaboration for Heterogeneous Artificial Intelligence Generated Content Service Provisioning
WU Mengru, GAO Yu, ZHAO Bo, XU Bo, SUN Hao, GUO Lei
Available online  , doi: 10.11999/JEIT251300
Abstract:
  Objective  In the Artificial Intelligence of Things (AIoT), Edge Servers (ESs) provide intelligent content generation services to AIoT devices by utilizing cached Artificial Intelligence Generated Content (AIGC) models. However, the limited computing resources and caching capacity of ESs make it difficult to support the large-scale caching demands of heterogeneous AIGC services. To address this issue, a communication, computation, and caching resource collaboration scheme is proposed based on a combined cloud-edge and edge-edge collaborative framework. The scheme considers three representative AIGC services: lightweight AIGC services, computation-intensive AIGC services, and preprocessing-based AIGC services. The objective is to minimize the total AIGC service latency through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions.  Methods  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. First, an AIGC service-oriented AIoT system model is established to incorporate both cloud-edge and edge-edge collaboration. An optimization problem is then formulated to minimize the total latency of AIGC services through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions. Because the formulated problem is non-convex, an Alternating Optimization (AO) algorithm is proposed. The original problem is decomposed into three subproblems. These subproblems are solved using the Successive Convex Approximation (SCA) method, Karush-Kuhn-Tucker (KKT) conditions, and an improved Harris Hawks Optimization (HHO) algorithm.  Results and Discussions  Simulation experiments compare the proposed joint optimization scheme with three baseline methods: Particle Swarm Optimization (PSO), fixed resource allocation, and random offloading and caching. First, the convergence of the proposed AO algorithm is verified (Fig. 2). The results show that the algorithm converges rapidly within a limited number of iterations across different subproblems. Second, increasing transmission bandwidth significantly reduces the total AIGC service latency (Fig. 3). This occurs because each device obtains more bandwidth resources for task transmission, and the ES can allocate more bandwidth to deliver generated content in the downlink. Furthermore, the total AIGC service latency decreases as the ES storage capacity increases for all schemes (Fig. 4). Greater storage capacity enables the ES to store more AIGC models, which reduces the transmission delay between the ES and the cloud server. Moreover, when the required floating-point operations per bit increase, the total AIGC service latency rises significantly across all schemes (Fig. 5). Finally, the total AIGC service latency decreases as the maximum transmit power of the Base Station (BS) increases (Fig. 6). This occurs because higher BS transmit power improves the downlink signal-to-noise ratio, which increases the downlink transmission rate and reduces overall service latency. The proposed scheme demonstrates better performance than the baseline schemes, particularly under high computational demand.  Conclusions  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. The objective is to minimize total AIGC service latency through joint optimization of the transmit power of AIoT devices and BSs, computing resource allocation, AIGC model deployment, and service offloading decisions under computation and caching resource constraints. Because the formulated problem is a mixed-integer nonlinear programming problem, an efficient AO algorithm is developed. The original optimization problem is decomposed into three subproblems, which are solved using the SCA algorithm, KKT conditions, and the HHO algorithm, respectively. Simulation results show that the proposed algorithm reduces the total AIGC service latency compared with the baseline schemes.
Adversarial Attacks on 3D Target Recognition Driven by Gradient Adaptive Adjustment
LIU Weiquan, SHEN Xiaoying, LIU Dunqiang, SUN Yanwen, CAI Guorong, ZANG Yu, SHEN Siqi, WANG Cheng
Available online  , doi: 10.11999/JEIT251264
Abstract:
  Objective   Robust environmental perception is essential for intelligent driving systems. Light Detection And Ranging (LiDAR) provides high-resolution 3D point cloud data and serves as a core information source for object detection and recognition. However, deep learning models for 3D point cloud recognition show notable vulnerability to adversarial attacks. Small, imperceptible perturbations can cause severe classification errors and threaten system safety. Existing attack methods have improved the Attack Success Rate (ASR), but the perturbations they generate often lack concealment, create outliers, and show poor imperceptibility because they do not adequately preserve the geometric structure of point clouds. This reduces their suitability for realistic security evaluation of optoelectronic perception systems. Developing an attack method that maintains a high success rate while preserving geometric consistency and imperceptibility is therefore critical. This study addresses this need by proposing a framework that incorporates point cloud geometry into perturbation generation.  Methods   A Gradient Adaptive Adjustment (GAA) adversarial attack method for 3D point cloud recognition is proposed. The framework (Fig. 2) includes three coordinated modules. The 3D Point Cloud Salient Region Extraction module evaluates decision-level vulnerability using Shapley value analysis to identify and rank point subsets with the strongest influence on classifier output. Perturbations are then concentrated in these sensitive regions. A curvature-weighted gradient mechanism integrates local geometric priors. For each point in the salient region, a local covariance matrix is computed from its k-nearest neighbors. Principal component analysis generates eigenvalues and eigenvectors, which are used to compute a curvature measure. A Gaussian kernel function produces curvature-dependent weights that are applied to backpropagated gradients. This suppresses perturbations in high-curvature areas and encourages them in low-curvature regions to preserve local shape morphology. A principal curvature direction constrained 0ptimization module further refines the perturbation direction. The weighted gradient is projected onto the principal curvature directions, and the projection components are fused using coefficients derived from the corresponding eigenvalues. This aligns the perturbation with natural geometric trends and avoids unnatural deformation. An adaptive optimization algorithm then minimizes a multi-objective loss balancing attack success, geometric similarity (via chamfer distance and hausdorff distance), and perturbation sparsity. The adversarial point cloud is iteratively updated based on the saliency map, curvature-weighted gradients, and principal direction constraints.  Results and Discussions   Experiments on ModelNet40, ShapeNetPart, and KITTI were conducted using PointNet, DGCNN, and PointConv. The GAA method showed strong performance. On ModelNet40 with PointNet, it achieved a 97.69% ASR with an average of 28 perturbed points, outperforming ten baselines such as AL-Adv (92.92% ASR, 40 points) and Kim et al. (89.38% ASR, 36 points) (Table 1). It also produced lower geometric distortion, as indicated by smaller Chamfer Distance and Hausdorff Distance values. Visual results (Fig. 4) show that GAA produces fewer outliers and more natural adversarial point clouds compared with methods such as AL-Adv. The method generalized well across architectures, reaching 99.78% ASR on DGCNN and 96.91% on PointConv (Table 2), with similar performance on ShapeNetPart (Table 3). Ablation experiments on the number of salient regions (K) showed consistent improvements in ASR and reduced geometric distortion as K increased from 1 to 6 (Table 4, Fig. 5), confirming the advantage of targeting multiple critical regions. Tests on the KITTI dataset demonstrated strong performance in real-world, noisy environments. The method maintained high ASRs, such as 99.33% on PointNet, with limited perturbations (Table 5). An ablation study on K indicated that K=4 offers an effective balance between success rate and perturbation cost for PointNet (Table 6).  Conclusions   This study presents a GAA method for adversarial attacks on 3D point cloud recognition. By combining a Shapley value-based saliency analyzer, a curvature-weighted gradient mechanism, and a principal curvature direction constraint, the method generates adversarial examples that achieve high attack success while preserving geometric consistency. Experiments show that GAA minimizes perceptual distortion and perturbs fewer points across datasets and models. The method provides a practical tool for vulnerability analysis and supports the development of more robust and secure optoelectronic perception systems for intelligent driving. Future work will examine robustness under adverse conditions and assess physical-world implications.
Blind Parameter Estimation Method for PSK Modulated Frequency-Hopping Signals Based on Improved Maximum Likelihood
ZHANG Tianhao, ZHANG Yushu, XU Zhongqiu, TANG Xinyi, DANG Wenhua, LI Guangzuo
Available online  , doi: 10.11999/JEIT260005
Abstract:
  Objective  Blind parameter estimation of non-cooperative Frequency-Hopping (FH) signals is a critical task in electronic reconnaissance and countermeasures. Estimation methods based on time-frequency analysis typically suffer from limited resolution or high computational complexity. Furthermore, methods based on compressive sensing rely heavily on the consistency between the predefined dictionary and the actual signal characteristics, and the estimation precision will be significantly compromised by grid mismatch or modulation-induced energy dispersion. Maximum Likelihood (ML)-based methods offer the advantage of high theoretical estimation accuracy with relatively low computational complexity. However, existing studies typically assume an ideal unmodulated signal model with a single frequency transition. Consequently, these ML-based methods suffer from severe model mismatch when processing FH signals with digital modulation, such as Phase Shift Keying (PSK), or multi-hop signals. Moreover, the conventional iterative solution of ML-based methods is prone to divergence or trapping in local optima. To address these limitations, this paper proposes an improved ML-based method for the blind parameter estimation of PSK-modulated FH signals.  Methods  To handle received multi-hop signals, a signal slicing technique based on the Short-Time Fourier Transform (STFT) is proposed to extract slices containing individual frequency transitions. Subsequently, to mitigate the model mismatch caused by digital modulation in conventional ML-based methods, a model-matching signal extraction approach based on the ML objective function is developed for PSK-modulated FH signals. Furthermore, a weighted iterative solving algorithm for ML estimation is designed to enhance convergence, thereby achieving robust and accurate estimation of frequency-hopping parameters.  Results and Discussions  To validate the effectiveness of the model-matching signal extraction approach, ablation experiments were carried out under various modulation schemes, including binary PSK (BPSK), quadrature PSK (QPSK), and 8-ary PSK (8PSK). The results indicate that the proposed approach (Group D) significantly reduces the Mean Square Error (MSE) of hopping frequency estimation compared to that without the proposed extraction (Group ND). These results demonstrate that the proposed method effectively mitigates the model mismatch (Fig. 5). Simulation results also illustrate that the designed weighted iterative algorithm achieves superior convergence performance compared with linear weighting and non-weighting schemes (Fig. 6). Moreover, the experiments verify the algorithm's insensitivity to initial frequency offsets, showing that it tolerates offsets of up to 2 MHz at SNR of -10 dB with little performance degradation (Fig. 7). Finally, comparative analysis with representative existing methods indicates that the proposed method outperforms the others in terms of estimation accuracy (Fig. 8).  Conclusions  To achieve blind parameter estimation for PSK-modulated FH signals, this paper proposes an improved ML-based method. By utilizing a signal slicing technique based on the STFT, the proposed method successfully extends the applicability of the ML-based estimator to continuous multi-hop signals. To mitigate the model mismatch induced by PSK modulation, a model-matching signal extraction approach is developed to isolate valid signal segments that conform to the ML model. Furthermore, a weighted iterative algorithm incorporating a dynamic weighting function is introduced to address the instability of the conventional iterative ML solver. Simulation results confirm that the proposed method effectively eliminates model mismatch and ensures superior convergence performance with insensitivity to initial frequency offsets. Moreover, it is shown to achieve high estimation precision for both hopping frequencies and hopping times.
A Semantic-Enhanced Cybersecurity Named Entity Recognition Approach Oriented to Lightweight Adaptation of Large Language Models
HU Ze, XU Tongwu, YANG Hongyu
Available online  , doi: 10.11999/JEIT251260
Abstract:
  Objective  Named Entity Recognition (NER) in the field of cybersecurity is a fundamental technology supporting threat intelligence analysis, vulnerability management, and security incident response. However, this field generally faces challenges such as dense technical terms, scarce labeled data, dynamic changes in entity categories, and highly complex semantic features, which make traditional deep learning models and existing Large Language Models (LLMs) significantly inadequate in terms of domain adaptability and semantic fusion capability. To address the aforementioned key issues while also considering the need for lightweight model deployment, this paper aims to construct a cybersecurity NER approach that can enhance domain semantic representation, improve the ability to identify rare entities, and apply to low-resource environments, providing a reliable technical path for intelligent threat analysis in cybersecurity scenarios.  Methods  To address the complex semantic features of cybersecurity texts, this paper proposes a semantically enhanced, lightweight, and LLMs-adaptable cybersecurity NER approach. The proposed approach uses LLM2Vec to achieve bidirectional semantic reconstruction of large model decoders and combines Low-Rank Adaptation (LoRA) for low-rank fine-tuning, so as to maintain deep semantic encoding capability while significantly reducing the amount of parameter updates. To address the challenges of sparse keywords and severe noise interference in cybersecurity texts, a sparse gated attention mechanism is introduced to strengthen keyword-focused feature extraction by dynamically selecting high-contribution cybersecurity terms through global gating and sparse inference. A SecRoBERTa-based semantic enhancement component is introduced, which utilizes a domain-pre-trained model to generate similar word embeddings, optimizes feature robustness in small-sample scenarios, and alleviates the challenges of identifying out-of-vocabulary words and low-frequency terms. Finally, a masked conditional random field is employed to constrain label transitions and guarantee BIO-compliant output sequences, achieving robust and consistent entity boundary prediction.  Results and Discussions  Extensive experiments were conducted on two public cybersecurity datasets, DNRTI and APTNER. The proposed approach achieved an F1 score of 91.91% on DNRTI, surpassing the previous state-of-the-art model by 2.14%. On APTNER, it reached an F1 score of 80.37%, outperforming the best baseline by 2.97%. Ablation studies confirmed the contribution of each key component: the Sparse Gated Attention mechanism improved F1 by 3.57% over standard Multi-Head Attention on DNRTI; the semantic enhancement module contributed a 2.32% F1 gain; and the MCRF (Masked Conditional Random Field) layer provided a 10.63% F1 improvement over traditional CRF (Conditional Random Field). The model also demonstrated efficient training and inference characteristics, aligning with its lightweight design goals.  Conclusions  This paper proposes a lightweight adaptation approach based on LLMs for NER in the cybersecurity domain, which effectively addresses the limitations of existing LLMs-based NER methods in domain adaptation and rare entity recognition. By integrating LLM2Vec and LoRA for lightweight fine-tuning, a sparse gated attention mechanism for domain feature fusion, and a SecRoBERTa-based semantic enhancement component for similar word precomputation, the proposed approach achieves high performance on DNRTI and APTNER datasets. The research provides an efficient technical path for NER tasks in low-resource cybersecurity scenarios and offers strong support for downstream tasks such as automated threat intelligence analysis.
FPGA Hybrid PLB Architecture for Highly Efficient Resource Utilization
WANG Yanlin, GAO Lijiang, YANG Haigang
Available online  , doi: 10.11999/JEIT260108
Abstract:
6-input look-up tables (LUTs) are frequently used in commercial Field-Programmable Gate Arrays (FPGAs) to build programmable logic blocks, while related experiments reveal that their average application in circuits is less than 30%, resulting in a significant waste of programmable resources. In this paper, the 6-input LUTs are fractured based on fracturable factors and recombined with different granularities to construct several new Hybrid Basic Logic Elements (HBLE). Based on HBLE, several novel Hybrid Programmable Logic Block (HPLB) architectures are proposed. Then the Programmable Logic Blocks (PLB) of Xilinx is replaced by several innovative HPLB architectures. Concurrently, a statistical evaluation algorithm for the mapped netlist is proposed. Finally, several HPLB architectures are experimentally verified and evaluated as appropriate. Experimental evaluations of the three enhanced architectures show that the HPLBs achieve an average area reduction of more than 30% when compared to Xilinx’s PLBs without adding more input ports. The hybrid HPLB architectures constructed with a fracturable factor N=3 produces the best optimization results when taking into account both HPLB utilization and area optimization. Based on the MCNC and VTR benchmarks, resource consumption increased by an average of 8.27% and 27.64%, respectively, thereby improving FPGA logic efficiency.  Objective  Currently, modern commercial FPGA architectures employ 6-LUTs as the fundamental building blocks for Basic Logic Elements (BLEs). Only about 30% of the Logic Elements (LEs) in the circuit are ultimately translated to 6-LUTs when mapping 6-LUT BLEs, according to experimental results. Nevertheless, more than half of the logic resources are wasted when 6-LUTs implement functions with inputs smaller than 6. Programmable resources will unavoidably be significantly wasted as a result. A circuit design mapped to 100 4-LUTs can be mapped to 78 6-LUTs during 6-LUT mapping studies, according to experimental data, with the {6,5,4,3,2}-LUT function distribution being {23,32,17,9,13}. The findings indicate that only around 25% of the 6-LUTs are ultimately mapped to 6-input functions, with the remaining 6-LUTs being underutilized. This illustrates even more how inefficient technical mapping is for LUTs with large input K.Methods The fracturable factor N, which is the number of sub-LUTs that may be obtained from a single LUT, characterizes the fracturable and reconfigurable nature of LUT architectures in FPGAs. Motivated by this, we decompose a 6-LUT into several granularities according to the fracturable factor in order to address the previously described problem of low resource utilization. Three novel hybrid-granularity divisible logic (HBLE) structures are created by connecting and reconfiguring the resultant sub-LUTs with additional input ports and multiplexer modules. We shall now investigate how FPGA performance is optimized by these three HBLE topologies. We shall now investigate how FPGA performance is optimized by these three HBLE topologies. One undivided 6-LUT and one divisible 6-LUT, divided into two 5-LUTs with a divisibility factor N=2, make up the HBLE2 structure. One undivided 6-LUT and one divisible 6-LUT, divided into one 5-LUT and two 4-LUTs, with a divisibility factor N=3, are included in the HBLE3 structure. One undivided 6-LUT and one divisible 6-LUT, which divides into four 4-LUTs with a divisibility factor N=4, make up the HBLE4 structure. Adder units are supported by all three HBLE structures, allowing for both latched and direct combinational logic output. Additionally, they allow direct latched output by avoiding combinational logic. A Hybrid Programmable Logic Block (HPLB) is a novel structure created by merging several HBLEs. The MCNC circuit set and the VTR circuit set, the two most well-known academic circuit benchmarks (BMs), are chosen for experimental assessment. A Xilinx Virtex-7 FPGA is used to map each circuit set. The mapped netlist is then used to tally the kinds and numbers of LUTs that were utilized. The minimum number of CLBs needed is found once the data has been arranged using the corresponding greedy algorithms. Since each Xilinx CLB has eight 6-LUTs, the greedy approach uses # Total LUT Number / 8 to determine the smallest number of CLBs needed following BM mapping. In order to guarantee similar conditions, each structure also needs to be sorted using the greedy algorithm after Xilinx’s CLB structure is replaced with the HPLB structure suggested in this research. This results in the bare minimum of HPLBs needed. It is not possible to use every LUT in the mapped CLBs during actual packing owing to routing constraints. As a result, the smallest value that may be achieved in a theoretical optimization scenario is represented by the optimized result that is acquired following greedy algorithm restructuring.  Results and Discussions  The average number of HPLBs needed for both HPLB2 and HPLB3 structures drops by about 8% when CLB structures are swapped out for HPLBs in order to map the MCNC circuit set. However, the number of HPLBs needed increases by more than 30% on average as a result of the HPLB4 structure. The needed count is smaller when HPLBs are used in place of CLBs for mapping the VTR circuit set. On average, the HPLB2 and HPLB4 counts drop by less than 10%, whereas the HPLB3 count drops by around 30%. This enables SRAM scheduling and complete input pin use. On the other hand, because of resource waste, the uniform CLB structure results in higher CLB requirements when implementing functions with a tiny LUT input K. The HPLB4 structure performs worse than the HPLB3 structure, according to post-mapping HPLB counts. Both the MCNC and VTR circuit sets achieve average area reduction ratios over 30%, according to analysis of post-mapping area optimization. All three HPLB structures attained area optimization ratios of about 31% on the MCNC test set. Different optimization effects were seen in the VTR test circuit set: HPLB2 produced an average area reduction of 30.63%, whereas HPLB4 produced an average decrease of 51.21%. The HPLB2 structure produced a 45.22% area reduction, even though its optimization effect was marginally less than that of HPLB4. A thorough examination of the area optimization results showed that a higher divisibility factor N produces more noticeable benefits for integrating small-scale LUTs in circuits, resulting in higher area reduction ratios from the enhanced architectures.  Conclusions  In order to solve the issue of low resource utilization in 6-LUTs, this research proposes three split granularity-based HPLB enhancement architectures. In addition to establishing an assessment procedure and matching algorithms for the enhanced structures, these HPLBs take the place of Xilinx’s CLB structure in order to examine the new structure’s benefits in resource utilization. Based on the proportion differences of different LUTs in the post-mapping netlist, evaluation experiments using the MCNC and VTR circuit test suites show that, although HPLB4 achieves significant area optimization, it requires additional HPLBs, resulting in increased interconnect area. While both HPLB2 and HPLB3 structures obtain average area optimizations over 30%, HPLB3 produces a significantly greater HPLB count and area optimization than HPLB2 as the test circuit scale grows. Thus, after replacing the CLB structure, the HPLB3 structure provides a more balanced optimization impact, greatly improving the utilization of programmable resources when taking into account the combined aspects of HPLB usage count and area optimization.
Efficient and Verifiable Ciphertext Retrieval Scheme Based on Trusted Execution Environment
WU Axin, FENG Dengguo, ZHANG Min, CHI Jialin, YI Yuling
Available online  , doi: 10.11999/JEIT251358
Abstract:
The ciphertext retrieval mechanism enables retrieval functionality over encrypted data. Symmetric Searchable Encryption (SSE) is a critical branch of ciphertext retrieval. However, due to considerations such as saving computing power, cloud servers may return incorrect or incomplete results. Moreover, attackers can also exploit these leaked information from search and access patterns to reconstruct the keyword details. Therefore, it is necessary and meaningful to protect the privacy of search and access patterns while achieving result verifiability. Nevertheless, existing verifiable SSE schemes that support search and access pattern privacy typically rely on keyword traversal mechanisms and their verification mechanisms are inefficient, which impose high computational and communication overheads on users. To address the above performance bottlenecks, this paper introduces an efficient and verifiable ciphertext retrieval scheme based on Trusted Execution Environment (TEE). To improve the efficiency of ciphertext retrieval, this scheme employs the collaborative implementation of hardware-level security isolation and oblivious data rearrangement to achieve keyword trapdoor size independent of the size of the keyword dictionary. Meanwhile, the correctness of the returned results is verified by embedding random numbers and blinding polynomial constant terms. Thanks to these designs, the scheme achieves significant efficiency improvements. Specifically, firstly, this scheme ensures that the size of keyword trapdoors depends solely on the number of query keywords, not the global dictionary size, effectively minimizing communication and computational costs. Secondly, this scheme requires storing only two random numbers to enable verifiability, substantially minimizing local storage overhead for users. Thirdly, the adoption of techniques, such as enabling data users to retrieve results via single-server and single-round interaction and leveraging symmetric homomorphic encryption, further enhances operational efficiency. Additionally, confidential computing within TEE weakens the security assumptions and trust level towards TEE. After formally proving the security of the proposed scheme using simulation-based methods, this paper has conducted a comprehensive performance evaluation. The evaluation results confirm that this scheme is significantly more efficient than other schemes with the same functionalities.
Performance Analysis and Rapid Prediction of Long-range Underwater Acoustic Communications in Uncertain Deep-sea Environments
CHEN Xiangmei, TAI Yupeng, WANG Haibin, HU Chenghao, WANG Jun, WANG Diya
Available online  , doi: 10.11999/JEIT251244
Abstract:
  Objective  In complex and dynamically changing deep-sea environments, the performance of underwater acoustic communications shows substantial variability. Feedback-based channel estimation and parameter adaptation are impractical in long-range scenarios because platform constraints prevent reliable feedback channels and the slow propagation of sound introduces significant delay. In typical long-range systems, environmental dynamics are often ignored and communication parameters are selected heuristically, which frequently leads to mismatches with actual channel conditions and causes communication failures or reduced efficiency. Predictive methods able to assess performance in advance and support feed-forward parameter adjustment are therefore required. This study proposes a deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications under uncertain environmental conditions to enable efficient and reliable parameter–channel matching without feedback.  Methods  A feed-forward method for underwater acoustic communication performance analysis and rapid prediction is developed using deep-learning-based sound-field uncertainty estimation. A neural network is first used to estimate probability distributions of Transmission Loss (TL PDFs) at the receiver under dynamic environments. TL PDFs are then mapped to probability distributions of the Signal-to-Noise Ratio (SNR PDFs), enabling communication performance evaluation without real-time feedback. Statistical channel capacity and outage capacity are analyzed to characterize the theoretical upper limits of achievable rates in dynamic conditions. Finally, by integrating the SNR distribution with the bit-error-rate characteristics of a representative deep-sea single-carrier communication system under the corresponding channel, a rate–reliability prediction model is constructed. This model estimates the probability of reliable communication at different data rates and serves as a practical tool for forecasting link performance in highly dynamic and feedback-limited underwater acoustic environments.  Results and Discussions  The method is validated using simulation data and sea trial data. The TL PDFs predicted by the deep learning model show strong consistency with the traditional Monte Carlo (MC) method across multiple receiver locations (Fig. 6). Under identical computational settings, deep-learning-based TL PDF prediction reduces computation time by 2\begin{document}$ \sim $\end{document}3 orders of magnitude compared with the MC method. The chained mapping from TL PDFs to SNR PDFs and then to channel capacity metrics accurately represents the probabilistic features of communication performance under uncertain conditions (Fig. 7 and Fig. 8). The rate–reliability curves derived from the deep-learning-based TL PDFs are highly consistent with MC-based results. In the high sound-intensity region, prediction errors for reliable communication probabilities across data rates range from 0.1% to 3%, and in the low sound-intensity region errors are approximately 0.3% to 5% (Fig. 12). Sea trial results further indicate that predicted rate–reliability performance agrees well with measured data. In the convergence zone, deviations between predicted and measured reliability probabilities at each rate range from 0.9% to 4%, and in the shadow zone from 1% to 9% (Fig. 18). Under a 90% reliability requirement, the maximum achievable rates predicted by the method match the measurements in both the convergence and shadow zones, demonstrating accuracy and practical applicability in complex channel environments.  Conclusions  A deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications in uncertain deep-sea environments is developed and validated. The framework builds a chained mapping from environmental parameters to TL PDFs, SNR PDFs, and communication performance metrics, enabling quantitative capacity assessment under dynamic ocean conditions. Predictive “rate–reliability’’ profiles are obtained by integrating probabilistic propagation characteristics with the performance of a representative deep-sea single-carrier system under the corresponding channel, providing guidance for parameter selection without feedback. Sea trial results confirm strong agreement between predicted and measured performance. The proposed approach offers a technical pathway for feed-forward performance analysis and dynamic adaptation in long-range deep-sea communication systems, and can be extended to other communication scenarios in dynamic ocean environments.
Inverse Design of a Silicon-Based Compact Polarization Splitter-Rotator
HUI Zhanqiang, ZHANG Xinglong, HAN Dongdong, LI Tiantian, GONG Jiamin
Available online  , doi: 10.11999/JEIT250858
Abstract:
  Objective  The Polarization Splitter-Rotator (PSR) is a key device used to control the polarization state of light in Photonic Integrated Circuits (PICs). Device size has become a major constraint on integration density in PICs. Traditional design methods are time-consuming and tend to yield larger device footprints. Inverse design, by contrast, determines structural parameters through optimization algorithms according to target performance and enables compact devices to be obtained while maintaining functionality. This strategy is now applied to wavelength and mode division multiplexers, all-optical logic gates, power splitters, and other integrated photonic components. The objective of this work is to use inverse design to address size limitations in silicon-based PSRs by combining the Momentum Optimization algorithm with the Adjoint Method. This combined approach improves the integration level of PICs and provides a feasible pathway for the miniaturization of other photonic devices.  Methods  The design region is defined on a 220 nm Silicon-on-Insulator (SOI) wafer and is discretized into 25×50 cylindrical elements. Each element has a 50 nm radius, a 150 nm height, and an initial relative permittivity of 6.55. The adjoint method is used to obtain gradient information across the design region, and this gradient is processed with the Momentum Optimization algorithm. The relative permittivity of each element is then updated according to the processed gradient. During optimization, the momentum factor is dynamically adjusted with the iteration number to accelerate convergence, and a linear bias is applied to guide the permittivity toward the values of silicon and air as the iterations progress. After optimization, the elements are binarized based on their final permittivity: values below 6.55 are assigned to air, whereas values above 6.55 are assigned to silicon. This results in a structure containing irregularly distributed air holes. To compensate for performance loss introduced during binarization, the etching depth of air holes with pre-binarization permittivity between 3 and 6.55 is optimized. Adjacent air holes are merged to reduce fabrication errors. The final device consists of air holes with five radii, among which three larger-radius types are selected for further refinement. Their etching radii and depths are optimized to recover remaining performance loss. Device performance is evaluated through numerical analysis. Calculated parameters include Insertion Loss (IL), Crosstalk (CT), Polarization Extinction Ratio (PER), and bandwidth. Tolerance analysis is also conducted to assess robustness under fabrication variations.  Results and Discussions   A compact PSR is designed on a 220 nm SOI wafer with dimensions of 5 μm in length and 2.5 μm in width. During optimization, the momentum factor in the Momentum Optimization algorithm is dynamically adjusted. A larger momentum factor is applied in the early stage to accelerate escape from local maxima or plateau regions, whereas a smaller momentum factor is used in later iterations to increase the weight of the current gradient. Compared with other optimization strategies, this algorithm requires only 20%~33% of the iteration count needed by alternative methods to reach a Figure of Merit (FOM) of 1.7, which improves optimization efficiency. Numerical analysis shows that the device achieves stable performance across the 1 520~1 575 nm wavelength range. The IL remains low (TM0 < 1 dB, TE0 < 0.68 dB), and the CT is effectively suppressed (TM0 < –23 dB, TE0 < –25.2 dB). The PER is high (TM0 > 17 dB, TE0 > 28.5 dB). Tolerance analysis indicates strong robustness to fabrication variations. Within the 1 520~1 540 nm range, performance remains stable under etching depth offsets of ±9 nm and etching radius offsets of ±5 nm, demonstrating reliable manufacturability.  Conclusions   Numerical analysis demonstrates that combining the adjoint method with the Momentum Optimization algorithm is a feasible strategy for designing an integrated PSR. The design principle relies on controlling light propagation through adjustments to the relative permittivity, which determine the distribution and placement of air holes to achieve polarization splitting and rotation. Compared with traditional design approaches, inverse design uses the design region more efficiently and enables a more compact device structure. The proposed PSR is markedly smaller and shows enhanced fabrication tolerance. It is suitable for future large-scale PICs and provides useful guidance for the miniaturization of other photonic devices.
Conditional Generative Adversarial Networks-based Channel Estimation for ISAC-RIS System
LIU Yu, ZHENG Zelin, LIU Gang
Available online  , doi: 10.11999/JEIT251168
Abstract:
  Objective  In RIS-assisted ISAC systems, accurate channel estimation is crucial to ensure reliable operation. Although traditional deep learning methods can partially address the channel estimation problem, their generalization ability and estimation accuracy remain limited in complex multi-user channel environments. To tackle these challenges, this paper proposes a two-stage channel estimation method based on Conditional Generative Adversarial Network(CGAN) for RIS-assisted multi-user ISAC systems, aiming to enhance both the accuracy and stability of channel estimation.  Methods  This paper proposes a two-stage channel estimation method based on CGAN for estimating the SAC channels in RIS-assisted multi-user ISAC systems. By adjusting the switching states of the RIS, the overall estimation problem is decomposed into subproblems, enabling sequential estimation of the direct and reflected channels. Within the proposed CGAN framework, the adversarial training between the generator and discriminator allows the model not only to learn the mapping relationship between the observed signals and the true channels but also to optimize the output according to the discriminator’s feedback, thereby effectively improving both training efficiency and estimation accuracy.  Results and Discussions  Extensive simulation experiments were conducted to verify the effectiveness of the proposed method. First, the estimation performance of the SAC channel under different SNR conditions was compared. The results demonstrate that the proposed CGAN-based method achieves significantly better NMSE performance than the LS benchmark and traditional models such as FNN and ELM (Fig. 4). Then, the impact of increasing the number of antennas and RIS elements on SAC channel estimation performance was investigated. Compared with the LS benchmark, the proposed CGAN method consistently maintains superior performance under various SNR conditions (Figs. 5 and 6).  Conclusions  This paper investigates the channel estimation problem in RIS-assisted multi-user ISAC systems and proposes a two-stage channel estimation method based on CGAN. By adjusting the switching states of the RIS and employing adversarial training between the generator and discriminator networks, the proposed method achieves accurate estimation of the SAC channel. Simulation results demonstrate that, under various SNR conditions and channel dimensions, the CGAN-based estimation method exhibits strong generalization capability and significantly outperforms the benchmark schemes in estimation accuracy. Therefore, it shows great potential as an effective solution for enhancing system stability and efficiency.
Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning in Wireless Networks
LIU Jingyuan, MA Ke, XU Runchen, CHANG Zheng
Available online  , doi: 10.11999/JEIT251221
Abstract:
  Objective  Multimodal Federated Learning (MFL) uses complementary information from multiple modalities, yet in wireless edge networks it is restricted by limited energy and frequent missing modalities because many clients store only images or only reports. This study presents Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning (CREEMFL), which applies selective completion and joint communication–computation optimization to reduce training energy under latency and wireless constraints.  Methods  CREEMFL completes part of the incomplete samples by querying a public multimodal subset, and processes the remaining samples through zero padding. Each selected user downloads the global model, performs image-to-text or text-to-image retrieval, conducts local multimodal training, and uploads model updates for aggregation. An energy–delay model couples local computation and wireless communication and treats the required number of global rounds as a function of retrieval ratios. Based on this model, an energy minimization problem is formulated and solved using a two-layer algorithm with an outer search over retrieval ratios and an inner optimization of transmission time, Central Processing Unit (CPU) frequency, and transmit power.  Results and Discussions  Simulations on a single-cell wireless MFL system show that increasing the ratio of completing text from images improves test accuracy and reduces total energy. In contrast, a large ratio of completing images from text provides limited accuracy gain but increases energy consumption (Fig. 3, Fig. 4). Compared with four representative baselines, CREEMFL achieves shorter completion time and lower total energy across a wide range of maximum average transmit powers (Fig. 5, Fig. 6). For CREEMFL, increased system bandwidth further reduces completion time and energy consumption (Fig. 7, Fig. 8). Under different user modality compositions, CREEMFL also attains higher test accuracy than local training, zero padding, and cross-modal retrieval without energy optimization (Fig. 9).  Conclusions  CREEMFL integrates selective cross-modal retrieval and joint communication–computation optimization for energy-efficient MFL. By treating retrieval ratios as variables and modeling their effect on global convergence rounds, it captures the coupling between per-round costs and global training progress. Simulations verify that CREEMFL reduces training completion time and total energy while preserving classification accuracy in resource-constrained wireless edge networks.
Finite-time Adaptive Sliding Mode Control of Servo Motors Considering Frictional Nonlinearity and Unknown Loads
ZHANG Tianyu, GUO Qinxia, YANG Tingkai, GUO Xiangji, MING Ming
Available online  , doi: 10.11999/JEIT250521
Abstract:
  Objective  Ultra-fast laser processing with an infinite field of view requires servo motor systems with superior tracking accuracy and robustness. However, such systems are highly nonlinear and affected by coupled unknown load disturbances and complex friction, which constrain the performance of conventional controllers. Although Sliding Mode Control (SMC) exhibits inherent robustness, traditional SMC and observer designs cannot achieve accurate finite-time disturbance compensation under strong nonlinearities, thus limiting high-speed and high-precision trajectory tracking. To address this limitation, a novel finite-time adaptive SMC approach is proposed to ensure rapid and precise angular position tracking within a finite time, satisfying the stringent synchronization requirements of advanced laser processing systems.  Methods  A novel control strategy is developed by integrating an adaptive disturbance observer fused with a Radial Basis Function Neural Network (RBFNN) and finite-time Sliding Mode Control (SMC). First, the unknown load disturbance and complex frictional nonlinear dynamics are combined into a unified "lumped disturbance" term, improving model generality and the ability to represent real operating conditions. Second, a finite-time adaptive disturbance observer is constructed to estimate this lumped disturbance. The observer utilizes the universal approximation capability of the RBFNN to learn and approximate the dynamic characteristics of unknown disturbances online. Simultaneously, a finite-time adaptive law based on the error norm is introduced to update the neural network weights in real time, ensuring rapid and accurate finite-time estimation of the lumped disturbance while reducing dependence on precise model parameters. Based on this design, a finite-time SMC is developed. The controller uses the observer’s disturbance estimation as a feedforward compensation term, incorporates a carefully formulated finite-time sliding surface and equivalent control law, and introduces a saturation function to suppress control input chattering. A suitable Lyapunov function is then constructed, and the finite-time stability theory is rigorously applied to prove the practical finite-time convergence of both the adaptive observer and the closed-loop control system, guaranteeing that the system tracking error converges to a bounded neighborhood near the origin within finite time.  Results and Discussions  To verify the effectiveness and superiority of the proposed control strategy, a typical Permanent Magnet Synchronous Motor (PMSM) servo system model is constructed in the MATLAB environment, and a simulation scenario with desired trajectories of varying frequencies is established. The proposed method is comprehensively compared with the widely used Proportional–Integral (PI) control and the advanced method reported in reference [7]. Simulation results demonstrate the following: 1. Tracking performance: Under various reference trajectories, the proposed controller enables the system to accurately follow the target trajectory with a tracking error substantially smaller than that of the PI controller. Compared with the method in reference [7], it achieves smoother responses and smaller residual errors, effectively eliminating the chattering observed in some operating conditions of the latter. 2 Disturbance rejection and robustness: The adaptive disturbance observer based on the RBFNN rapidly and effectively learns and compensates for the lumped disturbance composed of unknown load variations and frictional nonlinearities. Even in the presence of these disturbances, the proposed controller maintains high-precision trajectory tracking, demonstrating strong disturbance rejection and robustness to system parameter variations. 3. Control input characteristics: Compared with the reference methods, the control signal of the proposed approach quickly stabilizes after the initial transient phase, effectively suppressing chattering caused by high-frequency switching. The amplitude range of the control input remains reasonable, facilitating practical actuator implementation. 4. Comprehensive evaluation: Based on multiple error performance indices, including Integral Squared Error (ISE), Integral Absolute Error (IAE), Time-weighted Integral Absolute Error (ITAE), and Time-weighted Integral Squared Error (ITSE), the proposed controller consistently outperforms both PI control and the method in reference [7]. It demonstrates comprehensive advantages in suppressing transient errors rapidly and reducing overall error accumulation. The method also improves steady-state accuracy and achieves a balanced response speed with effective noise attenuation. 5. Observer performance: The RBFNN weight norm estimation converges rapidly and stabilizes at a low level after initial adaptation, confirming the effectiveness of the proposed adaptive law and the learning efficiency of the observer.  Conclusions  A finite-time sliding mode control strategy with an adaptive disturbance observer is proposed for servo systems used in ultra-fast laser processing. The method models unknown load disturbances and frictional nonlinearities as a lumped disturbance term. An adaptive observer, integrating an RBF neural network with a finite-time mechanism, accurately estimates this disturbance for real-time compensation. Based on the observer, a finite-time SMC law is formulated, and the practical finite-time stability of the closed-loop system is theoretically proven. Simulations conducted on a permanent magnet synchronous motor platform confirm that the proposed approach achieves superior tracking accuracy, robustness, and control smoothness compared with conventional PI and existing advanced methods. This work offers an effective solution for achieving high-precision control in nonlinear systems subject to strong disturbances.
Breakthrough in Solving NP-Complete Problems Using Electronic Probe Computers
XU Jin, YU Le, YANG Huihui, JI Siyuan, ZHANG Yu, YANG Anqi, LI Quanyou, LI Haisheng, ZHU Enqiang, SHI Xiaolong, WU Pu, SHAO Zehui, LENG Huang, LIU Xiaoqing
Available online  , doi: 10.11999/JEIT250352
Abstract:
This study presents a breakthrough in addressing NP-complete problems using a newly developed Electronic Probe Computer (EPC60). The system employs a hybrid serial–parallel computational model and performs large-scale parallel operations through seven probe operators. In benchmark tests on 3-coloring problems in graphs with 2,000 vertices, EPC60 achieves 100% accuracy, outperforming the mainstream solver Gurobi, which succeeds in only 6% of cases. Computation time is reduced from 15 days to 54 seconds. The system demonstrates high scalability and offers a general-purpose solution for complex optimization problems in areas such as supply chain management, finance, and telecommunications.  Objective   NP-complete problems pose a fundamental challenge in computer science. As problem size increases, the required computational effort grows exponentially, making it infeasible for traditional electronic computers to provide timely solutions. Alternative computational models have been proposed, with biological approaches—particularly DNA computing—demonstrating notable theoretical advances. However, DNA computing systems continue to face major limitations in practical implementation.  Methods  Computational Model: EPC is based on a non-Turing computational model in which data are multidimensional and processed in parallel. Its database comprises four types of graphs, and the probe library includes seven operators, each designed for specific graph operations. By executing parallel probe operations, EPC efficiently addresses NP-complete problems.Structural Features:EPC consists of four subsystems: a conversion system, input system, computation system, and output system. The conversion system transforms the target problem into a graph coloring problem; the input system allocates tasks to the computation system; the computation system performs parallel operations via probe computation cards; and the output system maps the solution back to the original problem format.EPC60 features a three-tier hierarchical hardware architecture comprising a control layer, optical routing layer, and probe computation layer. The control layer manages data conversion, format transformation, and task scheduling. The optical routing layer supports high-throughput data transmission, while the probe computation layer conducts large-scale parallel operations using probe computation cards.  Results and Discussions  EPC60 successfully solved 100 instances of the 3-coloring problem for graphs with 2,000 vertices, achieving a 100% success rate. In comparison, the mainstream solver Gurobi succeeded in only 6% of cases. Additionally, EPC60 rapidly solved two 3-coloring problems for graphs with 1,500 and 2,000 vertices, which Gurobi failed to resolve after 15 days of continuous computation on a high-performance workstation.Using an open-source dataset, we identified 1,000 3-colorable graphs with 1,000 vertices and 100 3-colorable graphs with 2,000 vertices. These correspond to theoretical complexities of O(1.3289n) for both cases. The test results are summarized in Table 1.Currently, EPC60 can directly solve 3-coloring problems for graphs with up to n vertices, with theoretical complexity of at least O(1.3289n).On April 15, 2023, a scientific and technological achievement appraisal meeting organized by the Chinese Institute of Electronics was held at Beijing Technology and Business University. A panel of ten senior experts conducted a comprehensive technical evaluation and Q&A session. The committee reached the following unanimous conclusions:1. The probe computer represents an original breakthrough in computational models.2. The system architecture design demonstrates significant innovation.3. The technical complexity reaches internationally leading levels.4. It provides a novel approach to solving NP-complete problems.Experts at the appraisal meeting stated, “This is a major breakthrough in computational science achieved by our country, with not only theoretical value but also broad application prospects.” In cybersecurity, EPC60 has also demonstrated remarkable potential. Supported by the National Key R&D Program of China (2019YFA0706400), Professor Xu Jin’s team developed an automated binary vulnerability mining system based on a function call graph model. Evaluation of the system using the Modbus Slave software showed over 95% vulnerability coverage, far exceeding the 75 vulnerabilities detected by conventional depth-first search algorithms. The system also discovered a previously unknown flaw, the “Unauthorized Access Vulnerability in Changyuan Shenrui PRS-7910 Data Gateway” (CNVD-2020-31406), highlighting EPC60’s efficacy in cybersecurity applications.The high efficiency of EPC60 derives from its unique computational model and hardware architecture. Given that all NP-complete problems can be polynomially reduced to one another, EPC60 provides a general-purpose solution framework. It is therefore expected to be applicable in a wide range of domains, including supply chain management, financial services, telecommunications, energy, and manufacturing.  Conclusions   The successful development of EPC offers a novel approach to solving NP-complete problems. As technological capabilities continue to evolve, EPC is expected to demonstrate strong computational performance across a broader range of application domains. Its distinctive computational model and hardware architecture also provide important insights for the design of next-generation computing systems.
Personalized Federated Learning Method Based on Collation Game and Knowledge Distillation
SUN Yanhua, SHI Yahui, LI Meng, YANG Ruizhe, SI Pengbo
Available online  , doi: 10.11999/JEIT221203
Abstract:
To overcome the limitation of the Federated Learning (FL) when the data and model of each client are all heterogenous and improve the accuracy, a personalized Federated learning algorithm with Collation game and Knowledge distillation (pFedCK) is proposed. Firstly, each client uploads its soft-predict on public dataset and download the most correlative of the k soft-predict. Then, this method apply the shapley value from collation game to measure the multi-wise influences among clients and quantify their marginal contribution to others on personalized learning performance. Lastly, each client identify it’s optimal coalition and then distill the knowledge to local model and train on private dataset. The results show that compared with the state-of-the-art algorithm, this approach can achieve superior personalized accuracy and can improve by about 10%.
The Range-angle Estimation of Target Based on Time-invariant and Spot Beam Optimization
Wei CHU, Yunqing LIU, Wenyug LIU, Xiaolong LI
Available online  , doi: 10.11999/JEIT210265
Abstract:
The application of Frequency Diverse Array and Multiple Input Multiple Output (FDA-MIMO) radar to achieve range-angle estimation of target has attracted more and more attention. The FDA can simultaneously obtain the degree of freedom of transmitting beam pattern in angle and range. However, its performance is degraded due to the periodicity and time-varying of the beam pattern. Therefore, an improved Estimating Signal Parameter via Rotational Invariance Techniques (ESPRIT) algorithm to estimate the target’s parameters based on a new waveform synthesis model of the Time Modulation and Range Compensation FDA-MIMO (TMRC-FDA-MIMO) radar is proposed. Finally, the proposed method is compared with identical frequency increment FDA-MIMO radar system, logarithmically increased frequency offset FDA-MIMO radar system and MUltiple SIgnal Classification (MUSIC) algorithm through the Cramer Rao lower bound and root mean square error of range and angle estimation, and the excellent performance of the proposed method is verified.
Satellite Navigation
Research on GRI Combination Design of eLORAN System
LIU Shiyao, ZHANG Shougang, HUA Yu
Available online  , doi: 10.11999/JEIT201066
Abstract:
To solve the problem of Group Repetition Interval (GRI) selection in the construction of the enhanced LORAN (eLORAN) system supplementary transmission station, a screening algorithm based on cross interference rate is proposed mainly from the mathematical point of view. Firstly, this method considers the requirement of second information, and on this basis, conducts a first screening by comparing the mutual Cross Rate Interference (CRI) with the adjacent Loran-C stations in the neighboring countries. Secondly, a second screening is conducted through permutation and pairwise comparison. Finally, the optimal GRI combination scheme is given by considering the requirements of data rate and system specification. Then, in view of the high-precision timing requirements for the new eLORAN system, an optimized selection is made in multiple optimal combinations. The analysis results show that the average interference rate of the optimal combination scheme obtained by this algorithm is comparable to that between the current navigation chains and can take into account the timing requirements, which can provide referential suggestions and theoretical basis for the construction of high-precision ground-based timing system.