Latest Articles

Articles in press have been peer-reviewed and accepted, which are not yet assigned to volumes/issues, but are citable by Digital Object Identifier (DOI).
Display Method:
Accelerated Broadband Electromagnetic Scattering Analysis via ACA-Driven Measurement Matrix Interpolation
WANG Zhonggen, WU Chenggang, NIE Wenyan, SUN Yufa
Available online  , doi: 10.11999/JEIT260392
Abstract:
  Objective  Broadband electromagnetic scattering analysis is widely used in radar target recognition, stealth technology, and microwave imaging. Although the Method of Moments (MoM) provides high computational accuracy, it incurs substantial computational and memory costs for electrically large or geometrically complex targets because full impedance matrices must be constructed and solved. Existing acceleration techniques, including the MultiLevel Fast Multipole Method (MLFMM) and Adaptive Cross Approximation (ACA), reduce the computational burden but still require repeated matrix construction and equation solving at every frequency during wideband analysis. Methods such as Asymptotic Waveform Evaluation (AWE), Model-Based Parameter Estimation (MBPE), and impedance matrix interpolation have been proposed to reduce this redundancy. However, AWE is prone to error accumulation over wide frequency bands, MBPE requires expensive initial sampling, and conventional impedance matrix interpolation still requires the computation of full high-dimensional impedance matrices at the sampling frequencies. More recently, Compressive Sensing Method of Moments (CS-MoM) and its extension, CS-HBFM, have improved wideband analysis by employing Hyper-Basis Functions (HBFs). By constructing Characteristic Mode Basis Functions (CMBFs) only once at the highest frequency, CS-HBFM eliminates repeated basis-function generation. Nevertheless, existing CS-HBFM methods rely on nondeterministic random or uniform sampling, require expensive large-scale matrix-vector products, and repeatedly reconstruct and solve impedance equations throughout the frequency sweep.  Methods  A CS-ACA-MMI framework is proposed for broadband electromagnetic scattering analysis by combining dual ACA decomposition with Measurement Matrix Interpolation (MMI). First, CMBFs are constructed at the highest frequency, and dominant HBFs are selected according to the Modal Significance (MS) criterion. ACA is then applied to the full impedance matrix to extract deterministic row indices corresponding to the dominant Rao-Wilton-Glisson (RWG) basis functions. These indices are reused throughout the frequency band, eliminating nondeterministic sampling and repeated index extraction. Second, four sampling frequencies are selected using Chebyshev-Lobatto nodes. Low-dimensional measurement matrices are constructed directly from the extracted row indices, avoiding the generation of full high-dimensional impedance matrices. The measurement impedance elements at the sampling frequencies are corrected according to the geometric distance, interpolated to the target frequency, and then restored to the actual measurement impedance elements, thereby eliminating repeated construction of measurement matrices during frequency sweeping. Third, ACA is applied to the far-field component of the interpolated measurement matrix, converting large-scale matrix-vector products into low-dimensional matrix multiplications. The near-field sensing matrix is obtained directly by multiplying the measurement matrix by the basis functions, enabling rapid construction of the complete sensing matrix. Finally, the dense linear system is transformed into an overdetermined system under the compressive sensing framework, and the least-squares method is used to reconstruct the current coefficients, from which the broadband Radar Cross Section (RCS) is calculated. The Root Mean Square Error (RMSE) is used to evaluate numerical accuracy. Three representative targets, namely a cylinder, a slotted cone, and an almond, are analyzed. Broadband RCS, numerical accuracy, total computation time, and single-frequency measurement-matrix memory consumption are compared with those obtained using MoM and CS-HBFM to validate the proposed framework.  Results and Discussions  Three numerical examples, including a perfect electric conductor cylinder, a slotted cone, and an almond, are used to validate the proposed CS-ACA-MMI framework. The ACA-extracted row indices are concentrated near geometric boundaries and structural junctions, demonstrating the physical validity of the deterministic sampling strategy (Fig. 2). Parametric studies show that appropriate ACA thresholds and four sampling frequencies provide the best balance between computational efficiency and numerical accuracy (Figs. 35). The broadband RCS predicted by the proposed framework agrees closely with the MoM results over the entire frequency band (Figs. 68), and the RMSE remains low, demonstrating high numerical accuracy. Compared with CS-HBFM, the proposed framework reduces the total computation time by 93.4% for the cylinder, 96.7% for the slotted cone, and 81.0% for the almond (Table 2). These improvements result from deterministic index reuse, MMI, and dual ACA acceleration, which substantially reduce the computational cost of broadband frequency-sweeping analysis.  Conclusions  A CS-ACA-MMI framework is proposed by integrating ACA with MMI for efficient broadband electromagnetic scattering analysis. The proposed framework eliminates repeated matrix construction and equation solving during frequency sweeping while overcoming the nondeterministic sampling strategy and the high computational and memory costs of conventional CS-HBFM. Dominant row indices extracted by ACA at the highest frequency provide a deterministic measurement-matrix construction strategy and a stable physical basis for broadband interpolation. By shifting the interpolation target from full impedance matrices to low-dimensional measurement matrices, the computational complexity and redundant matrix construction are substantially reduced. A second ACA decomposition further accelerates sensing-matrix construction by converting large-scale matrix-vector products into low-dimensional matrix multiplications. Numerical results demonstrate that the proposed framework achieves numerical accuracy comparable to that of MoM while reducing total computation time by more than 81% and decreasing single-frequency measurement-matrix memory consumption by up to 65%. Because only the measurement matrices at four sampling frequencies need to be stored, the overall memory requirement is further reduced.
Space-time-coding metasurface Enabled Integrated Design of Radar Communication and Electromagnetic Stealth
ZHANG Ming, WANG Zhe, WANG Boya, YANG Lin, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260536
Abstract:
  Objective  To address the complexity and limited modulation of existing reconfigurable metasurfaces integrating radiation and stealth, this study proposes a single-layer space-time-coding metasurface that dynamically switches between beam scanning and RCS reduction. By periodically modulating meta-atom states in the time domain, the phase responses of incident and harmonic waves are differentiated, enabling dynamic integration of radiation and stealth without complex feeding networks or multilayer structures. The design achieves precise beam scanning and excellent RCS reduction, providing a simple, low-cost solution for integrating radar communication and electromagnetic stealth.  Methods  This metasurface adopts a metal-dielectric-metal structure, with each meta-atom integrating a PIN diode. Electromagnetic simulations are conducted using CST Microwave Studio. At the optimal operating frequency, the meta-atom exhibits a 1-bit tunable reflection phase and a near-lossless co-polarized reflection amplitude. By feeding space-time-coding sequences into the metasurface model, dynamic control of beam scanning and RCS reduction is achieved. A prototype is fabricated using PCB technology, and its performance is validated through echo scattering measurements conducted with a vector network analyzer and a frequency offset option.  Results and Discussions  Simulations and measurements confirm that the proposed space-time-coding metasurface enables precise beam steering and excellent RCS reduction. At 10.0 GHz, the meta-atoms provide a 180° reflection phase difference and near-lossless amplitude. In radiation mode, harmonic beams are steered to angles of +2°, +14°, +32°, +46°, –15°, –28°, and –44°, with side-lobe levels for the ±3rd to fundamental harmonics suppressed to –10.59 dB. In scattering mode, the RCS of the fundamental echo is reduced by 11.56 dB compared to a copper plate. Additionally, a –10 dB RCS reduction is achieved from 9.9 to 10.1 GHz, peaking at 14.83 dB at 9.95 GHz. Experimental results validate the design’s integrated radiation-scattering capability and dynamic control flexibility.  Conclusions  This study presents a design method for a space-time-coding metasurface that enables dynamically integrated manipulation of radiation and scattering. By employing meta-atoms integrated with PIN diodes, the proposed method periodically modulates the operating states of the meta-atoms on a single-layer metasurface to achieve differentiated distributions of the wavefront phases of various harmonics. The proposed scheme significantly reduces design complexity and manufacturing costs. The successful realization of beam scanning and RCS reduction highlights the great potential of this technology in the integrated design of radar communication and electromagnetic stealth.
A Reinforcement Learning Driven Power Allocation Algorithm for Collocated MIMO Radar
HUANG Jieyu, XIE Junwei, ZHANG Haowei, FENG Weike, HAN Weihang
Available online  , doi: 10.11999/JEIT260695
Abstract:
  Objective  Traditional optimization-based power allocation algorithms for collocated MIMO radar have two fundamental limitations. First, they optimize tracking performance only for the next time step and therefore lack a full-time-horizon view of the power allocation process. This myopic strategy cannot achieve optimal multi-target tracking accuracy over extended periods, particularly when target trajectories vary substantially. Second, these algorithms rely on iterative nonlinear constrained optimization, resulting in high computational complexity. Therefore, they cannot satisfy the real-time requirements of dynamic battlefield environments where target states change rapidly. To address these limitations, this paper proposes a Reinforcement Learning (RL)-driven power allocation algorithm. Unlike conventional methods, the proposed approach formulates the power allocation problem as a Markov Decision Process (MDP) that maximizes long-term cumulative tracking accuracy. The algorithm adaptively allocates limited transmit power among multiple beams according to the current system state, balancing immediate tracking performance with long-term cumulative tracking accuracy.  Methods  The Posterior Cramér-Rao Lower Bound (PCRLB) is employed to quantify the theoretical lower bound of the tracking error for each target. The state space is constructed by combining the motion states (position and velocity) of all targets with the normalized PCRLB from the previous allocation step. The action space consists of discrete transmit power levels for each beam, subject to the total power budget and individual beam power constraints. All feasible power allocation vectors are enumerated and encoded to reduce the action-space dimensionality. The reward function is defined as the negative weighted sum of the normalized PCRLB, encouraging the agent to minimize tracking errors. The power allocation process is formulated as an MDP and solved using the Dueling Double Deep Q-Network (D3QN) algorithm. The D3QN framework incorporates three major enhancements: (1) a double-network architecture comprising an online Q-network and a target Q-network to improve training stability; (2) a dueling architecture that decomposes the Q-value into a state-value function and an action-advantage function to improve action discrimination; and (3) off-policy learning with experience replay to improve the use of historical trajectories. An ε-greedy strategy is adopted for exploration, with ε gradually decreasing during training. After offline training, the learned network directly generates real-time transmit power allocation decisions from the current system state without iterative optimization.  Results and Discussions  Simulations are conducted using three targets following the Constant Velocity (CV) model. Fixed power allocation yields the lowest tracking accuracy because of inefficient resource utilization. The traditional optimization method, which minimizes the instantaneous tracking error, achieves moderate tracking performance but remains myopic. When the discount factor \begin{document}$ \gamma =0 $\end{document}, the D3QN algorithm achieves performance comparable to that of the traditional optimization method because both optimize only immediate rewards. In contrast, when \begin{document}$ \gamma =0.99 $\end{document}, the D3QN algorithm significantly improves full-time-horizon tracking accuracy. The resulting power allocation strategy allocates more transmit power to distant, low-Signal-to-Noise Ratio (SNR) targets at earlier stages while reducing redundant power assigned to nearby high-SNR targets. The training curves show that \begin{document}$ \gamma =0.99 $\end{document} achieves a higher steady-state cumulative reward, although convergence exhibits greater oscillation because of the increased difficulty of estimating long-term returns. Furthermore, the trained D3QN network generates transmit power allocation decisions almost instantaneously, whereas the traditional optimization method must solve a constrained optimization problem at every time step, providing a substantial real-time computational advantage.  Conclusions  This paper proposes an RL-driven power allocation algorithm for collocated MIMO radar multi-target tracking that overcomes the myopic behavior and high computational complexity of conventional optimization methods. The proposed algorithm constructs the state space and reward function using the PCRLB, models the power allocation process as an MDP, and solves it using the D3QN algorithm. Simulation results demonstrate that, with an appropriate discount factor (\begin{document}$ \gamma =0.99 $\end{document}), the proposed approach significantly improves full-time-horizon tracking accuracy. This improvement results from the agent’s ability to learn a long-term optimal policy that proactively allocates transmit power to future distant, low-SNR targets. Furthermore, the trained network enables real-time decision-making through direct forward propagation, substantially reducing computational latency compared with iterative optimization. This work provides a new approach for intelligent radar resource management in complex battlefield environments.
Hierarchical Attention Mechanism-based Path Planning for Multi-UAV Inspection
FEI Bowen, XING Wenjie, LIU Daqian
Available online  , doi: 10.11999/JEIT260192
Abstract:
  Objective  In modern power inspection, the use of multiple Unmanned Aerial Vehicles (UAVs) for cooperative inspection is an efficient but challenging task. Existing multi-UAV path planning methods often have limited cooperative scheduling capability. They also fail to accurately capture the topological relationships among heterogeneous nodes, especially those between device nodes and charging stations under strict energy constraints. To address these limitations, this paper proposes Hierarchical Attention mechanism-based Path Planning for multi-UAV Inspection (HAPPI). The objective is to minimize the total flight distance of the UAV fleet while ensuring that all device nodes are inspected and all UAVs safely return to the base station under energy and visit-count constraints.  Methods  The multi-UAV power inspection problem is first formulated as a combinatorial optimization problem with energy constraints. It is then modeled within a Markov Decision Process (MDP) framework. To solve this problem, HAPPI adopts an encoder-decoder architecture with a customized hierarchical attention mechanism. The encoder uses a multi-level attention design to model three types of node relationships. Self-attention among device nodes is used to learn spatial proximity and visit-order preferences. Cross-attention between device nodes and charging stations is used to model energy supply-demand relationships. Self-attention among charging stations is used to explicitly capture the topological structure of the charging-station network. This hierarchical design enables the model to distinguish functional differences and dependencies among heterogeneous nodes. The decoder integrates the global graph embedding, the embedding of the last visited node, and the current remaining energy of the UAV to generate a context vector. A single-head attention mechanism is then used to compute compatibility scores for all candidate nodes. A masking strategy excludes infeasible nodes, including visited nodes, unreachable nodes, nodes that would prevent the UAV from reaching a charging station, and premature returns to the base station. The final node is selected from a probability distribution generated by softmax, which supports both greedy and sampling decoding strategies. The policy network is trained using reinforcement learning, and a baseline network is used to stabilize training. Policy-gradient optimization is used to minimize the expected total path length (Fig. 2).  Results and Discussions  Extensive simulations are conducted on three problem scales: T20C2, with 20 device nodes and 2 charging stations; T60C6, with 60 device nodes and 6 charging stations; and T100C10, with 100 device nodes and 10 charging stations. The training results show that HAPPI achieves faster convergence and a lower final cost than the baseline Attention Model (AM) and Heterogeneous Attention-based Deep Reinforcement Learning (HADRL) methods (Fig. 4). In the comprehensive performance comparison, HAPPI with sampling obtains the shortest total path lengths on T60C6 and T100C10, with values of 6.21 and 8.41, respectively. It outperforms five classical metaheuristic algorithms and two deep reinforcement learning baselines on these two larger-scale scenarios (Table 1). On T20C2, HADRL with sampling achieves the shortest path length, whereas HAPPI remains highly competitive. Overall, HAPPI reduces the total path length by approximately 12% on average compared with the baseline methods. The visualization results show that HAPPI generates routes with fewer route crossovers and a more balanced workload among UAVs, improving safety and efficiency (Fig. 6 and Fig. 7). The single-UAV path length distribution further confirms the superior load-balancing capability of HAPPI across all problem scales (Fig. 8).  Conclusions  This paper presents HAPPI, a hierarchical attention mechanism-based deep reinforcement learning method for cooperative path planning in multi-UAV power inspection scenarios with multiple charging stations. By explicitly modeling spatial relationships among device nodes, energy dependencies between device nodes and charging stations, and the internal topology of the charging-station network, HAPPI improves information aggregation and constraint satisfaction. Experimental results across different problem scales show that HAPPI achieves higher planning quality, greater computational efficiency, and stronger generalization than heuristic and learning-based comparison methods. Future work will extend this framework to multi-objective optimization that considers time, risk, and energy trade-offs, and will further validate the method using real-world inspection data.
Co-Frequency Interference Analysis and Dynamic Simulation Validation of Satellite-Direct-to-Device Systems Against Terrestrial IMT Networks in Cross-Border Scenarios
LIU Quan, ZHAO Weisong, XIAO Na, SONG Yanjun, ZHOU Meng, ZHANG Zhili, WANG Jinhai, WANG Lichong
Available online  , doi: 10.11999/JEIT260263
Abstract:
  Objective   Satellite-Direct-to-Device (SD2D) systems that reuse terrestrial IMT spectrum may generate harmful downlink interference to incumbent IMT networks in neighboring administrations, particularly in cross-border deployments where SD2D downlinks overlap the receive bands of both IMT user equipment (UE) and IMT Base Stations (BSs). A practical coexistence methodology is therefore required to (i) translate IMT receiver protection criteria into explicit Power Flux Density (PFD) and Equivalent Power Flux Density (EPFD) constraints and (ii) validate these constraints using a dynamic simulation framework so that they can be converted into enforceable geographic coordination measures, such as minimum isolation distances. This study focuses on the dominant interference path, namely SD2D downlink interference to IMT receivers, and establishes a traceable workflow from deterministic protection limits to dynamic simulation validation and the corresponding minimum isolation distances.  Methods  A cross-border scenario is modeled in which Country A deploys an SD2D system and Country B operates a terrestrial IMT network. Two representative downlink frequencies, 1995 MHz and 2 190 MHz, are evaluated for two representative Starlink configurations, Starlink-1 and Starlink-2. The IMT network is modeled using ITU-R- and 3GPP-compliant parameters, with an I/N protection threshold of –6 dB and a target percentile κ (baseline κ=99.5%) for both IMT UEs and BSs. Satellite transmit antennas follow the ITU-R S.1528 reference pattern, IMT BS receive antennas follow the ITU-R F.1336 sector pattern, and IMT UEs are modeled with omnidirectional antennas. A back-lobe blockage model is incorporated into both satellite and BS antenna patterns to account for rear-side shielding. Signal propagation follows the ITU-R P.619 model, using free-space path loss as the conservative baseline, while an optional clutter-loss term is incorporated through a clutter-occurrence probability. Deterministic protection limits are derived by calculating the maximum permissible aggregate PFD for IMT UE protection and the maximum permissible aggregate EPFD for IMT BS protection. A dynamic simulation framework then validates these limits and searches for the required minimum isolation distances (Fig. 4). Co-channel beam isolation angles are optimized using the C/I Complementary Cumulative Distribution Function (CCDF), and a segmented search algorithm determines the minimum UE- and BS-side isolation distances together with the corresponding κ-percentile PFD/EPFD statistics.  Results and Discussions  The deterministic analysis yields a maximum permissible aggregate PFD of –102.72 dBW/m2/MHz at 2 190 MHz for IMT UEs and a maximum permissible aggregate EPFD of –129.53 dBW/m2/MHz at 1995 MHz for IMT BSs (Fig. 3). For Starlink-1, the C/I design criterion yields a minimum co-channel beam isolation angle pair of (12°, 12°) (Fig. 5). Dynamic simulation shows that, under the representative baseline configuration with an I/N threshold of –6 dB and κ=99.5%, the minimum isolation distances are 195 km for UE protection and 290 km for BS protection (Fig. 6, Fig. 7, and Table 4). The resulting coordination isolation distance is therefore 290 km, and the simulated κ-percentile PFD and EPFD agree with the deterministic protection limits, with a residual margin below 0.5 dB. For Starlink-2, the optimized co-channel beam isolation angles increase to (15°, 15°), and the corresponding minimum isolation distances increase to 272 km for UEs and 420 km for BSs under the same baseline configuration (Table 5). These baseline distances should be interpreted as representative values for the specified simulation configuration rather than unique, strictly converged results. Stability verification shows that, under different sampling intervals, simulation durations, and random seeds, the UE- and BS-side minimum isolation distances remain within 195~210 km and 290~300 km, respectively, for Starlink-1, and within 266~290 km and 370~420 km, respectively, for Starlink-2 (Table 7). Sensitivity analysis for Starlink-1 further indicates that the required minimum isolation distance is governed by the upper tail of the aggregate I/N distribution (Table 6). Increasing κ from 99.5% to 100% increases the UE- and BS-side minimum isolation distances from 195/290 km to 304/560 km. Clutter attenuation substantially reduces the UE-side minimum isolation distance, decreasing it to 173 km when the clutter-occurrence probability is 0.5, while producing little change in BS protection. Polarization reuse increases the UE- and BS-side minimum isolation distances to 222 km and 360 km, respectively, whereas increasing the number of co-channel beams to 16 increases the BS-side minimum isolation distance to 330 km. The minimum service elevation angle and the link establishment strategy are identified as the dominant operational factors. Changing the minimum service elevation angle from 10° to 35° changes the required UE- and BS-side minimum isolation distances from 340/460 km to 101/150 km, whereas replacing the Sat-MaxElevation strategy with the UE-MaxElevation strategy reduces them to 80/180 km.  Conclusions   The proposed workflow converts IMT receiver protection criteria into deterministic protection limits expressed as PFD and EPFD constraints and validates them using a dynamic simulation framework. Under an I/N threshold of –6 dB and κ=99.5%, the baseline and stability analyses jointly indicate representative UE- and BS-side minimum isolation-distance ranges of 195~210 km and 290~300 km for Starlink-1 and 266~290 km and 370~420 km for Starlink-2, rather than unique, strictly converged values. Sensitivity analysis further shows that κ only changes the statistical criterion used to extract tail events from the sample set, whereas clutter attenuation primarily benefits IMT UEs. In contrast, the minimum service elevation angle, polarization reuse, the number of co-channel beams, and the link establishment strategy reshape the worst-case interference geometry and can produce substantial, and sometimes non-monotonic, changes in the required minimum isolation distances. The proposed framework establishes a traceable link between IMT receiver protection criteria and enforceable border coordination measures.
Recent Advances in Remote Sensing Image-Text Retrieval Driven by Vision-Language Foundation Models
WU Hui, ZHAO Yan, ZHANG Peirong, HOU Yingyan, QI Xiyu, WANG Lei
Available online  , doi: 10.11999/JEIT260189
Abstract:
  Significance  Remote Sensing Image-Text Retrieval (RS-TIR) connects large-scale Earth observation imagery with natural-language queries and has become an important interface for geospatial intelligence systems. Compared with conventional content-based retrieval, RS-TIR allows users to search for scenes, objects, spatial layouts, and functional regions through semantic descriptions rather than handcrafted visual cues. This capability is increasingly needed in natural resource monitoring, urban governance, disaster response, environmental assessment, and on-demand retrieval from rapidly growing satellite archives. However, RS-TIR remains challenging. Remote sensing imagery is captured from nadir or near-nadir perspectives, shows strong rotation invariance, and contains extreme scale variation, ranging from tiny vehicles to large airports. It also requires domain-specific semantic descriptions, such as land-use attributes, spatial distributions, and geoscientific relations. Meanwhile, high-quality image-text annotations remain limited relative to the scale of remote sensing data. These properties widen the cross-modal semantic gap between images and language and limit the generalization ability of traditional cross-modal retrieval methods. Against this background, this review examines how Vision-Language Foundation Models (VLMs) reshape RS-ITR through large-scale contrastive pre-training, stronger transferable representations, and more flexible multimodal interaction mechanisms. It also explains why remote sensing adaptation is needed and why a focused synthesis of architectures, datasets, alignment mechanisms, and future directions is timely for this field.  Progress   The technical development of RS-ITR is reviewed from three complementary perspectives. First, this review summarizes the domain-specific challenges that shape the task, including visually isotropic topology with extreme scale variation, professional and fine-grained textual semantics, and the compounded cross-modal semantic gap between overhead imagery and natural-language descriptions (Fig. 3). The overall survey structure is then presented to show the logical progression from task formulation to future challenges (Fig. 1). From a methodological perspective, RS-ITR has evolved from handcrafted visual descriptors and shallow semantic mapping to deep representation learning, and then to VLM-driven paradigms with stronger generalization and zero-shot transfer capability (Fig. 4, Table 2). Early methods rely on color, texture, shape, and hash-based retrieval. However, they struggle to model high-level geospatial semantics and complex scene composition. Deep learning methods improve retrieval by learning joint embedding spaces, adopting dual-encoder or interaction-based architectures, and using multi-scale feature fusion and region-aware matching. These methods improve semantic consistency, but they still depend heavily on labeled data and often show limited robustness in open or cross-sensor scenarios. Second, this review summarizes the benchmark ecosystem used to evaluate these methods. Representative datasets range from small-scale test sets, such as Sydney-Caption and UCM-Caption, to mainstream benchmarks, such as RSICD and RSITMD, and recent large-scale training resources, such as RS5M and SkyScript (Table 1). These datasets show a clear transition from small manually annotated corpora to web-scale or automatically generated image-text pairs. This transition supports domain pre-training and large model adaptation. Third, this review analyzes the core VLM techniques that now drive progress in RS-ITR. The model spectrum and representative architecture families are systematically summarized, including contrastive dual-encoder models, multimodal interaction models, and remote sensing foundation models integrated with large language models (Fig. 5, Fig. 6, Table 3). Domain adaptation routes are further grouped into continued remote sensing pre-training, parameter-efficient transfer learning, adapter-based tuning, prompt learning, and instruction tuning. At the semantic alignment level, this review focuses on contrastive joint embedding, fine-grained multi-scale alignment, and the use of remote sensing priors, such as spatial topology and geolocation. Performance comparisons on RSICD and RSITMD show that remote sensing VLMs, especially RemoteCLIP, GeoRSCLIP, iEBAKER, and LRSCLIP, yield consistent gains in mean Recall (mR) and overall retrieval robustness (Table 4). In parallel, this review tracks the extension of retrieval capability into unified multi-task remote sensing models, in which retrieval, grounding, segmentation, and reasoning begin to share a common multimodal representation space.  Conclusions  Several conclusions are drawn from the comparative analysis. First, VLMs establish a dominant paradigm for RS-ITR because they narrow the cross-modal semantic gap and improve transferability across datasets and scenes. Second, no single architecture is universally optimal. Dual-encoder models remain attractive for large-scale retrieval because of their efficiency, whereas interaction-based or instruction-enhanced models provide finer semantic alignment at a higher computational cost. Third, domain adaptation is indispensable. Continued pre-training on remote sensing image-text corpora, parameter-efficient tuning, and prompt-based adaptation consistently outperform direct reuse of internet-trained VLMs. This finding indicates that remote sensing imagery differs too strongly from natural-image distributions for generic pre-training alone to be sufficient. Fourth, the most effective recent methods do not improve performance through scale alone. They also exploit remote sensing-specific information, including multi-scale structures, foreground objects, explicit keyword reasoning, and spatial priors. Finally, this review shows that the field is shifting from isolated retrieval models toward more general geospatial multimodal systems. Retrieval is no longer treated only as a matching task. It is also becoming a key capability that supports question answering, instruction following, knowledge augmentation, and coordinated reasoning in remote sensing applications.  Prospects   Future research is expected to advance in four closely related directions. The first direction is the unified representation of multi-source heterogeneous data, especially the integration of optical imagery with Synthetic Aperture Radar (SAR), hyperspectral data, thermal infrared observations, and multi-temporal acquisitions. The second direction is knowledge-enhanced retrieval, in which geospatial priors, land-use rules, remote sensing terminology, and external knowledge bases are incorporated into multimodal alignment and retrieval-augmented reasoning. The third direction is lifelong and open-world learning. Real deployment requires models to remain reliable under seasonal variation, sensor updates, regional domain shifts, cloud contamination, and newly emerging categories, while avoiding catastrophic forgetting. The fourth direction is efficiency and deployability. Practical remote sensing systems often operate under tight computational budgets. Therefore, lightweight tuning, sparse computation, token reduction, model compression, and on-orbit and edge inference will become increasingly important. Interactive and explainable retrieval is also likely to gain importance. It allows analysts to refine queries through dialogue and inspect the image regions or semantic cues that support retrieval decisions. Overall, continued progress in data construction, domain adaptation, semantic alignment, and efficient multimodal modeling is expected to make RS-ITR a more robust infrastructure capability for Earth observation applications.
Deep Side-Channel Attack Method Integrating Convolutional Block Attention Mechanism and Triplet Metric Learning
XU Yang, LI Kaibin, HE Xingxing
Available online  , doi: 10.11999/JEIT260140
Abstract:
  Objective  Side-Channel Attack (SCA) is one of the primary threats to the physical security of cryptographic chips, and deep learning methods for secret key recovery have attracted considerable attention in the field of SCA. However, existing deep learning-based side-channel attack methods have limited capability to focus on critical leakage intervals during feature extraction, particularly for long traces with high-dimensional noise. Therefore, irrelevant background noise interferes with feature extraction, leading to reduced attack efficiency and slower convergence of Guessing Entropy (GE). To address these limitations, a deep side-channel attack method integrating the Convolutional Block Attention Module (CBAM) and triplet loss is proposed to improve the extraction of weak leakage features under complex noise conditions and enhance secret key recovery efficiency.  Methods  CBAM is embedded into a Convolutional Neural Network (CNN) to construct an adaptive feature extraction network. CBAM consists of a Channel Attention Module (CAM) and a Spatial Attention Module (SAM). CAM adaptively recalibrates feature-channel weights to emphasize leakage-related features with a high Signal-to-Noise Ratio (SNR), whereas SAM identifies Points Of Interest (POI) in the temporal domain and suppresses background noise outside the leakage intervals. After attention-based feature refinement, triplet loss is adopted as the optimization objective to optimize the distribution of embedding features, encouraging compact intra-class clusters while maximizing inter-class separation in the embedding space. Finally, a multivariate Gaussian template attack is performed using the optimized embedding features to recover the secret key. The overall framework is illustrated in (Fig. 2).  Results and Discussions  The proposed method is evaluated on two public benchmark datasets, ASCAD and AES_HD, using GE and the minimum number of attack traces required for GE to converge to 1 (\begin{document}$ {T}_{{\mathrm{GE0}}} $\end{document}) as evaluation metrics. On the ASCAD dataset, the proposed method requires only 144 attack traces in the ASCAD_f (HW) scenario, representing a 51.0% reduction compared with the conventional CNN model. In the ASCAD_f (ID) scenario, only 61 attack traces are required, corresponding to a 68.0% reduction. In the ASCAD_r dataset, GE converges with 176 attack traces under the HW leakage model and 137 attack traces under the ID leakage model, outperforming representative methods, including RL-SCA and Metric Learning (Table 2 and Fig. 3). On the low-SNR AES_HD dataset, the proposed method requires 1,219 attack traces, outperforming MHA and NLS while maintaining smooth and stable GE convergence (Table 2 and Fig. 4). Furthermore, desynchronization experiments demonstrate that the proposed method maintains accurate localization of effective leakage points under severe desynchronization noise, indicating strong robustness to time-domain jitter (Table 3). Ablation studies further verify the synergistic effect of the proposed architecture and confirm the effectiveness of its core components (Table 4).  Conclusions  A deep side-channel attack method integrating CBAM and triplet-loss-based metric learning is proposed. The CBAM module enables the network to adaptively focus on leakage-related features, improving feature extraction over conventional CNN-based methods. Triplet loss enhances the discriminability of embedding features, thereby improving template matching accuracy. Experimental results on the ASCAD and AES_HD datasets demonstrate that the proposed method substantially reduces the number of attack traces required for successful secret key recovery and accelerates GE convergence. The proposed method consistently outperforms existing mainstream approaches under fixed-key, random-key, and low-SNR conditions. Future work will focus on improving robustness under more severe desynchronization conditions and enhancing generalization in small-sample scenarios.
A WiFi Multi-link Collaborative Human Tracking Method for Smart Home
PAN Houcheng, CAI Yushuang, YAO Junmei, ZHANG Tingting
Available online  , doi: 10.11999/JEIT260267
Abstract:
  Objective  WiFi-based indoor passive human tracking has attracted increasing attention for smart home applications because Internet of Things (IoT) devices are widely interconnected through existing WiFi infrastructures. Existing approaches estimate the Angle of Arrival (AoA) or Time of Flight (ToF) of the target-reflected path to achieve tracking. However, these methods are fundamentally limited by the difficulty of separating weak target-reflected signals from dense multipath propagation when commercial WiFi devices provide only a small antenna array and limited bandwidth under existing communication protocols. Therefore, Doppler Frequency Shift (DFS)-based approaches that exploit multiple links have become more practical. Dead Reckoning (DR) is widely adopted in these methods, but dynamic link selection for signal fusion remains challenging. Furthermore, the performance of DR-based methods degrades substantially when device-location errors are present. Existing methods also generally employ a single-transmitter–multi-receiver architecture, in which Access Point (AP) broadcasts downlink WiFi signals to STAtions (STA). It complicates data aggregation and limits the use of multiple receive antennas of AP's. To address the challenges of multi-link signal fusion and device-location uncertainty, a Particle Filter (PF)-based method is proposed. Furthermore, a multi-transmitter–single-receiver architecture is adopted, in which multiple STAs transmit uplink packets to a single AP. This architecture simplifies data aggregation while exploiting the AP’s multi-antenna capability.  Methods  In the IEEE 802.11 standard, the wireless channel is estimated at the receiver using pilot signals embedded in WiFi packets, and Channel State Information (CSI) is continuously obtained. Human motion perturbs multipath propagation and produces time-varying changes in CSI. Therefore, CSI serves as the primary sensing signal because it implicitly captures target motion. In practice, raw CSI is affected by amplitude and phase impairments. Accordingly, signal preprocessing is first performed before Doppler extraction. Subsequently, multi-link DFS measurements are fused using PF, in which the target state is sequentially propagated and updated according to likelihoods derived from DFS measurements. This probabilistic framework naturally enables dynamic multi-link fusion because unreliable links receive lower weights rather than being deterministically discarded. Moreover, the effects of device-location errors are incorporated into the measurement-noise model, improving robustness to device-location uncertainty. When prior knowledge of device locations is unavailable, device self-localization is first performed by jointly estimating the Line-of-Sight (LoS) AoA and ToF between the AP and each STA. Preliminary device self-localization experiments (Fig. 5) demonstrate a median STA localization error of approximately 0.56 m (Table 1), providing reliable initialization for the subsequent tracking algorithm.  Results and Discussions  A prototype system is implemented using a multi-transmitter–single-receiver architecture with commercial Intel AX200/AX201 WiFi cards. CSI is collected over a 20 MHz channel centered at 5.24 GHz using PicoScenes, with IEEE 802.11ac packets transmitted at 100 or 200 Hz. Each CSI sample contains measurements from 57 subcarriers. Meanwhile, Fine Time Measurement (FTM) is performed on Channel 11 at 2.4 GHz using iw and hostapd. In the prototype system, each STA transmitter is equipped with a single antenna, while the AP receiver employs two antennas. Experiments are conducted using one AP and three or four STAs (Fig. 7). WiTraj and PITrack, both based on DR, are used as baseline methods. When device self-localization is required, the proposed method achieves a median tracking error of 0.47 m, compared with 1.78 m for WiTraj and 1.77 m for PITrack (Fig. 9), representing an accuracy improvement of approximately 70% over conventional DR-based methods. Error-injection experiments with random device-location perturbations show that, unlike conventional DR-based methods, which are highly sensitive to device-location errors, the proposed method remains robust to device-location uncertainty (Fig. 10). Additionally, computational complexity analysis shows that, although execution time increases with the particle count (Table 3), tracking performance reaches a stable level beyond a moderate particle count (Fig. 11). Therefore, real-time operation can be achieved by selecting an appropriate particle count. Finally, experiments in complex environments demonstrate that the proposed method consistently achieves higher tracking accuracy than the baseline methods (Fig. 12).  Conclusions  A PF-based method is proposed for multi-link collaborative passive human tracking using commercial WiFi devices. By probabilistically fusing measurements from multiple links, robust tracking is achieved even when device-location errors are present. When device locations are unavailable, device self-localization is achieved by combining CSI-based estimation with FTM measurements, providing the geometric information required for tracking. A prototype system based on a multi-transmitter–single-receiver architecture is developed to simplify data aggregation while exploiting the AP’s multi-antenna capability. Experimental results demonstrate that the proposed method achieves sub-0.5 m median tracking error when device-location errors are present, representing an improvement of approximately 70% over conventional DR-based methods. The proposed method also exhibits strong robustness to device-location uncertainty and supports real-time implementation when an appropriate particle count is selected. Future work will extend the framework to multi-person tracking and evaluate its performance under more realistic deployment conditions.
Decoupled Learning for Long-tailed Oracle Bone Character Recognition Based on Adaptive Difficulty Sampling
SUN Junwei, GUAN Suyan, CHEN Xinyu, WANG Kun, CAI Yuanqiang
Available online  , doi: 10.11999/JEIT260327
Abstract:
  Objective  Oracle Bone Character (OBC) recognition is challenged by an extreme long-tailed distribution and substantial intra-class variation. Conventional deep learning methods are often dominated by head classes, whereas existing approaches tend to overfit tail classes or fail to account for differences in learning difficulty across classes. To address these limitations, a two-stage decoupled learning framework is proposed to improve the recognition of tail and difficult classes while preserving the discriminative capability of head classes.  Methods  The proposed framework decouples feature representation learning from classifier optimization. In the first stage, the backbone network is trained using a mixed data augmentation strategy that combines CutMix and RandAugment with Label-Distribution-Aware Margin (LDAM) loss to learn robust feature representations and alleviate the effect of intra-class variation. In the second stage, the backbone network is frozen, and only the classifier is optimized. An adaptive difficulty sampling strategy is proposed to dynamically assign sampling weights according to historical and current class-level training difficulty. The classifier is further optimized using a Class-Balanced LDAM (CBL) loss, which combines class-balanced weighting with LDAM to refine decision boundaries for long-tailed classification.  Results and Discussions  Experiments on the highly imbalanced OBC306 dataset demonstrate that the proposed method achieves an overall accuracy of 94.34% and an average class accuracy of 89.89%. Compared with the Inception-v4 baseline, the proposed method improves the average class accuracy by 19.61%. Comparisons with representative long-tailed OBC recognition methods further demonstrate superior overall performance. Comprehensive ablation studies verify the effectiveness of the mixed data augmentation strategy and the adaptive difficulty sampling strategy in improving the recognition of rare and difficult characters. Parameter sensitivity analysis and qualitative error analysis further confirm the robustness and effectiveness of the proposed framework.  Conclusions  The proposed two-stage decoupled learning framework effectively addresses long-tailed OBC recognition by balancing the learning priorities of head, tail, and difficult classes. The mixed data augmentation strategy improves feature robustness, whereas the adaptive difficulty sampling strategy and the Class-Balanced LDAM loss jointly optimize classifier learning and refine decision boundaries without degrading head-class recognition performance. The proposed framework provides an effective solution for the digital recognition of Oracle Bone Characters and offers technical support for low-resource ancient character recognition.
Construction of a DNA Strand Displacement Memristor and Its Filter Circuit Characteristics
WANG Yanfeng, CHEN Guanzhou, SUN Ce, SUN Junwei
Available online  , doi: 10.11999/JEIT260283
Abstract:
  Objective  Filter circuits are widely used in modern control and signal-processing systems for noise suppression and signal integrity enhancement. Conventional Resistor-Capacitor (RC) filters are widely applied, but their fixed parameters limit adaptability and miniaturization in emerging molecular and nanoscale computing platforms. To address these limitations, DNA Strand Displacement (DSD) technology is integrated with memristor theory to develop tunable multistable molecular filter circuits. This study aims to design and validate first- and second-order low-pass filter circuits based on the dynamic response and state-dependent behavior of a DSD-based memristor. The proposed filters are designed to improve frequency selectivity, parameter adaptability, and system stability compared with traditional filter architectures. This approach is intended for molecular signal processing, integrated biocircuits, and adaptive filtering systems that require compact size and reconfigurability.  Methods  The method consists of four stages. First, core DSD reaction modules, including sine, cosine, integration, addition, and multiplication modules, are designed to construct a programmable multistable memristor model. Second, square-wave and sinusoidal input signals are generated through DSD reactions to evaluate the memristor response under different frequencies and amplitudes. Third, the memristor is embedded into low-pass filter structures to construct first- and second-order DSD-based memristor filter circuits. Fourth, simulations are performed using Visual DSD for molecular dynamics analysis and MATLAB for circuit-level analysis. Circuit performance is evaluated using transfer functions, Nyquist plots, Bode diagrams, and time-domain comparisons with classical RC filters. This combined simulation strategy verifies both molecular feasibility and circuit functionality.  Results and Discussions  The DSD-based memristor exhibits multistable behavior and converges to six stable equilibrium points under different initial conditions (Fig. 8). Its hysteresis characteristics further confirm the state-dependent memory behavior of the designed molecular memristor (Fig. 7). The first-order DSD-based memristor filter circuit provides stable attenuation for square-wave and sinusoidal input signals. Its output amplitudes are consistently higher than those of the traditional RC filter across the tested frequencies (Table 3). The second-order DSD-based memristor filter circuit further reduces signal delay and improves stability, especially under high-frequency inputs (Table 4). Frequency-response analyses show that the cutoff frequency can be dynamically tuned by adjusting DSD reaction rates and initial concentrations (Figs. 9 and 11). Time-domain simulations further confirm the filtering performance of the first- and second-order circuits (Figs. 10 and 12). Reliability analysis indicates that lower initial copy numbers increase stochastic molecular noise, whereas higher initial copy numbers make the output distribution closer to the deterministic response and improve the probability of successful filtering. These results verify the feasibility of DSD-memristor integration for adaptive molecular filtering.  Conclusions  A DSD-based memristor with multistable characteristics and its corresponding first- and second-order low-pass filter circuits are designed and validated. Compared with traditional RC architectures, the proposed filters show improved output stability, parameter tunability, and frequency adaptability. By combining DSD technology with memristor theory, this study provides a reconfigurable molecular-scale filtering framework for signal-processing applications. The results provide a basis for future work on adaptive molecular circuits, intelligent filtering, and nanoelectronic system design. Further studies should focus on experimental validation, real-time tuning strategies, sequence optimization, anti-interference design, signal amplification, and circuit integration.
Non-Orthogonal PSWFs Signal Detection Method Based on Adaptive Time-Space Feature Fusion
CHEN Wenhua, MAO Zhongyang, LU Faping, SUN Ye, GAO Yixuan
Available online  , doi: 10.11999/JEIT260024
Abstract:
  Objective   To address B5G/6G demands for high spectral efficiency and transmission reliability, non-orthogonal Prolate Spheroidal Wave Function (PSWF) modulation has drawn extensive research interest for its strong time-frequency energy concentration. However, severe mutual interference in PSWF signal multiplexing degrades conventional detection performance in complex channels. Limited by ideal channel assumptions or single-modal feature extraction, existing methods cannot fully exploit signal temporal-spatial information and lack adaptability to dynamic interference. This work proposes an intelligent detection architecture with adaptive temporal-spatial feature fusion (ATSFF) for accurate, robust non-orthogonal PSWF signal detection.  Method   A dual-path parallel framework extracts complementary features from time and transform domains. A Gated Recurrent Unit (GRU) extracts deep temporal features and captures long-range dependencies from 1D received signals. The other path converts 1D signals into 2D representations via the Gramian Angular Difference Field (GADF), and extracts hierarchical spatial features using ResNet50. An adaptive probability-weighted fusion mechanism dynamically adjusts feature contributions based on sub-network prediction uncertainty to generate robust detection outputs.  Results and Discussion   Simulations on a 32-class non-orthogonal PSWF dataset (Fig. 2) show the proposed ATSFF outperforms conventional coherent detection, cross-term suppression detection, Approximate Message Passing Interleave Division Multiple Access (AMP-IDMA) and Temporal Sparse Bayesian Learning Least Squares (TMSBL-LS) across the full Signal-to-Noise Ratio (SNR) range. t-SNE visualization (Fig. 4) confirms better inter-class separability and intra-class compactness of fused features. At a bit error rate of 4×10–5, it achieves 0.2 dB gain (Fig. 6) over coherent detection. ATSFF has higher network overhead and weaker real-time performance than linear detection, yet it supports GPU batch inference with fixed single-sample computation cost, suiting accuracy-sensitive communication scenarios.  Conclusions   Targeting interference issues in non-orthogonal PSWF signal detection, this paper presents an adaptive temporal-spatial feature fusion detection method. It realizes dual-modal feature extraction via GRU and ResNet50, and adopts a prediction-uncertainty-based adaptive weighted fusion mechanism to integrate dual-path strengths, significantly improving detection accuracy and robustness. Simulations validate its superior performance and stable channel adaptability, providing an efficient solution and design reference for intelligent detection of non-orthogonal PSWF signals.
Load Optimization of Inverter Air-Conditioning Clusters Driven by Constraint Surface Projection and Spatial-Fitness Synergy
ZHENG Bowen, PAN Mingming, WANG Lei, LIU Chang, ZHENG Qingrong, TANG Zhuofan, ZHAO Jianli
Available online  , doi: 10.11999/JEIT260149
Abstract:
  Objective  Supply-demand imbalances in modern distribution networks are intensified by the increasing penetration of distributed renewable energy and frequent extreme high-temperature events. Large-scale Inverter Air-Conditioning (IAC) clusters can be aggregated as virtual energy storage resources for Demand Response (DR), providing an effective way to improve grid flexibility. However, existing dispatch strategies are often limited by the curse of dimensionality. Conventional penalty-function-based soft constraints also fail to strictly satisfy aggregate power equality constraints and may introduce steady-state errors. This paper develops an optimization framework in which grid-side power commands are accurately tracked while user thermal discomfort is reduced and fairness among heterogeneous users is maintained.  Methods  A multi-objective optimization framework based on an Equivalent Thermal Parameter (ETP) model is established to describe the thermodynamic states of heterogeneous buildings. To balance collective comfort and individual fairness, a composite fitness function is designed by integrating a weighted mean-squared error term, a fairness variance term, and a maximum violation suppression term. To eliminate the steady-state errors of traditional penalty-based methods, a Spatial-Fitness Adaptive Particle Swarm Optimization (SFA-PSO) algorithm is proposed. A geometric constraint surface projection mechanism maps particles strictly onto the power-conservation hyperplane, thereby satisfying the aggregate power equality constraint. In addition, the learning factors are dynamically adjusted through a Spatial-Fitness Adaptive (SFA) strategy. This strategy measures the mismatch between a particle’s fitness rank and spatial distance rank, which helps prevent premature convergence in high-dimensional search spaces.  Results and Discussions  Extensive continuous scheduling simulations are conducted in a complex dynamic environment. The environment includes multi-source thermal disturbances, a bidirectional communication packet loss rate of 1%, and Part Load Ratio (PLR) values of 20%, 50%, and 80%. First, ablation experiments confirm that constraint surface projection guarantees power tracking accuracy. Traditional penalty-based methods, such as Penalty Particle Swarm Optimization (Penalty-PSO), produce steady-state power deviations of approximately 10^-1 kW. By contrast, SFA-PSO limits aggregate power tracking errors to within 10^-9 kW (Fig. 3). The SFA strategy also prevents the premature convergence observed in Physical Particle Swarm Optimization (Phy-PSO). It enables continuous fitness reduction, especially in low-load scenarios with narrow feasible regions (Fig. 4). This improvement is attributed to the dynamic evolution of the learning factors. The cognitive factor remains high at the early stage to promote global exploration. It then decreases as the social factor increases, which strengthens local exploitation and improves convergence precision (Fig. 5). Second, continuous dynamic scheduling performance is evaluated through a 6-hour simulation during the peak load period from 12:00 to 18:00. The dispatch interval is 5 min, yielding 72 decision steps. Under tight peak-load constraints, Genetic Algorithm (GA) and Whale Optimization Algorithm (WOA) show severe power-limit violations because their population update rules do not cooperate well with the projection mechanism. By contrast, SFA-PSO maintains strict constraint satisfaction (Fig. 7). SFA-PSO remains at the lowest fitness level throughout the real-time evolution curves, indicating strong robustness against environmental thermal noise and uplink and downlink communication packet loss (Fig. 8). Quantitatively, compared with eight baseline algorithms, including Social Learning Particle Swarm Optimization (SLPSO), Competitive Swarm Optimizer (CSO), and Dynamic State Cluster-Based Particle Swarm Optimization (DSCPSO), SFA-PSO achieves the best overall performance. It obtains an average fitness of 904, a minimum fitness of 243, and the lowest standard deviation of 551 (Table 2). Finally, scalability analyses across cluster sizes from 100 to 1,000 nodes further validate the high-dimensional optimization capability of SFA-PSO. In all scale scenarios, SFA-PSO shows the strongest optimization capacity. It achieves rapid initial descent within the first 20 iterations and maintains continuous exploration in later stages (Fig. 9). Although the projection and SFA mechanisms increase computational time by 30% to 50% compared with basic Particle Swarm Optimization (PSO) (Fig. 6), the absolute optimization time remains stable at approximately 1.5 seconds even for a 1,000-node cluster (Fig. 9). This computational overhead is acceptable for minute-level control cycles and meets the real-time dispatch requirements of modern smart grids.  Conclusions  The proposed SFA-PSO algorithm effectively addresses the steady-state error of traditional soft-constraint methods in aggregate power control. By ensuring accurate tracking of dispatch commands and mitigating high-dimensional search traps, it provides a robust and scalable solution for flexible scheduling of large-scale IAC loads in smart grids. It also maintains a practical balance between grid-side regulation and user-side comfort. The method still has limitations. The constraint projection mechanism depends on the host algorithm, which restricts cross-algorithm generalization. High-precision tracking also increases computational cost. Future work will focus on adaptive constraint handling and lightweight algorithm design. Coordinated scheduling for heterogeneous loads, such as electric vehicles and energy storage, will also be investigated.
Resource Allocation Optimization in Dual-RIS Cooperative Rate-Splitting Multiple Access Networks
CHEN Yuang, WU Chang, PENG Mingyu, LU Hancheng
Available online  , doi: 10.11999/JEIT260171
Abstract:
  Objective  In Rate-Splitting Multiple Access (RSMA) systems, the achievable common-stream rate is limited by the user with the weakest channel quality. This constraint reduces scalability, robustness, and user fairness in dense 6G networks. Existing cooperative RSMA architectures partly reduce this bottleneck, but they remain constrained by fixed channel conditions and limited interference management. To address these issues, this paper proposes a dual Reconfigurable Intelligent Surface (RIS) cooperative RSMA system. Two cooperatively deployed RISs create additional controllable propagation paths through cascaded double reflection. The objective is to maximize the system sum rate by jointly optimizing Base Station (BS) Beamforming (BF), Rate Splitting (RS) strategies, and dual-RIS phase configurations, thereby improving spectral efficiency, robustness, and user fairness under users’ Quality of Service (QoS) constraints.  Methods  A tractable system model is developed for the dual-RIS cooperative RSMA system. The model captures cascaded multi-link channels, multi-node channel structures, and interference coupling. Based on this model, a joint optimization problem is formulated to maximize the system sum rate by optimizing BS BF, RS strategies, and the discrete phase shifts of both RISs. Because of strong variable coupling and non-convexity, a low-complexity Alternating Optimization (AO) algorithm is designed. The original problem is decomposed into three subproblems: BS-side RIS phase optimization, user-side RIS phase optimization, and BS BF optimization. Semidefinite Relaxation (SDR) and Successive Convex Approximation (SCA) are used to transform these subproblems into tractable convex forms, which are then solved iteratively with fast convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed dual-RIS cooperative RSMA system. The proposed AO algorithm converges within six iterations under different numbers of RIS reflecting elements, and it reaches 97.8% of the steady-state sum rate within three iterations when M = 170 (Fig. 3). Compared with SOPS and RPS, the proposed phase-configuration scheme obtains 10.6% and 31.8% sum-rate gains when M = 190, respectively (Fig. 4). The proposed RSMA scheme also outperforms NOMA and SDMA by 10.0% and 14.6%, respectively (Fig. 5). Under M = 160 and b = 4, dual-RIS cooperation provides an 11.9% sum-rate gain over the single-RIS scheme, and its performance is close to the CPS upper bound (Fig. 6). Balanced allocation of reflecting elements between the two RISs further improves the sum rate (Fig. 7). The proposed BF strategy also outperforms ZF and RBF, achieving 33.2% and 336.5% gains at a transmit power of 30 dBm, respectively (Fig. 8). Under different RIS cooperation modes, the proposed joint optimization scheme achieves the best overall performance (Fig. 9). These results show that dual-RIS cooperative RSMA improves common-stream decoding, interference suppression, robustness, and user fairness.  Conclusions  This paper investigates a dual-RIS cooperative RSMA communication system. The proposed architecture improves common-stream decoding while mitigating complex interference. To maximize the system sum rate, BS BF vectors, RS vectors, and the discrete phase matrices of two RISs are jointly optimized. A low-complexity AO algorithm based on SDR and SCA is developed to solve the strongly coupled non-convex problem. Simulation results show that the proposed dual-RIS cooperative RSMA scheme achieves clear sum-rate gains over advanced benchmark schemes. Compared with the single-RIS mode and SDMA, it obtains 11.9% and 14.6% rate gains, respectively, while improving system robustness and user fairness.
A Multi-Dimensional Scenario-Based Evaluation Method for Deep Learning Side-Channel Analysis Using a Multi-Attribute Decision Model
GU Zepeng, CHEN Lin, CAI Juesong, YAN Yingjian
Available online  , doi: 10.11999/JEIT260198
Abstract:
  Objective  Deep Learning Side-Channel Analysis (DL-SCA) has substantially improved the effectiveness of attacks against protected cryptographic implementations. However, the transition of DL-SCA models from research to practical deployment is limited by the lack of systematic, fair, and scenario-specific evaluation methods. Existing evaluations mainly rely on Guessing Entropy (GE) and Success Rate (SR), while overlooking practical factors such as resource overhead and environmental adaptability. Moreover, inconsistent hyperparameter optimization leads to unfair model comparisons and provides limited quantitative guidance for model selection under different deployment constraints, including resource-constrained devices, high-noise environments, and real-time applications. This paper proposes a systems engineering-based evaluation framework that enables comprehensive, quantitative, and scenario-specific assessment of DL-SCA models.  Methods  A multi-dimensional, scenario-based evaluation framework is developed using systems engineering principles. First, a hierarchical evaluation index system is established, comprising three criteria—attack effectiveness, resource overhead, and environmental adaptability—and six evaluation metrics: GE, SR, training time (TC), peak memory consumption (MC), model complexity (MoC), and noise robustness (Rob). Second, a standardized evaluation process based on the V-model is designed to ensure fair comparison. Each candidate model, including a Multi-Layer Perceptron (MLP), Convolutional Neural Network (CNN), and CNN-LSTM hybrid model, undergoes independent hyperparameter optimization using grid search before multi-dimensional performance evaluation. Third, a hybrid Criteria Importance Through Intercriteria Correlation-Analytic Hierarchy Process (CRITIC-AHP) Multi-Attribute Decision-Making (MADM) framework is developed. The CRITIC method derives objective weights from the statistical characteristics of the evaluation data, whereas the AHP method incorporates scenario-specific preferences through pairwise comparison matrices. The objective and subjective weights are fused to generate scenario-specific weights. Finally, a Multi-dimensional Attack Performance Metric (MAPM) is defined as the weighted sum of normalized evaluation metrics using the fused weights, providing a composite score for each model under a specific deployment scenario.  Results and Discussions  The proposed framework is validated using the ASCAD fixed-key dataset. After independent hyperparameter optimization, the three model architectures are evaluated using all six metrics. The CRITIC method produces the objective weight vector W critic = [0.17, 0.19, 0.15, 0.21, 0.14, 0.14]. Four representative deployment scenarios—Resource-Constrained, High-Performance, High-Noise, and Real-Time—are then defined, and the corresponding AHP preference weights are fused with the objective weights to generate the final scenario-specific weights. For example, MC receives the highest weight (0.52) in the Resource-Constrained scenario, whereas Rob dominates the High-Noise scenario with a weight of 0.57. The resulting MAPM scores (Table 9, Fig. 9, and Fig. 10) clearly differentiate the strengths of the evaluated models and demonstrate the scenario-specific decision capability of the proposed framework. CNN achieves the highest score in the High-Performance scenario (0.894), MLP ranks first in the Real-Time scenario (0.758) because of its shortest training time, and the CNN-LSTM hybrid model performs best in the High-Noise scenario (0.863) because of its superior noise robustness despite higher resource overhead. These results demonstrate that no single model is optimal across all deployment scenarios and that MAPM provides a clear and quantitative basis for model selection under specific deployment constraints.  Conclusions  This paper proposes a systems engineering-based, multi-dimensional evaluation framework to address the major limitations of current DL-SCA model assessment. By integrating a hierarchical evaluation index system, a standardized V-model evaluation process, and a hybrid CRITIC-AHP Multi-Attribute Decision-Making (MADM) framework, the proposed method quantitatively balances the trade-offs among attack effectiveness, resource overhead, and environmental adaptability. Experimental results obtained using the ASCAD benchmark demonstrate that the framework provides clear, quantitative, and scenario-specific guidance for model selection. The proposed Multi-dimensional Attack Performance Metric (MAPM) provides a practical decision basis for selecting DL-SCA models under diverse deployment constraints, narrowing the gap between academic attack development and practical model deployment. Future work will extend the framework to additional model architectures and datasets, improve evaluation automation, and validate its effectiveness in practical deployment environments.
Construction and Performance Analysis of Optimal Low-Hit-Zone Frequency Hopping Sequence Sets
TIAN Xinyu, CHEN Xiaoyu, ZHANG Jitao
Available online  , doi: 10.11999/JEIT260343
Abstract:
  Objective  ElectroMagnetic Interference (EMI) is a critical factor limiting the reliability of synchronization systems. Existing Fifth-Generation (5G) synchronization schemes extensively employ Zadoff-Chu (ZC) sequences to distinguish users through cyclic shifts. However, finite sequence lengths and limited orthogonal resources create substantial capacity bottlenecks in high-density access scenarios. To address these challenges, this paper investigates the problem from two perspectives. At the system level, a synchronization framework is developed by integrating Frequency Hopping (FH) with ZC sequences. By jointly exploiting code, time, and frequency-domain resources, the proposed framework improves concurrent access capability for local clusters while enhancing robustness against complex EMI through frequency diversity. At the sequence-design level, a class of multi-subset Low-Hit-Zone (LHZ) Frequency Hopping Sequence (FHS) sets is constructed to provide an efficient sequence allocation scheme for local-cluster synchronization.  Methods  Based on the theoretical framework proposed by Cai et al., the sequence mapping mechanism is reconstructed, and a disjoint Cyclic Perfect Mendelsohn Difference Family (CPMDF) is introduced to construct FHS sets that are optimal with respect to the Peng-Fan bound. The generating units are further expanded through Cartesian products, and a column-incoherent partitioning strategy is proposed to construct multi-subset LHZ FHS sets. It is proved that every nonempty subset satisfies the Peng-Fan-Lee bound with equality. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure. At the system level, an FH-ZC synchronization architecture is developed by combining predefined FH patterns with the frequency-domain correlation properties of ZC sequences for subband signal detection. A Peak-to-SideLobe Ratio (PSLR) decision metric and an early-termination strategy are adopted to evaluate synchronization preamble detection under interference. Furthermore, a multi-user simulation model is established to evaluate synchronization detection performance under accumulated co-channel collisions and EMI.  Results and Discussions  The proposed construction generates an FHS set that is optimal with respect to the Peng-Fan bound and a class of multi-subset LHZ FHS sets in which every nonempty subset is optimal with respect to the Peng-Fan-Lee bound. Example 2 demonstrates the construction procedure and the intra-subset and inter-subset Hamming correlation properties of the proposed multi-subset LHZ FHS sets. Table 1 shows that, under the same frequency-resource constraints, the proposed construction generates more sequences than existing methods under the compared parameter settings, indicating higher sequence-resource utilization. Table 2 compares the parameters of the proposed sequence sets with representative constructions reported previously and demonstrates that the proposed multi-subset optimal sequence family provides a new parameter combination. To the best of our knowledge, an optimal sequence family with a multi-subset structure has not been reported previously. Figures 2 and 3 demonstrate that the proposed FH-ZC synchronization architecture achieves a higher synchronization detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI. Figure 4 shows that the synchronization detection probability decreases as the number of active users increases because accumulated co-channel collisions degrade synchronization performance. Compared with Global-LHZ-FH-ZC, Clustered-LHZ-FH-ZC provides higher synchronization detection robustness by better matching the local-cluster competition structure characterized by strong intra-cluster competition and weak inter-cluster coupling.  Conclusions  To satisfy the sequence-capacity requirements of massive-access scenarios, this paper proposes a class of multi-subset LHZ FHS sets. By expanding the generating sequence sets through Cartesian products and partitioning subsets using a column-incoherent strategy, the proposed construction achieves both a large family size and optimal LHZ performance. The proposed multi-subset structure is well suited to local-cluster synchronization and substantially improves sequence family size and sequence-resource utilization, thereby providing a richer sequence resource pool for high-density multi-user systems. Simulation results under the considered physical-layer model demonstrate that the proposed LHZ FHS subsets reduce the effect of frequency collisions during multi-user synchronization detection. Furthermore, the FH-ZC synchronization scheme achieves a higher synchronization preamble detection probability than the conventional full-band Fixed-ZC baseline under subband-selective blocking interference caused by EMI.
Spatial-domain Anti-jamming for Unmanned Systems Under Limited Prior Information
PAN Zihao, ZHANG Bangning, ZHEN Pan, ZHU Bowen, WANG Ning, GUO Daoxing
Available online  , doi: 10.11999/JEIT260296
Abstract:
  Objective  Unmanned systems play an increasingly important role in emergency response, public safety, intelligent transportation, and other mission-critical applications. Reliable communications in complex electromagnetic environments are essential for autonomous operation. However, communication links are directly exposed to open, non-cooperative electromagnetic environments and are therefore vulnerable to intentional jamming and unintentional interference. In practical scenarios, prior information regarding the desired signal, jamming sources, and multipath propagation is often unavailable, substantially degrading the performance of conventional spatial-domain anti-jamming methods. To address this challenge, this paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information.  Methods  The proposed method first applies a spatial smoothing algorithm to the received signals to decorrelate coherent multipath components. Capon spatial spectrum estimation is then performed to detect the Direction Of Arrival (DOA) of potential incident signals. Spectrum peaks corresponding to individual incident signals are subsequently identified. A Covariance Matrix Reconstruction (CMR)-based beamforming algorithm is then applied by traversing all detected spectrum peaks to sequentially extract the signal associated with each peak, thereby separating the mixed signals. After signal separation, a signal classification method based on spectral similarity and time delay is employed. Kullback-Leibler (KL) divergence between the spectrum of each separated signal and the reference spectrum is calculated to identify jamming signals. The remaining communication signals are further classified into direct-path and multipath signals according to their relative time delays. Finally, different processing strategies are applied according to the identified signal type. Specifically, multipath signals are either suppressed as interference or coherently combined with the direct-path signal after time-delay and phase alignment.  Results and Discussions  Two simulation scenarios, including jamming only and combined jamming and multipath, are designed to evaluate the proposed method in terms of the output Signal-to-Interference-plus-Noise Ratio (SINR), beam pattern, Bit Error Rate (BER), and Error Vector Magnitude (EVM). Simulation results demonstrate that, under the jamming-only scenario, the proposed method achieves performance close to the theoretical optimum. The output SINR increases with the input Signal-to-Noise Ratio (SNR) at a fixed Jamming-to-Signal Ratio (JSR) (Fig. 3(a)) and remains nearly unchanged as JSR increases at a fixed SNR (Fig. 3(b)), indicating stable jamming suppression capability. The recovered time-domain waveform and spectrum remain highly consistent with the transmitted signal (Fig. 4). The BER curve nearly overlaps that of the optimal beamformer (Fig. 5). At \begin{document}$ {E}_{\rm b}/{N}_{0}=10\;{\mathrm{dB}} $\end{document}, the recovered Quadrature Phase-Shift Keying (QPSK) constellation closely matches the ideal constellation, achieving an EVM of –11.52 dB (Fig. 6). Under simultaneous jamming and multipath conditions, the proposed framework flexibly suppresses or exploits multipath signals. Compared with multipath suppression, multipath utilization further improves both the output SINR and BER (Fig. 7(a) and Fig. 7(b)). The corresponding beam pattern forms a beam toward the multipath direction rather than a null, demonstrating effective multipath exploitation (Fig. 7(c)).  Conclusions  This paper proposes a spatial-domain anti-jamming framework for unmanned systems operating under limited prior information. Using only the received mixed signals, the proposed framework estimates the directions of arrival, separates incident signals, and classifies them as direct-path, multipath, or jamming signals. Appropriate suppression or preservation strategies are then applied according to the identified signal type. Therefore, the framework flexibly suppresses or exploits multipath signals while preserving the direct-path signal and mitigating jamming. Simulation results demonstrate the effectiveness of the proposed method in terms of output SINR and demodulation accuracy, confirming reliable jamming suppression and communication performance even when prior information regarding the desired signal, jamming sources, and multipath propagation is unavailable. Future work will investigate the effects of array perturbations, intelligent jamming, and heterogeneous communication modes on the proposed framework and extend it to more complex unmanned-system communication environments.
Radiation-Hardened Ga2O3 MOSFET Design Featuring NiO Heterojunction and Comb-Shaped Gate Modulation
GAO Sheng, ZHANG Lin, WU Yanjun, WANG Qi, JING Liang
Available online  , doi: 10.11999/JEIT260396
Abstract:
  Objective  Gallium Oxide Metal-Oxide-Semiconductor Field-Effect Transistor (Ga2O3 MOSFET) is regarded as a promising power device for high-voltage applications, particularly in aerospace and satellite power systems, because of its ultra-wide bandgap and high critical breakdown field. However, the Conventional MOSFET (C-MOSFET) exhibits limited reliability in space radiation environments. Under off-state conditions, the electric field is highly concentrated near the gate edge. Heavy-ion irradiation generates dense electron-hole pairs along the ion track. Driven by the intense electric field, these carriers undergo avalanche multiplication through impact ionization, causing the drain current to increase sharply without recovery and ultimately leading to irreversible Single-Event Burnout (SEB) at relatively low drain bias. This failure mechanism severely limits the application of Ga2O3 MOSFETs in harsh radiation environments. Furthermore, the lack of reliable and efficient p-type doping restricts the implementation of conventional radiation-hardening techniques, including junction termination extension and junction isolation. Therefore, ionization-induced carriers readily accumulate in sensitive regions, increasing susceptibility to Single-Event Effect (SEE). The extremely low thermal conductivity of Ga2O3 further promotes local heat accumulation following heavy-ion irradiation, producing localized hot spots that increase the likelihood of thermal burnout. Existing hardening approaches, including field-plate optimization and dielectric engineering, provide only limited improvement. Moreover, the application of heterojunction structures for radiation hardening has rarely been investigated, and systematic hardening strategies have not yet been established. To address these limitations, this paper proposes a Comb-Shaped Gate Metal-Oxide-Semiconductor Field-Effect Transistor (CSG-MOSFET) incorporating a NiO heterojunction. The proposed structure redistributes the channel electric field, suppresses electric-field crowding at the conventional gate edge, and significantly improves SEB tolerance, providing an effective solution for Ga2O3 power devices operating in harsh radiation environments.  Methods  Technology Computer-Aided Design (TCAD) simulations are performed to evaluate the electrical characteristics and SEB performance of the proposed CSG-MOSFET in comparison with the C-MOSFET. The simulations incorporate high-field mobility, Shockley-Read-Hall recombination, Auger recombination, impact ionization, and heavy-ion models. Based on the charge-compensation effect of the p-NiO/n-Ga2O3 heterojunction, the proposed structure utilizes the extended depletion region formed at the heterointerface to redistribute the channel electric field. This heterojunction-induced depletion region improves electric-field uniformity and enhances SEB tolerance. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field away from the conventional gate edge, suppress local electric-field crowding, and improve device reliability under high-voltage and radiation conditions.  Results and Discussions  Simulation results demonstrate that the optimized Double Comb-Shaped Gate MOSFET (DCSG-MOSFET) significantly improves radiation hardness compared with the C-MOSFET. The SEB Threshold Voltage (VSEB) increases from 240 V to 2 280 V, while the Breakdown Voltage (BV) increases from 2 000 V to 3 500 V. Meanwhile, the specific on-resistance decreases. Therefore, the Baliga Figure of Merit (BFOM) and the SEB-based figure of merit are substantially improved. The NiO heterojunction and comb-shaped gate columns effectively redistribute the electric field, shifting the peak electric field from the conventional gate edge to the outer gate-column edge and suppressing local electric-field crowding. These improvements substantially enhance the radiation hardness of the device.  Conclusions  A radiation-hardened DCSG-MOSFET incorporating a NiO heterojunction is proposed and evaluated using TCAD simulations. The optimized structure significantly improves SEB tolerance while maintaining excellent electrical performance. Compared with the C-MOSFET, both VSEB and BV are substantially increased, demonstrating enhanced blocking capability. Charge compensation at the p-NiO/n-Ga2O3 heterojunction forms an extended depletion region that effectively redistributes the channel electric field and suppresses electric-field crowding near the conventional gate edge. Furthermore, the comb-shaped gate columns, operating together with the extended gate field plate, relocate the peak electric field to the outermost gate-column edge, thereby suppressing impact ionization induced by heavy-ion irradiation and effectively mitigating SEB. The reduced specific on-resistance further improves the BFOM and the SEB-based figure of merit. These results demonstrate that the proposed DCSG-MOSFET is a promising candidate for power electronic applications in harsh radiation environments, including aerospace and satellite systems.
THz Ultra-Massive MIMO Channel Estimation via a Noise-Conditioned Fixed-Point Network
ZHANG Huawei, NIU Yaning, JIANG Zhanjun, LIU Yingting
Available online  , doi: 10.11999/JEIT260420
Abstract:
  Objective  Terahertz (THz) Ultra-Massive Multiple-Input Multiple-Output (UM-MIMO) systems are expected to support future high-capacity wireless communications. However, accurate channel estimation remains challenging under hybrid near-/far-field propagation and Array-of-SubArrays (AoSA) architectures, where limited Radio-Frequency (RF) chains, low Signal-to-Noise Ratio (SNR), noise uncertainty, and structural perturbations degrade compressed observations. Existing compressed sensing, Bayesian inference, and deep unfolding methods generally rely on fixed statistical assumptions, which limit their cross-SNR generalization under varying noise conditions and statistical mismatches. To address these limitations, this paper proposes a Noise-Conditioned Fixed-Point Network (FPN-NCAS) for robust THz UM-MIMO channel estimation. The proposed method aims to improve estimation accuracy, cross-SNR generalization, robustness, and iterative stability by incorporating noise-aware nonlinear recovery.  Methods  FPN-NCAS is developed within an Orthogonal Approximate Message Passing (OAMP)-based fixed-point unfolding framework. A coarse noise power estimate is obtained from repeated pilot differences and injected into the nonlinear recovery module as an explicit conditioning variable. After each linear update, the vector-domain estimate is reshaped into an AoSA-aligned tensor to exploit subarray-level structural priors. The nonlinear recovery chain consists of three components. Token-Gate performs lightweight subarray-level reliability pre-calibration to suppress unreliable responses under low-SNR and structurally inconsistent conditions. Block Shrink serves as the core noise-conditioned block-sparse proximal operator, in which a smooth dual-threshold mechanism balances strong denoising at low SNR with structural preservation at medium-to-high SNR. G-HMTD(Gated Hybrid Multi-scale Transformer Denoiser) further refines residual errors by combining Local Multi-scale Enhancement and Global Context Modeling. Bridge relaxation and nonlinear residual scaling are also incorporated to improve inter-stage stability.  Results and Discussions  Simulation results demonstrate that FPN-NCAS consistently outperforms LS, OAMP, ISTA-Net+, FPN-OAMP, and FPN-OTFN over the 0~20 dB SNR range (Fig. 6). At SNR = 0 dB, FPN-NCAS achieves Normalized Mean Square Error (NMSE) gains of approximately 3.0 dB and 1.9 dB over FPN-OAMP and FPN-OTFN, respectively. At SNR = 20 dB, the gains increase to approximately 4.5 dB and 2.6 dB (Fig. 6(a)). The convergence curves show that FPN-NCAS reaches a stable plateau after approximately three layers at SNR = 5 dB and five layers at SNR = 15 dB, demonstrating stable fixed-point iterative behavior (Fig. 6(b) and Fig. 6(c)). Analysis of repeated pilots shows that four repeated pilot pairs introduce only 3.13% additional pilot overhead while reducing the relative standard deviation of the coarse noise power estimate to 25.00%. Under moderate noise power mismatch, NMSE degradation remains within 0.3 dB. The proposed method also maintains strong robustness under colored Gaussian noise, impulsive noise, near-/far-field distribution shifts, variation in the number of propagation paths, AoSA subarray shuffling, and RF amplitude/phase mismatch (Fig. 7 and Fig. 8). Ablation studies indicate that Block Shrink provides the largest performance gain, whereas Token-Gate and G-HMTD further improve performance through structural calibration and residual refinement (Fig. 9).  Conclusions  This paper proposes FPN-NCAS for noise-conditioned fixed-point channel estimation in THz UM-MIMO systems. By integrating repeated-pilot-based noise conditioning, AoSA-aware feature reshaping, Token-Gate calibration, Block Shrink recovery, and G-HMTD refinement, the proposed method improves NMSE performance, robustness, and iterative stability under different SNR conditions and non-ideal scenarios. The improved performance is achieved at the cost of higher inference complexity. Future work will focus on lightweight implementations and extensions to wideband, multi-user, and hardware-impaired THz communication systems.
Decision Learning Correction Network: Fusion Classification of Hyperspectral Images and LiDAR Data
WANG Haoyu, LIU Nuofei, CHENG Yuhu, LIU Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260362
Abstract:
  Objective  HyperSpectral Images (HSI) and Light Detection and Ranging (LiDAR) provide complementary information for land-cover classification. HSI captures rich spectral information for material discrimination, while LiDAR provides elevation and structural information for spatial characterization. However, most existing fusion methods treat multimodal fusion as a one-shot static aggregation process, implicitly assuming that a fixed fusion strategy is applicable to all pixels and regions. This assumption is difficult to satisfy in complex remote sensing scenes, where class-boundary and cross-modal heterogeneous regions exhibit high information density but account for only a small proportion of samples (Fig. 1). To address this limitation, this paper proposes a Decision Learning Correction Network (DLCN) that reformulates static HSI-LiDAR fusion as a context-dependent sequential decision-making process.  Methods  The proposed DLCN consists of feature extraction, fusion decision learning, and classification. First, HSI and LiDAR are processed through two parallel branches to extract spectral and spatial features and elevation and structural features, respectively. The extracted features are then concatenated to form the current state and are fed into an Actor-Critic framework. The Actor network generates fusion actions to adaptively adjust modality contributions, while the Critic network evaluates the long-term value of each action for classification. To improve learning from difficult samples, a key-sample-oriented sampling module assigns higher sampling probabilities to samples with larger modal fidelity loss. Meanwhile, a modal fidelity constraint mechanism evaluates spectral fidelity, feature consistency, structural preservation, and resolution matching, and corrects destructive actions during fusion. Through this closed-loop framework, DLCN performs dynamic generation, evaluation, and correction of fusion actions, thereby producing high-quality fusion features for classification (Fig. 2).  Results and Discussions  Experiments are conducted on the Houston2013, Trento, and MUUFL datasets. DLCN achieves the highest Overall Accuracy (OA) of 97.85%, 99.58%, and 94.38% on the three datasets, respectively, outperforming CHNet, DSymFuser, mPMCL, MEDFN, S3F2Net, and MSAF. The classification maps demonstrate that DLCN effectively reduces misclassification in class-boundary, mixed land-cover, and structurally complex regions, producing results that more closely match the ground-truth maps across all three datasets (Figs. 35). Ablation studies further demonstrate that the value-guided policy optimization mechanism, key-sample-oriented sampling module, and modal fidelity constraint mechanism each improve classification performance. Compared with the baseline models, the complete DLCN consistently increases OA on Houston2013, Trento, and MUUFL, validating the effectiveness of the proposed decision-learning-correction framework. Time-step analysis shows that DLCN progressively improves classification accuracy while maintaining stable spectral-angle variation during sequential decision making (Fig. 6). Furthermore, DLCN achieves inference times of 1.32 s, 0.86 s, and 2.23 s on the three datasets, respectively, ranking first among the compared methods. These results indicate that the additional computation introduced by the Actor-Critic decision framework and modal fidelity constraint mechanism is effectively translated into improved classification performance without imposing excessive computational cost.  Conclusions  This paper proposes a DLCN for HSI and LiDAR fusion classification. Unlike conventional static fusion methods, DLCN formulates multimodal fusion as a sequential decision-making process and adaptively adjusts fusion strategies according to the local context. Its closed-loop framework enables fusion actions to be generated, evaluated, and corrected throughout the decision process, thereby producing high-quality fusion features for classification. Experimental results demonstrate that DLCN produces more accurate classification maps in heterogeneous remote sensing scenes, and the time-step analysis further confirms the stability of the sequential decision-making process. Future work will focus on more fine-grained feature representation and more robust policy optimization to improve model generalization in complex remote sensing scenes.
A Nested Multi-scroll Memristive Hopfield Neural Network and Its Hardware Implementation
WANG Zhe, WAN Qiuzhen, ZHOU Pan, RAO Huhui
Available online  , doi: 10.11999/JEIT260516
Abstract:
  Objective  In recent years, memristors have been employed to emulate neuronal synapses with dynamically adjustable synaptic weights, enabling the construction of Memristive Hopfield Neural Networks (HNNs). Compared with conventional HNNs, Memristive HNNs more accurately reproduce the nonlinear dynamical behavior of biological neural systems. Multi-scroll attractors have attracted considerable attention in secure communication because of their complex topological structures and strong state-space ergodicity. However, previous studies have primarily focused on conventional multi-scroll attractors with single structural patterns, whereas multi-scroll attractors with special structures remain largely unexplored. Therefore, this paper proposes a nested multi-scroll Memristive HNN system that generates nested multi-scroll attractors, thereby overcoming the limitations of conventional single-structure multi-scroll attractors.  Methods  A Four-Dimensional (4D) Memristive HNN system is constructed from a three-neuron HNN by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of neuron 2. Equilibrium-point and stability analyses are performed to investigate the regulatory effects of the Memristive self-connected synapse coupling strength and system initial conditions on the system dynamics. The number of multi-scroll attractors is regulated by adjusting the memristor control parameters. Building on this framework, a Multi-level Logic Pulse current (IMLP) is introduced to construct a nested multi-scroll Memristive HNN system. The proposed system generates nested multi-scroll attractors with enhanced dynamical complexity. Finally, the MATLAB numerical simulation results are validated through Multisim circuit simulations and Field-Programmable Gate Array (FPGA)-based hardware experiments.  Results and Discussions  The results demonstrate that regulating the Memristive self-connected synapse coupling strength enables the proposed 4D Memristive HNN system to exhibit period-doubling bifurcations and chaotic behavior, as illustrated by the bifurcation diagrams and Lyapunov exponent spectra (Fig. 3). Various types of coexisting attractors are generated under different coupling strengths (Fig. 4). By adjusting the memristor control parameters, multi-scroll attractors with different numbers of scrolls are generated through one-directional extension (Figs. 58). After the introduction of the IMLP, the proposed nested multi-scroll Memristive HNN system generates nested multi-scroll attractors while preserving the controllable scroll-number extension property (Figs. 810). Spectral Entropy (SE) analysis demonstrates that the IMLP increases the dynamical complexity of the proposed system compared with the original 4D Memristive HNN system (Figs. 9 and 10). The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments confirms the physical realizability of the proposed nested multi-scroll Memristive HNN system (Figs. 1214).  Conclusions  A 4D Memristive HNN system is constructed by incorporating a multi-segment nonlinear magnetically controlled memristor into the Memristive self-connected synapse of a three-neuron HNN. Equilibrium-point and stability analyses reveal the regulatory effects of the Memristive self-connected synapse coupling strength and the evolution of coexisting attractors associated with different initial conditions. The results show that the system enters chaos through the period-doubling route to chaos and generates single-scroll and double-scroll chaotic attractors. The number of multi-scroll attractors is continuously increased by adjusting the memristor control parameters. Furthermore, introducing the IMLP produces a nested multi-scroll Memristive HNN system capable of generating nested multi-scroll attractors with increased dynamical complexity. The strong agreement among MATLAB numerical simulations, Multisim circuit simulations, and FPGA-based hardware experiments validates the physical realizability of the proposed nested multi-scroll Memristive HNN system.
Research on Ka-band Enhanced Active Load Modulation Ultra-wideband High-efficiency Doherty Power Amplifier
YANG Lin, YAN Chengyu, WANG Yanping, ZHANG Ming, WANG Baozhu, HAN Qi, HE Yuhang, HOU Weimin, LI Kang
Available online  , doi: 10.11999/JEIT260514
Abstract:
  Objective  The Ka-band has become a key frequency band for satellite communications, placing stringent requirements on the millimeter-wave power amplifier, a core component of the transmitter, to provide high efficiency, compact size, and broadband operation. To maximize spectral efficiency, millimeter-wave satellite communication signals typically exhibit a high Peak-to-Average Power Ratio (PAPR), making high back-off efficiency particularly important. Although the Doherty Power Amplifier (DPA) is widely adopted because of its high efficiency under power back-off conditions, its operating bandwidth is inherently limited. In addition, both saturated efficiency and back-off efficiency degrade substantially at millimeter-wave frequencies. Therefore, extending the operating bandwidth while maintaining high efficiency remains a major challenge for millimeter-wave DPAs used in satellite communication transmitters.  Methods  An ultra-wideband enhanced active load modulation technique is proposed to overcome the trade-off between bandwidth and back-off efficiency in DPAs. The proposed method achieves optimal load impedance modulation for both the carrier and peaking amplifiers over an ultra-wide frequency range by introducing an Impedance Tunable Bias Network (ITBN) and a dual-drive impedance control mechanism. These techniques improve load modulation while extending the load modulation bandwidth, thereby enhancing both efficiency and bandwidth in the millimeter-wave DPA architecture. Furthermore, a broadband phase-compensation technique is integrated into an unequal power division network to achieve sufficient load modulation across the entire operating band while accurately compensating for the phase difference between the carrier and peaking paths. The proposed ultra-wideband phase-compensated power division network further extends the high-efficiency operating bandwidth while reducing the chip area.  Results and Discussions  To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was designed and fabricated using a 0.15-μm GaN process. Across the 24~33 GHz frequency band, corresponding to a relative bandwidth of 31.6%, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB with a gain flatness of ±1 dB. The measured saturated output power is 28.9~31.0 dBm, with a Power-Added Efficiency (PAE) of 25.2%~33.5% at saturation and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating over similar frequency bands, the proposed design achieves the highest reported small-signal gain and saturated PAE while maintaining high saturated output power and high 6 dB back-off PAE. Furthermore, it occupies the smallest chip area among reported three-stage DPA MMICs.  Conclusions  An ultra-wideband enhanced active load modulation method is proposed to achieve sufficient active load modulation through multi-frequency impedance tuning across the entire operating band. A novel ultra-wideband phase-compensated unequal power division network is also proposed to reduce the chip area while maintaining accurate phase compensation. To validate the proposed method, a millimeter-wave ultra-wideband high-efficiency DPA was fabricated using a 0.15-μm GaN process. Measurement results demonstrate that, over the 24~33 GHz frequency band, the fabricated chip achieves a small-signal gain of 21.8~23.8 dB, a saturated output power of 28.9~31.0 dBm, a saturated PAE of 25.2%~33.5%, and a 6 dB back-off PAE of 15.5%~19.8%. Compared with previously reported GaN DPAs operating in similar frequency bands, the proposed design achieves a maximum relative bandwidth of 31.6%, the highest reported small-signal gain and saturated PAE, high saturated output power, and high 6 dB back-off PAE, while occupying the smallest chip area among reported two- and three-stage DPA MMICs. These measurement results validate the proposed method and demonstrate its strong potential for millimeter-wave satellite communication transmitters.
Iterative Parameter Estimation Method for Energy Detection Threshold in Ambient Backscatter
QU Wenfeng, YE Yinghui, SHI Liqin, LU Guangyue
Available online  , doi: 10.11999/JEIT260418
Abstract:
  Objective  In Ambient Backscatter Communication (AmBC) systems, a low-complexity Energy Detector (ED) is commonly employed at the reader to recover symbols transmitted by the tag. The detection performance of ED depends strongly on the accurate setting of the detection threshold, which is determined by the average received signal power corresponding to tag symbols “1” and “0”. Existing parameter estimation methods assume that the two symbols are transmitted with equal probability and therefore divide the sorted received signal power samples into two equal groups. However, because the number of transmitted symbols is finite, the actual numbers of symbols “1” and “0” are generally unequal. Therefore, equal partitioning introduces sample misclassification, causing the estimated threshold to deviate from its optimal value and reducing detection performance. To address this limitation, an iterative threshold parameter estimation method is proposed to reduce the parameter estimation bias caused by sample misclassification and improve the accuracy of detection threshold estimation.  Methods  An iterative threshold parameter estimation method is proposed to overcome the sample misclassification introduced by conventional sorting-based grouping. Because the initial detection threshold obtained by the sorting-based grouping method provides reliable decisions for most received samples, these initial decisions are used as the basis for sample reclassification. The received signal samples are then reclassified to iteratively update the threshold parameters, progressively refining the detection threshold. The proposed method is evaluated through simulations under three representative ambient radio-frequency source conditions: complex Gaussian, Phase-Shift Keying (PSK), and Quadrature Amplitude Modulation (QAM) sources.  Results and Discussions  Simulation results show that, over a wide range of Signal-to-Noise Ratio (SNR) values, the proposed iterative method substantially reduces the Bit Error Rate (BER) compared with the conventional sorting-based grouping method and approaches the theoretical lower bound obtained with perfect parameter estimation. At a given SNR, the proposed method improves BER by approximately 0.1, 1.5, and 1.3 orders of magnitude under complex Gaussian, PSK, and QAM sources, respectively (Fig. 2). These results demonstrate that the proposed iterative method effectively corrects sample misclassification and reduces the performance loss caused by parameter estimation bias. Moreover, most of the performance gain is achieved after only one iteration, indicating rapid convergence with minimal additional computational overhead. Under different numbers of sampling points, BER improvements of approximately 0.5, 1.6, and 1.1 orders of magnitude are achieved under complex Gaussian, PSK, and QAM sources, respectively (Fig. 3). These results indicate that using the initial decisions for sample reclassification effectively reduces the estimation bias introduced by fixed equal partitioning, thereby improving detection performance under limited-sample conditions. Under different Relative Channel Difference (RCD) values, BER improvements of approximately 0.56 and 0.7 orders of magnitude are achieved under complex Gaussian and QAM sources, respectively (Fig. 4). As the RCD increases, the separation between the received signal power distributions becomes more pronounced, improving the accuracy of the initial decisions and enabling more reliable sample reclassification. This positive feedback process further refines the parameter estimates and improves detection performance.  Conclusions  An iterative threshold parameter estimation method is proposed to address the sample misclassification introduced by conventional sorting-based grouping in ED. The proposed method uses the initial decisions to reclassify the received signal samples and iteratively update the threshold parameters. In addition, closed-form expressions for the detection threshold and BER under QAM sources are derived. Simulation results demonstrate that the proposed method effectively reduces parameter estimation bias with only one iteration while maintaining robust performance under limited-sample and varying channel conditions. Significant BER improvements are achieved with minimal additional computational overhead, making the proposed method well suited for practical, high-reliability AmBC systems.
LLMA-GCN: A Semantically Enhanced Hierarchical Spatiotemporal Graph Convolutional Network for Skeleton-based Action Recognition
JIA Guimin, ZHOU Xilong
Available online  , doi: 10.11999/JEIT260154
Abstract:
  Objective  Current methods that introduce the Large Language Model (LLM) into skeleton-based action recognition face three main limitations. Semantic guidance remains decoupled from spatial topology learning, temporal modeling lacks hierarchical semantic support, and traditional classification paradigms have limited generalization. To address these issues, this paper proposes LLMA-GCN, a semantically enhanced graph convolutional framework. The proposed framework integrates LLM-derived semantic prior knowledge with graph convolution to improve spatiotemporal feature learning and action classification.  Methods  LLMA-GCN uses visual skeleton data and semantic inputs in a dual-branch framework. Frozen LLMs and prompt engineering are used to precompute the joint semantic adjacency matrix and action text prototypes. The framework consists of three main components: an LLM-based hybrid graph topology learning strategy, a hierarchical visual sequence encoder based on the LLM Refinement Block (LRB), and an action text-prototype-guided decision learning mechanism. These components enable semantic guidance in graph topology learning, hierarchical spatiotemporal feature extraction, and visual-text alignment.  Results and Discussions  Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD I show that LLMA-GCN achieves competitive or superior performance compared with state-of-the-art methods. Ablation studies confirm the key roles of the hybrid graph topology, the LRB, and the action text-prototype-guided decision learning mechanism. Model-complexity analysis further indicates the potential of the proposed framework for practical application.  Conclusions  By fusing the joint semantic adjacency matrix and the physical adjacency matrix, LLMA-GCN enables action semantics to directly guide graph convolution and improves semantic perception. The LRB embeds semantic information into hierarchical spatiotemporal feature extraction, which strengthens the modeling of complex actions. The action text-prototype-guided decision learning mechanism further shifts skeleton-based action recognition from purely visual classification to text-prototype-guided visual-text alignment. Overall, LLMA-GCN provides a robust and generalizable framework for skeleton-based action recognition through deep fusion of visual and semantic features.
A Phase Transition Obstacle Avoidance Method for UAV Swarms Driven by Multistable Potential Fields
HE Ming, CHEN QiYang, HAN Wei, PAN Fan, MA YiSong
Available online  , doi: 10.11999/JEIT260357
Abstract:
  Objective  Unmanned Aerial Vehicle (UAV) swarms have demonstrated considerable potential for complex missions, such as search, surveillance, and disaster response, because of their distributed coordination and robustness. However, in dynamic environments with dense obstacles and rapidly changing risks, conventional swarm control methods often exhibit discontinuous behavior switching and control chattering, which reduce system stability and coordination efficiency. Existing approaches, including threshold-based switching and Artificial Potential Field (APF) methods with fixed potential weights, rely on abrupt transitions between behavioral modes, leading to oscillatory responses. To address these limitations, a phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. Swarm behavior evolution is modeled as a continuous phase transition process within a unified potential field framework, enabling smooth and adaptive transitions between formation flight and obstacle avoidance.  Methods  An environmental risk assessment model is first established by integrating static obstacle risk, dynamic obstacle risk, and inter-agent proximity risk. A distributed consensus protocol is then employed to establish global risk consensus. Subsequently, a morphology factor is generated through nonlinear mapping of the global risk consensus and is used as an order parameter to characterize the macroscopic swarm state. A unified time-varying potential field, comprising formation, obstacle avoidance, and navigation potentials, is constructed, and the relative weights of these potentials are continuously adjusted by the morphology factor. When the risk level is low, the system exhibits a monostable structure dominated by the formation and navigation potentials. As the risk increases, the potential field continuously evolves into a multistable structure dominated by the obstacle avoidance potential, thereby enabling distributed obstacle avoidance. A distributed consensus control law based on the negative gradient of the unified potential field is further developed. A damping term is incorporated to dissipate system energy and improve stability, while a dynamic compensation term addresses nonlinear dynamics. The control law depends only on local information, ensuring good scalability. The global uniform ultimate boundedness of the closed-loop system is established using Lyapunov theory.  Results and Discussions  Simulation results demonstrate that the proposed method enables the swarm to maintain a compact solid-phase swarm formation in low-risk regions and to transition smoothly to a dispersed liquid-phase swarm configuration when obstacles are encountered, followed by rapid formation recovery after obstacle avoidance. The pitch and roll angles of each UAV vary smoothly without abrupt changes, and both the UAV-to-obstacle distance and the inter-UAV separation remain above the prescribed safety threshold throughout the flight, ensuring collision-free operation. Statistical results obtained from 20 independent simulation runs show that, compared with the threshold-switching method, the proposed method reduces the rate of control input variation by approximately 26% and decreases the peak control input by approximately 18%. Compared with the bio-inspired diversion method, the average formation recovery time after obstacle avoidance is reduced by approximately 16%. Ablation experiments further demonstrate that removing the morphology-driven phase transition mechanism significantly increases trajectory oscillation and control oscillation, confirming the critical role of the multistable continuous phase transition mechanism in maintaining smooth swarm motion. In complex narrow-channel environments, the proposed method effectively avoids the local minimum problem encountered by conventional APF methods and generates smoother flight trajectories with substantially reduced oscillation.  Conclusions  A phase transition obstacle avoidance method for UAV swarms driven by multistable potential fields is proposed. By introducing a morphology factor and constructing a unified potential field framework, swarm behavior evolution is represented as a continuous phase transition process. The distributed control law enables smooth behavioral transitions while maintaining system stability and scalability. Simulation results demonstrate that the proposed method achieves better safety, smoother control, and higher coordination efficiency than conventional methods.
Low-Complexity Phase Ambiguity Resolution DOA EstimationAlgorithm for Composite Hierarchical Receiving Array Structure
CHEN Yiwen, DONG Yangze, CHEN Xiahua, LING Wenchang, XIONG Yiwen
Available online  , doi: 10.11999/JEIT260447
Abstract:
  Objective  Direction Of Arrival (DOA) estimation is a key technique for sonar target localization. As the demand for high-precision DOA estimation in complex environments continues to increase, the number of array elements used for estimation is steadily growing, leading to massive arrays. Although larger arrays improve DOA estimation accuracy and resolution, they also impose a substantial computational burden on conventional DOA estimation algorithms. To address this issue, a low-complexity composite hierarchical receiving array structure is constructed, and two fast phase ambiguity resolution algorithms are proposed: Composite HierArchical Global Nearest-Neighbor Matching (CHA-GNNM) and Composite HierArchical Cross-Correlation Covariance Merging (CHA-CCM).  Methods  The CHA-GNNM algorithm constructs multiple candidate solution sets by exploiting the auto-covariance and cross-covariance relationships among the subarrays within each group. The true solution in each candidate solution set is identified through nearest-neighbor matching based on source consistency, and the final DOA estimate is obtained through multilevel coherent combining. This approach achieves phase ambiguity resolution and angle matching with relatively low computational cost. However, because the correlation information among all array elements is not fully exploited, some estimation performance is sacrificed. To improve DOA estimation performance, the CHA-CCM algorithm reorganizes the composite hierarchical structure into evenly partitioned groups, which are regarded as several large subarrays. Multiple large candidate solution sets are first constructed from the cross-correlation relationships among these groups. Each group is then divided into multiple small subarrays, from which additional candidate solution sets are generated using the corresponding auto-covariance and cross-covariance relationships. A coarse DOA estimate is obtained through coprime clustering, followed by a more accurate initial DOA estimate derived from the small candidate solution sets. This initial DOA estimate is subsequently used to eliminate spurious solutions from the large candidate solution sets, yielding the final DOA estimate. Combined with a low-complexity covariance block-processing strategy, this approach avoids computationally expensive operations while improving DOA estimation accuracy.  Results and Discussions  Simulation results demonstrate that both proposed algorithms substantially reduce the computational burden as the number of array elements increases, while effectively achieving phase ambiguity resolution through the proposed composite hierarchical receiving array structure (Fig. 5). Compared with the conventional Root-MUSIC algorithm, CHA-GNNM achieves coarse DOA estimation with nearly four orders of magnitude lower computational complexity (Fig. 7), making it suitable for applications with stringent real-time requirements. In contrast, CHA-CCM requires only a modest increase in computational cost (Fig. 7) while achieving DOA estimation performance close to the Cramér-Rao Lower Bound (CRLB) above a certain signal-to-noise ratio threshold (Fig. 6). Therefore, a favorable balance is achieved between DOA estimation accuracy and computational complexity.  Conclusions  To address the rapid increase in computational complexity associated with massive arrays, a composite hierarchical receiving array structure is constructed for efficient DOA estimation. By hierarchically grouping the array elements, the proposed structure provides a new framework for low-complexity DOA estimation. Based on this structure, two fast DOA estimation algorithms are developed. Both algorithms achieve effective phase ambiguity resolution with low computational complexity by exploiting the structural differences among array groups and the consistency of observations from the same source across different groups, thereby enabling rapid DOA estimation. CHA-GNNM primarily exploits the phase relationships among subarrays to perform phase ambiguity resolution and angle matching through a simple computational procedure, making it suitable for applications requiring high computational efficiency and real-time processing. Because the cross-correlation information among all array elements is not fully exploited, some estimation performance is reduced under challenging signal conditions. To overcome this limitation, CHA-CCM reorganizes the composite hierarchical receiving array into evenly partitioned groups while preserving the low-complexity advantage of the hierarchical structure. Group-level cross-correlation information is further exploited so that the intrinsic relationships among different groups are more fully utilized. In addition, the signal processing procedure is simplified by eliminating unnecessary computational steps, thereby improving the robustness and accuracy of DOA estimation while maintaining manageable computational complexity. Compared with CHA-GNNM, CHA-CCM incurs only a small increase in computational cost and achieves a better balance between computational complexity and DOA estimation performance. Overall, the proposed composite hierarchical receiving array structure and the two fast DOA estimation algorithms provide an effective solution for efficient DOA estimation in massive arrays. CHA-GNNM is more suitable for applications with stringent real-time requirements, whereas CHA-CCM is better suited for applications requiring higher DOA estimation accuracy and robustness. The proposed structure achieves efficient phase ambiguity resolution and accurate DOA estimation and provides both theoretical significance and practical value for engineering applications of massive array signal processing.
A Hierarchical Cross-layer Closed-loop Learning Framework andCoordination Mechanism for Complex Multi-agent Systems
ZHANG Long, HUANG wenbo, LEI Zhen, FENG Xuanming, WANG Ying
Available online  , doi: 10.11999/JEIT260143
Abstract:
Complex Multi-Agent Systems (MAS) in dynamic and uncertain environments face challenges in unified modeling, adaptive coordination, and interpretable effectiveness evaluation. Existing methods usually address individual decision-making, inter-agent coordination, and high-level policy evolution separately. This separation leads to fragmented decision chains and weak cross-layer coupling. It also makes it difficult to explain how local learning gains are transformed into global effectiveness improvements under mission variation, observation disturbance, and structural damage. To address this issue, a Hierarchical Cross-layer Closed-loop Learning (HCCL) framework is proposed. The framework couples individual autonomy, system-level coordination, and system-of-systems learning to build a computable path from local policy optimization to overall effectiveness enhancement.   Methods   HCCL adopts a unified three-layer architecture. At the individual autonomy layer, each agent is modeled as a Partially Observable Markov Decision Process (POMDP) to describe decision-making under partial observability. At the system-level coordination layer, multi-agent coordination is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) and represented by a dynamic directed weighted coordination graph. A Graph Neural Network (GNN) is used to encode interaction dependencies, structural coupling, and joint value information. At the system-of-systems learning layer, a Meta-Decentralized Partially Observable Markov Decision Process (Meta-Dec-POMDP) is established to describe task-context adaptation and rule evolution. A cross-layer closed-loop mechanism is further designed. In the bottom-up behavior induction pathway, local state and capability features are aggregated into graph-level structural representations and supplied to the upper rule-learning process. In the top-down rule-shaping pathway, learned high-level rules are converted into control parameters and fed back to lower layers to regulate local policies and coordination relationships. Simulations are conducted under baseline, mission-variation, observation-disturbance, and structural-damage scenarios. The full HCCL model is compared with a non-closed-loop model and an upward-induction-only model. Interface ablation studies are also performed to analyze the contributions of cross-layer feature reporting, structural induction, and rule shaping.   Results and Discussions   The full HCCL model consistently outperforms the comparison models and ablated variants. In the baseline scenario, it achieves a task success rate of 88.6% and a comprehensive system effectiveness of 0.842. Under mission variation, it reduces the adaptation process to 16±2 rounds. Under structural damage, it achieves a recovery rate of 81.4% and restores coordination-structure stability to 0.742 within 20 steps. These results indicate that HCCL improves task performance, adaptation speed, and structural recovery. Ablation results show that removing any cross-layer interface reduces performance, while removing the top-down rule-shaping pathway causes the largest loss. This result indicates that upward structural perception alone is insufficient for sustained system-level improvement. The effectiveness gain mainly arises from closed-loop coupling between bottom-up behavior induction and top-down rule shaping, rather than from simple hierarchical stacking.   Conclusions   The HCCL framework is proposed for complex MAS by integrating POMDP-based individual autonomy modeling, Dec-POMDP- and graph-based coordination modeling, and Meta-Dec-POMDP-based rule evolution. Through bottom-up behavior induction and top-down rule shaping, HCCL provides a computable and interpretable path from local learning to overall effectiveness enhancement. Experimental results verify its advantages in task completion, adaptation, recovery, and coordination stability under multiple disturbances. Future work will focus on larger-scale heterogeneous systems, communication-constrained networking, online continual adaptation, and data-driven evaluation in realistic environments.
Off-grid Blind Near-Field Integrated Sensing And Communication: Algorithm Design and Lower Bound
YUAN Zhengdao, GUO Qinghua, HUANG Chongwen, GAO Dawei, MEI Fengtong, LIAO Guisheng
Available online  , doi: 10.11999/JEIT260404
Abstract:
  Objective  With the widespread deployment of extra-large-scale antenna arrays in 6G networks, user terminals are increasingly located in the near-field region. Existing Near-Field Integrated Sensing And Communication (NF-ISAC) algorithms face key challenges, including off-grid power leakage, severe model mismatch, and strong pilot dependence. These limitations make them unsuitable for low-overhead, high-performance 6G transmission. This paper aims to design an off-grid blind NF-ISAC algorithm and derive the theoretical performance bound for near-field sensing.  Methods  To overcome the limitations of analytical geometric steering vectors and adapt to more accurate electromagnetic propagation characteristics without closed-form expressions, an amplitude-phase separation method is first proposed. This method decomposes the nonlinear near-field steering vector into amplitude and phase terms, enabling high-precision characterization of the steering vector using a single-hidden-layer neural network. Second, the NF-ISAC problem is formulated as a constrained matrix factorization problem, and a corresponding factor graph model is constructed. The trained neural network is embedded into the factor graph as a function node. Message passing through the embedded neural network is then achieved, enabling joint blind coordinate sensing, channel estimation, and signal detection in a pilot-free manner. Finally, the Cramér-Rao Lower Bound (CRLB) for multi-user near-field joint distance and angle sensing in polar coordinates is derived based on the neural-network-fitted steering vector.  Results and Discussions  Extensive Monte Carlo simulations are conducted to evaluate the performance of the proposed algorithm. The simulation results show that the proposed algorithm achieves millimeter-level position sensing. Compared with existing mainstream algorithms, it improves both communication Bit Error Rate (BER) and sensing accuracy. The proposed algorithm achieves a 2~3 dB gain in sensing accuracy over the state-of-the-art near-field off-grid algorithm, and its performance is closest to the derived theoretical CRLB. These results indicate that the proposed algorithm effectively mitigates off-grid power leakage and model mismatch.  Conclusions  The proposed off-grid blind NF-ISAC algorithm overcomes the pilot dependence and model mismatch of existing NF-ISAC schemes. It achieves integrated high-precision sensing and reliable communication for near-field users in a pilot-free manner. The derived CRLB provides a theoretical benchmark for evaluating the sensing performance of NF-ISAC systems. This work provides technical support for the design of 6G NF-ISAC systems.
Explicit Discrimination-driven Automatic Unknown Class Clustering for Open-World Semi-Supervised Learning
SONG Jialun, DU Lan, CHEN Jian
Available online  , doi: 10.11999/JEIT251291
Abstract:
  Objective   Traditional target recognition is generally developed under the closed-set assumption, in which all test classes are assumed to be included in the training set. In real-world scenarios, however, unlabeled unknown classes are commonly encountered. Therefore, recognition systems should simultaneously recognize known classes and discover and cluster unknown classes. To address this challenge, a transductive Open-World Semi-Supervised Learning (OWSSL) method driven by explicit discrimination and automatic unknown-class clustering is proposed.   Methods   The proposed method is trained using a small set of labeled known-class samples and a large set of unlabeled test samples containing both known and unknown classes. It consists of two complementary modules. The Dynamic Known-Unknown Class Discrimination (DKUCD) module models the known-class boundary distribution using Extreme Value Theory (EVT) and progressively refines the distribution with high-confidence known-class samples during semi-supervised learning to improve known-unknown discrimination. The Neighbor Intersection-Over-Union Cluster Merging (NIOUCM) module automatically clusters high-confidence unknown-class samples by merging neighboring clusters according to their intersection-over-union relationships. The DKUCD and NIOUCM modules are optimized iteratively to improve discrimination and unknown-class clustering jointly.   Results and Discussions   Experiments conducted on the optical CIFAR-10 dataset and measured radar datasets demonstrate that the proposed method achieves accurate known-class recognition while effectively clustering unknown classes.   Conclusions   By explicitly discriminating between known and unknown classes and automatically estimating unknown-class clusters, the proposed method improves both known-class recognition and unknown-class clustering under open-world conditions.
Network Metric System and Scenario-Differentiated Analysis Driven by LLM Literature Mining
XU Qikun, LIU Yaxi, HAN Shuxian, ZHANG Huifeng, HUANGFU Wei
Available online  , doi: 10.11999/JEIT251120
Abstract:
  Objective  Network metrics provide the foundation for network design, operation, and optimization. Existing studies primarily focus on individual scenarios or representative metrics and lack unified extraction rules and reproducible workflows for large-scale, cross-scenario metric analysis. To address terminology ambiguity, scenario heterogeneity, and the quantification of complex metric relationships, this study proposes a reproducible domain-specific literature mining framework based on a Large Language Model (LLM). The framework automatically extracts and standardizes network metrics, annotates application scenarios, quantifies inter-metric relationships, and establishes a Service-Multiplexing-Versatility (SMV) analytical framework. Rather than providing a complete set of metric calculation methods, the SMV framework serves as a conceptual model for guiding multi-objective tradeoffs in network architecture design and lifecycle management.  Methods  An automated literature mining framework based on a multi-agent LLM architecture is developed (Fig. 1). A dataset comprising 583 articles published in IEEE/ACM Transactions on Networking during 2023–2024 is analyzed. The framework consists of three specialized agents. A terminology normalization agent maps aliases and synonymous expressions to standardized metric names. A scenario annotation agent assigns primary application scenario labels using high-information-density sections of each article. A correlation mining agent identifies the semantic direction and strength of relationships between metric pairs and quantifies these relationships as signed correlation coefficients ranging from −1 to +1. The reliability of the mining results is evaluated through dual-LLM cross-validation and manual sampling review (Fig. 2).  Results and Discussions  The proposed framework extracts 3,978 independent network metrics, of which 138 appear in more than 1% of the analyzed articles (Fig. 3). The metric frequency distribution exhibits a pronounced heavy-tailed distribution, with throughput (79.1%), end-to-end delay (74.6%), and packet error rate (59.5%) representing the most frequently studied metrics (Table 1). The core metric sets show strong scenario dependence (Fig. 4). For example, data center networks primarily emphasize throughput, end-to-end delay, and flow completion time, whereas Internet of Things (IoT) applications additionally prioritize energy consumption and network lifetime (Table 2). Furthermore, scenario-specific correlation matrices reveal markedly different coupling patterns among metrics (Figs. 5 and 6). In data center networks, throughput is strongly negatively correlated with flow completion time and queueing delay, reflecting the fundamental tradeoff associated with congestion control. In edge computing networks, end-to-end delay is negatively correlated with resource utilization, indicating the balance between real-time task offloading and resource utilization.  Conclusions  The strong coupling between network metrics and application scenarios indicates that future network architectures should be evaluated from a multidimensional perspective. Based on the extracted scenario-specific metric relationships, this study proposes the SMV analytical framework (Fig. 7). By jointly considering differentiated service quality requirements (Service), physical infrastructure cost and resource reuse (Multiplexing), and adaptive reconfiguration capability for emerging services and application scenarios (Versatility), the framework provides a theoretical basis for adaptive resource orchestration in AI-native networks. Future work will extend the current static literature mining pipeline into a continuously updated network metric knowledge base and further validate the engineering applicability of the SMV framework in programmable and multimodal networks.
A Lightweight Spatial-Spectral Dual-Branch Transformer Network for Classifying Polarized White Blood Cell Hyperspectral Images
YANG Yushi, YAN Jiaxuan, XIE Yi, QIU Lijia, HUANG Danfei
Available online  , doi: 10.11999/JEIT260124
Abstract:
  Objective  White Blood Cell (WBC) classification and morphology are essential in routine blood analysis and provide important information for disease diagnosis and health assessment. Current WBC classification mainly relies on hematology analyzers and manual microscopic examination. Hematology analyzers cannot acquire cellular images, limiting classification accuracy when abnormal cellular characteristics are present, whereas manual microscopic examination depends on operator experience and is susceptible to human error. Although automated WBC classification based on deep learning and computer vision has attracted considerable attention, methods using stained images are sensitive to staining conditions and image quality. In addition, conventional hyperspectral imaging has limited ability to distinguish WBC subtypes with highly similar morphological and spectral characteristics. To address these limitations, this study combines Polarized Hyperspectral Imaging (PHSI) with deep learning and proposes a lightweight classification framework for polarized hyperspectral WBC images, providing an efficient and reliable approach for clinical decision support.  Methods  A polarized hyperspectral microscopic imaging system is established to acquire images at multiple polarization angles. Based on Stokes Vector Theory, polarization parameters are calculated to generate multidimensional data cubes containing both light intensity and polarization-state information. Degree of Linear Polarization (DOLP) images are then computed to construct a PHSI dataset of WBCs. To exploit the multidimensional characteristics of PHSI, a Lightweight Spatial-Spectral Dual-Branch Transformer Network (LSDBT) is proposed. After preprocessing, the input data are fed into a dual-branch feature extraction module that separately extracts local spatial features and joint spatial-spectral features. An adaptive scaling factor is introduced to fuse the two feature streams and balance their contributions, enabling effective utilization of the multidimensional information contained in PHSI. A lightweight Swin Transformer backbone performs effective global feature modeling while reducing computational complexity. Global average pooling and a fully connected layer are used for classification. Model performance is evaluated using Overall Accuracy (OA), Precision, Recall, Specificity, and F1-Score. Ablation studies, comparative experiments, and feature visualization are conducted to validate the proposed method.  Results and Discussions  The DOLP spectra and PHSI visualizations of monocytes, lymphocytes, and neutrophils (Figures. 5 and 6) demonstrate different polarization characteristics that reflect the selective absorption and scattering of light by their internal structures. Compared with conventional intensity images, PHSI improves image contrast and provides additional polarization information that enhances discrimination among WBC types. The proposed LSDBT achieves an OA of 99.29% on the test set, with consistently high classification performance across all cell categories (Table 1). Analysis of the adaptive scaling factor shows that classification performance first improves and then decreases slightly as the scaling factor increases, with the optimal value of 3 providing the best balance between spatial and spatial-spectral features (Figure. 8). Ablation experiments (Tables 2 and 3) demonstrate that the dual-branch feature extraction module substantially improves classification performance, whereas the lightweight design greatly reduces computational complexity and model parameters with only a marginal reduction in accuracy. Compared with conventional hyperspectral imaging, the PHSI dataset achieves higher classification accuracy with all evaluated classifiers, indicating that polarization information provides complementary physical features that improve discrimination among WBC types (Table 4). Comparisons with representative methods show that LSDBT achieves the best overall classification performance across multiple evaluation metrics (Table 5). Furthermore, t-SNE visualization (Figure. 9) shows compact intra-class distributions and clear separation among different cell types, confirming the strong discriminative capability of the learned features. Although LSDBT does not have the lowest computational cost among the compared methods, it achieves the best balance between classification performance, model size, and computational efficiency (Table 6).  Conclusions  This study proposes a lightweight dual-branch Transformer network for polarized hyperspectral WBC classification. To the best of our knowledge, this is the first study to combine PHSI with deep learning for WBC classification. Comparative experiments with conventional hyperspectral imaging validate the superiority of PHSI for WBC classification. The proposed LSDBT integrates spatial and spatial-spectral information through a dual-branch feature extraction module and performs efficient global feature modeling using a lightweight Swin Transformer backbone. The network maintains high classification performance while substantially reducing computational complexity and model parameters. These results demonstrate that LSDBT provides an accurate and computationally efficient solution for automated WBC classification and supports the application of PHSI in cellular microscopic analysis and clinical auxiliary diagnosis.
Research on Adaptive Hybrid Beamforming Method for Massive MIMO LEO Satellite Communication Systems
XIAN Yongju, HUANG Xiaolong, XING Zhitong, LI Yun
Available online  , doi: 10.11999/JEIT260458
Abstract:
  Objective  With the growing demand for high-capacity and high-spectral-efficiency transmission in LEO satellite communications, massive MIMO has become a promising enabling technology. However, conventional fully digital beamforming is difficult to implement in practice due to the strict constraints on power consumption, hardware complexity, and payload cost of satellite platforms. Although partially connected hybrid beamforming can reduce hardware complexity, the conventional fixed subarray structure lacks flexibility and suffers from performance degradation, especially under low-resolution PSs constraints. To address this issue, this paper investigates adaptive antenna-RF chain mapping for hybrid beamforming design in LEO satellite massive MIMO systems.  Methods  This paper first establishes a system model for LEO satellite multi-user downlink massive MIMO hybrid beamforming and formulates a joint optimization problem with the objective of maximizing spectral efficiency. Considering the constant-modulus discrete phase constraints of low-resolution PSs, antenna-RF chain connection constraints, and transmit power constraint, the resulting problem is highly non-convex. To solve it, the original problem is transformed into an equivalent WMMSE formulation, and auxiliary variables are introduced to decouple the coupled variables. Based on this reformulation, a double-layer iterative optimization framework is developed by combining the BCD method and the PDD method. For adaptive antenna-RF chain mapping, the mapping problem is reformulated as a capacity-constrained linear assignment problem, and an optimal adaptive mapping method based on the Hungarian algorithm is proposed. Furthermore, to reduce the computational burden in large-scale antenna array scenarios, a low-complexity adaptive mapping method based on antenna priority sorting is developed.  Results and Discussions  Simulation results show that the proposed methods exhibit good convergence behavior. Specifically, the spectral efficiency increases rapidly in the initial iterations and then gradually converges, while the constraint violation decreases continuously, confirming the effectiveness of the proposed iterative optimization framework (Fig. 3). In terms of spectral efficiency, the proposed adaptive mapping methods consistently outperform the conventional fixed subarray and the existing greedy dynamic subarray scheme over different transmit powers and antenna scales. Among them, the Hungarian-based method achieving better spectral efficiency, whereas the antenna priority sorting-based method attains near-optimal performance with significantly reduced computational complexity (Figs. 4 and 5). As the number of PSs quantization bits increases, the system performance gradually approaches that of the continuous-PSs case, demonstrating the effectiveness of the proposed low-resolution PSs-based design (Fig. 6). In terms of energy efficiency, the proposed methods also outperform the conventional fully digital beamforming, fully connected, fixed subarray, and greedy dynamic subarray hybrid beamforming under different transmit powers and antenna scales (Figs. 7 and 8).  Conclusions  This paper proposes an adaptive hybrid beamforming design for LEO satellite massive MIMO systems under low-resolution PS constraints. By combining WMMSE reformulation with a PDD-BCD based optimization framework, joint design of digital precoding, analog precoding, and adaptive antenna-RF chain mapping is achieved. Simulation results demonstrate that the proposed methods provide superior performance in both spectral efficiency and energy efficiency. In particular, the Hungarian-based method provides better system performance, while the antenna priority sorting-based method achieves a favorable trade-off between performance and computational complexity. The proposed design provides an effective solution for high-performance hybrid beamforming in hardware-constrained LEO satellite massive MIMO systems.
Impact of Wireless Priors on the Computation and Energy Cost of MU-MIMO Precoding Learning
CONG Pengyu, HAN Shengqian, DENG Mingyu, LIU Shengjie, YANG Chenyang, SHEN Songhui
Available online  , doi: 10.11999/JEIT260388
Abstract:
  Objective  This paper studies the learning of downlink MU-MIMO (Multiuser Multi-Input Multi-Output) precoding policy from the perspective of computational complexity and energy consumption. Traditional numerical optimization achieves strong performance but incurs rapidly growing computational complexity with the number of base-station’s antennas and served users, leading to high inference latency and inference energy consumption. In recent years, deep learning has been widely used to reduce online computational cost, yet existing evaluations typically rely on training/inference time or FLOPs and lack direct energy/power measurements. More importantly, the computational cost of a deep model is tightly coupled with the network architecture, which should be designed to effectively exploit the prior knowledge of the precoding policy. The objective of this paper is thus to develop a network architecture that matches the multi-dimensional permutational properties of the policy, and to analyze how policy priors influence the complexity and energy through comprehensive hardware-based energy/power measurements and simulations.  Methods  We formulate the MU-MIMO precoding policy as a mapping from multi-user channel information to the optimal precoding matrix under a transmit power constraint. The optimal policy exhibits multi-dimensional joint permutation equivariance and invariance with respect to user indices, receive-antenna indices, and base-station antenna indices. To exploit this prior, we propose an Attention-based Graph Neural Network (AGNN) built on a hypergraph structure, whose update and aggregation procedures are designed to satisfy the required equivariance/invariance properties. An attention mechanism modeling inter-user interference is introduced to improve generalization to varying number of users. For broadband precoding, the input layer aggregates multi-subcarrier channel information into an expanded input representation. To quantify energy cost, we build a mixed CPU/GPU measurement framework that collects energy and power for GPU, CPU, and DRAM during training and inference. Simulations use 3GPP TR 38.901 UMa channel data with different antenna array sizes and bandwidth settings. We compare the proposed AGNN against numerical baselines (ZFBD+Greedy pairing) and two Transformer-based architectures, where only one-dimensional permutation properties are satisfied.  Results and Discussions  : The paper reveals two main findings. First, partial exploitation of policy priors leads to poor performance and high cost. In the MU-MISO scenario, the Transformer variants with one-dimensional permutation equivariance yield lower spectral efficiencies than the numerical baseline ZFBD+Greedy and much larger model sizes and inference FLOPs than AGNN. In contrast, AGNN, designed to match the multi-dimensional permutation properties, achieves higher spectral efficiency while reducing inference FLOPs by about one order of magnitude. Hardware measurements further show that AGNN reduces inference energy and power on both CPU and GPU. Second, in the MU-MIMO scenario with small- and large-scale settings, ZFBD+Greedy increases the system sum rate to 10.9 times, but raise inference FLOPs to 436.4 times, inference time to 20.8 times, and inference energy to 40.8 times. Conversely, AGNN increases the sum rate to 11.5 times while inference FLOPs rise to only 5.3 times; inference time and inference energy become dramatically smaller (0.03 times and 0.22 times, respectively). These results suggest that matching the precoding policy’s multi-dimensional permutation priors is an effective way to reduce the computational cost, including FLOPs, latency, and energy consumption, in large-scale 6G MU-MIMO systems.  Conclusions  This paper investigated how policy priors affect computational complexity and energy consumption in MU-MIMO precoding learning. By analyzing the multi-dimensional joint permutation equivariance/invariance properties of the optimal precoding policy, we designed an attention-based GNN (AGNN) that matches these properties. A hardware-aware measurement platform was built to obtain direct energy/power measurements for training and inference on CPU, GPU, and DRAM components. Simulations on 3GPP TR 38.901 channel datasets showed that Transformer architectures satisfying only one-dimensional permutation properties can lead to both inferior spectral efficiency and significantly higher computational/energy overhead. In contrast, AGNN achieves higher spectral efficiency while substantially reducing inference FLOPs, inference time, inference energy, and training complexity. As system size increases, traditional numerical methods incur rapidly growing energy costs relative to rate gains, whereas the proposed policy-prior learning method maintains low inference time and energy consumption. Overall, exploiting the prior knowledge of MU-MIMO precoding policy for network architecture design is a key enabler for energy- and complexity-efficient high-dimensional precoding optimization in future 6G networks.
Resource Allocation for Multi-UAV Relay Networks in 6G Semantic Communication
XIAO Liming, GUAN Zheng, LIU Jie, YU Jihong, CHEN Liyuan
Available online  , doi: 10.11999/JEIT260520
Abstract:
  Objective  Sixth-generation (6G) mobile networks aim to achieve global seamless coverage through space-air-ground integrated architectures. In this context, Unmanned Aerial Vehicles (UAVs) can serve as mobile aerial relay nodes to support massive ground user access in complex environments. However, traditional data-oriented communication paradigms incur high bandwidth overhead, making them difficult to apply under the limited spectrum resources of UAV networks. In addition, traditional centralized resource allocation methods struggle with real-time implementation due to highly dynamic network topologies and the complex coupling of multi-dimensional resources. To address these challenges, semantic communication has emerged as a new paradigm that extracts the essential meaning of information at the transmitter and reconstructs it at the receiver, thereby reducing redundant data traffic. Although semantic communication provides a feasible solution, existing semantic-driven resource allocation methods primarily focus on static ground networks or single-UAV scenarios, often overlooking the coverage limitations and co-channel interference in multi-UAV collaborative networks. Therefore, an intelligent joint resource optimization model and a distributed resource allocation framework are needed for multi-UAV relay networks. By jointly optimizing multi-dimensional resources, such a framework can improve semantic transmission efficiency and ensure long-term user fairness, thereby supporting intelligent resource scheduling in 6G integrated networks.  Methods  This study formulates a joint resource optimization model for multi-UAV relay networks under the semantic communication paradigm (Fig.1), coupling the number of semantic symbols, UAV trajectories, transmission power, and channel allocation. To evaluate semantic transmission quality, the proposed model introduces a Semantic Communication Quality of Service (SC-QoS) metric, which integrates Semantic Quantization Efficiency with normalized transmission delay. The optimization objective is formulated as a Mixed Integer Nonlinear Programming problem that maximizes the weighted sum of system-wide SC-QoS and long-term user fairness evaluated by Jain's fairness index. To solve this problem, a Two-Stage Hybrid Reinforcement Learning (TS-HRL) framework is proposed (Fig.2). The first stage employs a Capacity-Aware K-means (CA-K-means) algorithm for heuristic UAV pre-deployment. By introducing a dynamic distance compensation term based on residual capacity, this algorithm guides edge users toward low-load UAVs, achieving load balancing among UAV clusters while preserving spatial proximity. The second stage models the dynamic scheduling problem as a Decentralized Partially Observable Markov Decision Process and solves it using a Recurrent Independent Proximal Policy Optimization with Parameter Sharing (R-IPPO-PS) algorithm (Algorithm 1). This algorithm leverages Long Short-Term Memory (LSTM) networks to aggregate historical trajectories and observations, thereby capturing hidden environmental states. Furthermore, the multi-agent parameter-sharing mechanism reduces model complexity with respect to the agent scale, while the preference decoupling strategy converts discrete channel allocation decisions into continuous preference variables to facilitate model training.  Results and Discussions  The proposed TS-HRL framework is evaluated in a dynamic environment with randomly moving ground users. The convergence analysis shows that the proposed method achieves a higher initial reward and reaches a stable state with fewer iterations than the random deployment, memoryless allocation, and Enhanced Independent Soft Actor-Critic (EI-SAC) baselines (Fig.3). By leveraging LSTM-based historical information and CA-K-means topology initialization, the proposed method reduces invalid exploration and improves convergence stability. Compared with traditional bit-based communication, the semantic communication framework increases Semantic Spectral Efficiency (S-SE) by 3.6 times, reduces transmission delay by 92.1%, and improves fairness by 14.3% (Fig.4). Within the semantic communication paradigm, compared with the memoryless and heuristic schemes, the proposed method improves S-SE by 92.7% and 51.9%, reduces transmission delay by 13.5% and 57.7%, and improves fairness by 17.3% and 39.7%, respectively. Although the EI-SAC baseline achieves a fairness index of 0.93, its S-SE remains relatively low. In contrast, the proposed TS-HRL framework maintains a high fairness level while achieving an S-SE approximately 4.3 times higher than that of EI-SAC. Compared with random deployment, the proposed method improves S-SE by 11.3% with only a marginal 1.1% decrease in fairness, demonstrating a better trade-off between transmission efficiency and fairness. As the number of users increases from 10 to 40, most baselines show decreases in S-SE and fairness due to intensified co-channel interference and spectrum constraints (Fig.5). In contrast, the adaptive scheduling and global fairness reward mechanisms of the proposed method mitigate performance degradation and maintain the average transmission delay below 0.1 ms. These results indicate that the proposed method improves overall system performance while ensuring long-term user fairness.  Conclusions  This paper investigates joint resource allocation and trajectory planning in dynamic 6G multi-UAV relay networks by integrating the semantic communication paradigm. A joint optimization model coupling the number of semantic symbols, UAV trajectory, power control, and channel allocation is formulated to maximize system SC-QoS and long-term user fairness. To address the high-dimensional coupling of the formulated problem, a TS-HRL framework is proposed. This framework combines CA-K-means pre-deployment with the R-IPPO-PS algorithm to handle multi-dimensional resource scheduling under partial observability. Simulation results verify the convergence, stability, and scalability of the proposed method under different user densities. Through distributed cooperation among UAVs, the proposed method enhances semantic transmission efficiency and reduces transmission delay while ensuring long-term service fairness among ground users. These findings provide a viable approach for intelligent resource allocation in UAV-assisted semantic communication for future space-air-ground integrated networks.
LEO Satellite Multi-beam Multicast Precoding and User Grouping Joint Optimization Algorithm
GUO Lili, FENG Yimeng, YUAN Peihong, GAO Yue
Available online  , doi: 10.11999/JEIT260375
Abstract:
  Objective  In 6G LEO communication, multicast precoding is utilized to mitigate the significant inter-beam interference introduced by Full Frequency Reuse (FFR). However, traditional precoding algorithms are hindered by cubic computational complexity, making them unsuitable for massive MIMO. Additionally, existing user grouping strategies often fail to meet the fixed group size requirements of the DVB-S2X standard. Therefore, a joint optimization scheme is developed to integrate a low-complexity unsupervised deep learning model for precoding with improved user grouping algorithms to enhance sum rate and fairness.  Methods  To address the precoding challenge, an unsupervised deep learning model based on a hybrid CNN-LSTM architecture is proposed (Fig. 2). The Convolutional Neural Network (CNN) is utilized to extract spatial features from the Channel State Information (CSI), while the Long Short-Term Memory (LSTM) network captures their deep feature correlations. Unlike supervised learning, this model is trained by directly maximizing the sum rate as the loss function, subject to the Per-Antenna Constraint (PAC). For user grouping, two algorithms are developed to comply with the DVB-S2X standard. First, the CK-means algorithm is introduced, which modifies the standard K-means to ensure an equal number of users in each group. Second, FA-MAUG algorithm is proposed which prioritizes users with poor channel quality, thereby enhancing the overall robustness and fairness of the grouping result.  Results and Discussions  The intra-group similarity metric is employed to evaluate user grouping performance, which indicate that the CK-means algorithm achieves a similarity score approximately 0.1 higher than the MAUG algorithm and nearly 0.5 higher than random grouping across various group sizes (Fig. 3). This high similarity translates directly into better beamforming gain. In terms of throughput, the sum rate of the CNN-LSTM precoding combined with CK-means grouping outperforms traditional MMSE across different SNRs and total powers (Fig. 4, Fig. 5). The sum rate based on the proposed CNN-LSTM precoding scheme is on average 48.59% higher than that of the traditional MMSE algorithm and the gain of the CK-means algorithm compared to random grouping reaches 30.12% under different SNR conditions. Additionally, we analyzed the effects of the number of users per group, the number of groups, and the number of antennas on the system rate(Fig. 6, Fig. 7, Fig. 8), verifying the performance of the model in systems of different scales. Additionally, the complexity analysis confirms that the proposed precoding approach reduces the online overhead from the cubic level of traditional MMSE to a linear level relative to the number of antennas, which is critical for real-time deployment in massive MIMO satellite systems.  Conclusions  This paper presents a joint optimization framework for LEO satellite multicast systems, addressing the dual challenges of high computational complexity in precoding and lack of fairness in user grouping. Simulation results demonstrate that the proposed solution significantly enhances the system sum rate—by over 48.59% in typical SNR scenarios—compared to traditional methods. The approach provides a robust and scalable solution for the coordinated management of multi-beam interference in future 6G satellite communications. Future work will further explore the impact of hardware impairments and highly dynamic channel conditions on the generalization capabilities of the proposed model.
Fusing Global Perspective Rectification and Fine-Grained Semantic Decoupling for Language-Conditioned Robotic Grasping Detection
LIU Jin, LIU Zhitai, LI Zihan, SUN Yanjing, MIAO Yanzi, YUAN Xianfeng
Available online  , doi: 10.11999/JEIT260442
Abstract:
  Objective  Accurate grasping of target objects from language instructions is an essential prerequisite for service robots to achieve natural human-robot interaction. Existing research primarily relies on massive data-driven training or hierarchical feature fusion schemes for the cross-modal alignment between visual perception and textual instructions. However, these methods commonly overlook the heavy entanglement between target objects and background environments within low-level features, leading to a significant degradation in compositional generalization ability under cross-view and unseen background scenarios. To address these issues, this paper constructs a dual-view cross-scenario synchronous grasp detection and localization dataset, which is used to systematically evaluate and specifically improve the compositional generalization ability of existing models. Building upon this benchmark, we propose a simultaneous detection and localization network to simultaneously output the target location and the optimal grasping pose. Consequently, service robots equipped with the proposed network can perform object manipulation in real-world dynamic environments according to human instructions, providing theoretical foundations and technical support for embodied intelligence.  Methods  The structure of the proposed Simultaneous Grasp and Localization Network (SGL-Net) is illustrated (Fig. 1). Firstly, a cross-modal global context modulation module (CGCMM) is proposed. The module utilizes instruction priors to perform adaptive viewpoint correction and background suppression on visual channels during the early stages of feature extraction. Secondly, a word-pixel cross-modal alignment module (WPCAM) is designed to achieve fine-grained semantic decoupling via a flattened attention mechanism, further enhancing the model's comprehension in dynamic scenes. Finally, a unified decoder is employed to simultaneously output the target location and the optimal grasping pose.  Results and Discussions  Extensive quantitative and qualitative experiments are conducted on a restructured dual-view dataset, including bottom view and top view, and a real-world physical robotic platform. Comparative results demonstrate that SGL-Net significantly outperforms mainstream CNN-based and CLIP-based baseline methods in both grasp detection and object localization metrics (Table 2 and Table 3). Ablation studies explicitly validate the critical contributions of two designed modules in achieving fine-grained semantic alignment and scene decoupling (Table 4 and Table 5). Furthermore, qualitative visualizations (Fig. 4, Fig. 5, and Fig. 6) and real-world physical experiments (Fig. 7) indicate that SGL-Net exhibits strong viability for direct deployment in real-world physical environments. Overall, the network exhibits superior robust generalization capabilities and reliable practical deployment potential when confronting complex physical environments.  Conclusions  To address the challenge of insufficient cross-view and cross-domain generalization capability in robotic grasping operations, this paper constructs a corresponding validation dataset and proposes a simultaneous grasp detection and object localization network. By integrating a cross-modal global context modulation module and a word-pixel cross-modal alignment module, the proposed network achieves accurate localization of instruction-specified objects and prediction of optimal grasp poses. Experimental results from diverse testing datasets and real-world grasping experiments demonstrate the superior performance of the proposed network, highlighting its potential for deployment in practical scenarios. To further enhance robotic grasping capabilities in zero-shot scenarios, future work will focus on integrating Large Multimodal Models and realizing network adaptation through fine-tuning techniques.
Physics-Aware Reconstruction for Millimeter-Wave Radar Gait Recognition Under Complex Wearing Scenarios
HUANG Ling, QIU Liying, WANG Jiacheng, HAN Penglin, ZHOU Qingdi, YAN Huimei
Available online  , doi: 10.11999/JEIT260522
Abstract:
  Objective  Millimeter-wave (MMW) radar gait recognition has demonstrated considerable potential in the domain of non-contact biometric identification, primarily due to its inherent advantages in privacy preservation and resilience to variable lighting conditions. However, a significant challenge persists when applying these systems to unconstrained real-world environments where subjects may wear complex clothing, such as long coats or carry backpacks. These external covariates introduce non-stationary, high-frequency spectral components that result in severe spectral aliasing and masking of the intrinsic micro-Doppler (m-D) signatures of the human body. Traditional deep learning approaches usually treat radar spectrograms as generic image data and neglect the physical relationship between Doppler frequency shifts and human motion. Consequently, clothing-induced interference may overlap with motion-related signals, reducing recognition accuracy. Therefore, it is imperative to develop a framework that integrates radar physics and biomechanical properties to achieve effective signal decoupling and interference suppression. This study aims to provide a robust solution for radar-based gait recognition in complex scenarios by incorporating explicit physical constraints into the neural network architecture.  Methods  To mitigate the adverse effects of clothing-induced interference, this paper introduces PRISM-Net, a physics-aware reconstruction framework for millimeter-wave radar gait recognition. The framework is grounded in the biomechanical characteristics of human motion, specifically the observation that the human torso, representing the primary mass, generates stable and low-frequency Doppler shifts, whereas the movement of limbs produces higher-frequency, periodically alternating spectral components. (1) Physics-aware Frequency Structural Reconstruction: The proposed method diverges from conventional uniform processing by implementing a structural decoupling strategy. Based on the velocity distribution defined by biomechanical properties, the aliased original m-D spectrogram is partitioned into discrete frequency bands. The low-frequency band is dedicated to the torso, capturing stable and identity-persistent features, while the high-frequency band encompasses limb dynamics. This spatial-frequency partitioning enables the network to isolate the spectral regions most susceptible to clothing-induced clutter at the physical layer, thereby suppressing interference propagation. (2) Adaptive Weighted Attention Mechanism (WAM): An adaptive WAM module is integrated to regulate the signal-to-noise ratio (SNR) in the feature space. Given that clothing interference is dynamic and predominantly occupies high-frequency regions, the WAM adaptively evaluates the reliability of different frequency-derived features. In cases where the high-frequency limb features are corrupted by non-stationary noise from swinging garments, the WAM applies soft-thresholding to reduce their response weights. Simultaneously, the contribution of the more reliable torso-related features is enhanced. (3) Experimental Configuration and Protocol: The model is evaluated using the MMRGait-1.0 dataset, which provides a comprehensive set of radar gait signatures under various covariates. To ensure the assessment reflects real-world generalization, an open-set testing protocol is strictly followed. The training set comprises data from 74 subjects, while 47 entirely unseen subjects are used for evaluation. All radar spectrograms are processed to a standard resolution of 224 × 224 pixels. The network is optimized using the AdamW algorithm with a joint loss function consisting of cross-entropy and feature regularization components to ensure both classification accuracy and feature discriminability.  Results and Discussions  The experimental results demonstrate the effectiveness of integrating the biomechanical characteristics of human motion into the recognition process. As shown in Table 1, PRISM-Net achieves an average Rank-1 recognition rate of 89.7% in the 90-degree side-view scenario. In the coat (CT) occlusion scenario, the model maintains an accuracy of 85.1%. This is a 18.1% improvement over traditional lightweight models such as ShuffleNetV2 and a 7.4% improvement over the standard ResNet-18 baseline. The stability of the model was verified through ten independent repeated trials. As illustrated in the box plot in Fig. 1, PRISM-Net shows a low standard deviation of ±0.55%, whereas the baseline model exhibits a higher deviation of ±1.75%. An independent samples t-test results in a p-value of less than 0.001, confirming statistical significance. Ablation studies in Table 2 confirm the contribution of each module. Removing frequency decoupling reduces accuracy in the CT scenario to 75.5%, proving that physical structural separation is necessary to prevent feature distortion. The WAM module further provides a 4.2% accuracy gain by adaptively filtering noise. Regarding computational efficiency, Table 3 shows that PRISM-Net requires 11.33M parameters and 0.75 GFLOPs. Compared with heavy 3D-convolutional models that require over 10 GFLOPs, the proposed method achieves superior accuracy with significantly lower resource consumption. Finally, t-SNE visualization in Fig. 5 shows that PRISM-Net produces more compact intra-class distributions and clearer inter-class margins demonstrating improved feature discriminability.  Conclusions  This research demonstrates that the integration of the biomechanical characteristics of human motion into deep learning architectures improves the robustness of radar gait recognition. The proposed PRISM-Net framework effectively decouples motion components and suppresses clothing-induced noise through physics-aware frequency reconstruction and adaptive weighting. The high recognition accuracy and low computational complexity observed in the MMRGait-1.0 dataset suggest that this approach is suitable for real-time security applications on edge-computing devices.
FedFACO: Personalized Federated Learning Method Based on Fisher Information Matrix for Adaptive Aggregation and Client Collaborative Optimization
JIANG Wei-Jin, LIU Zhi-Hua, CUI Xin-Yu, XU Yu-Sheng, CHEN Shen-You, HU Jia-Long
Available online  , doi: 10.11999/JEIT260344
Abstract:
  Objective  Non-IID data heterogeneity remains one of the major challenges in personalized federated learning, as it often leads to inconsistent local optimization directions, insufficient global knowledge transfer, and degraded model personalization performance. To address these issues, this paper proposes FedFACO, a personalized federated learning method based on Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The proposed method aims to improve the adaptability of federated models to heterogeneous client distributions while maintaining effective knowledge sharing across clients. By introducing an adaptive aggregation mechanism and a collaborative optimization strategy, FedFACO provides a more principled way to balance global generalization and local personalization, which is particularly important in complex Non-IID federated environments.  Methods  FedFACO consists of two key components. First, an adaptive aggregation (AA) mechanism is employed to dynamically adjust fusion weights between global and local models based on client-specific states, generating personalized initializations aligned with local data. Second, a collaborative optimization (CO) mechanism is introduced, combining feature alignment with FIM-based client weighting. This enhances useful global knowledge transfer and suppresses low-quality updates. The FIM is utilized to quantify the information contribution of each update, ensuring reliable aggregation. The method is evaluated on MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet under Non-IID settings, and compared with representative baselines. Convergence behavior, dropout robustness, and sensitivity to low-quality updates are also examined.  Results and Discussions  Experimental results demonstrate that FedFACO consistently outperforms competitive baseline methods across all four benchmark datasets, achieving an average accuracy improvement of approximately 3.1% over mainstream approaches (Fig. 1, Table 2). On the more challenging Tiny-ImageNet dataset, FedFACO reduces the total training time required to reach convergence by approximately 4.8% compared with the best baseline (Table 3). Ablation studies confirm that performance is substantially improved by both AA and CO mechanisms, with their joint application yielding optimal accuracy (Table 6). Furthermore, FIM-guided weighting is shown to accurately quantify contribution quality in client dropout and asynchronous scenarios, significantly enhancing aggregation reliability (Fig. 2Fig. 3). Superior robustness is also demonstrated in malicious client scenarios (Fig. 4).  Conclusions  This paper presents FedFACO, a personalized federated learning method for Non-IID environments via Fisher information matrix-guided adaptive aggregation and client collaborative optimization. The method effectively balances global knowledge sharing and local personalization while enhancing training stability and robustness under heterogeneous client participation. Experimental results validate its effectiveness and superiority in accuracy, convergence efficiency, and robustness. Future work will focus on lightweight Fisher information approximations and adaptive triggering strategies to reduce computational overhead, as well as integration with privacy-preserving and security-defense mechanisms for deployment in resource-constrained and high-security environments.
Channel Estimation for MIMO-OFDM Based on Adaptive Transformer Network in High-Speed Mobile Scenarios
LIAO Xi, HE Xiangni, ZHANG Zhe, WANG Yang
Available online  , doi: 10.11999/JEIT260075
Abstract:
  Objective  Accurate channel state information (CSI) is essential for coherent detection, beamforming, and adaptive resource allocation in multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. In high-mobility environments, large Doppler shifts and multipath propagation jointly cause doubly selective fading, destroy subcarrier orthogonality, and aggravate inter-carrier interference. Consequently, conventional least squares (LS) and linear minimum mean square error (LMMSE) estimators suffer substantial performance degradation. Existing deep learning estimators provide strong nonlinear modeling capability, but often show insufficient adaptability to variations in signal-to-noise ratio (SNR), delay spread, and Doppler shift, or inefficiently incorporate physical channel priors. To address these problems, a Transformer network with adaptive feature modulation, termed AdaFiT, is proposed for channel estimation in high-mobility MIMO-OFDM systems.  Methods  The proposed AdaFiT framework performs channel estimation by jointly exploiting local time-frequency correlations, global dependencies, and explicit channel-aware adaptation. It takes least squares (LS) estimates at pilot positions, together with signal-to-noise ratio (SNR), delay spread, and maximum Doppler shift, as inputs. A separable two-dimensional linear upsampling module first interpolates the sparse pilot estimates to the full OFDM time-frequency grid by independently processing the real and imaginary components along the frequency and time dimensions. A convolutional feature enhancement module then extracts robust local representations: a complex feature mixing layer fuses multi-antenna real and imaginary components, while multi-scale convolutional blocks with channel attention capture short-range time-frequency correlations and suppress noise. Next, a feature-wise linear modulation (FiLM)-based channel-adaptive module embeds the three channel parameters through independent multilayer perceptrons and combines them into a channel-state representation. This representation generates scaling and shifting coefficients to dynamically recalibrate the block-embedded sequences according to varying channel statistics. Finally, the modulated sequences are processed by a Transformer encoder with learnable two-dimensional positional encoding to model long-range dependencies across the time-frequency grid and antenna dimensions. A residual reconstruction structure fuses the global features with locally enhanced representations, yielding accurate channel estimates with global consistency and preserved local details.  Results and Discussions  Simulation analyses are conducted based on the CDL-C and CDL-A channel models. The LS-based bilinear interpolation method, the LMMSE method, the AdaFortiTran model, and the AdaFiT model without the adaptive module are selected as comparison schemes. The mean squared error (MSE) performance of these models is compared under different signal-to-noise ratio (SNR), maximum Doppler shift, and delay spread conditions. Under the CDL-C channel model, AdaFiT achieves the lowest MSE over the entire SNR range (Fig. 3). At an SNR of 0 dB, its MSE is approximately 2.5 dB lower than that of AdaFortiTran, and the performance gain increases to approximately 7 dB at an SNR of 30 dB. Compared with the AdaFiT model without the adaptive module, AdaFiT achieves a maximum MSE gain of approximately 2.1 dB, which verifies the effectiveness of the proposed channel-adaptive feature modulation module. When the maximum Doppler shift increases from 200 Hz to 1400 Hz, AdaFiT consistently maintains the lowest MSE (Fig. 4). In the high-Doppler range of 10001400 Hz, AdaFiT outperforms the AdaFiT model without the adaptive module and AdaFortiTran by approximately 2 dB and 3.5 dB, respectively, demonstrating improved robustness against rapid channel variations. For delay spreads ranging from 100 ns to 700 ns, AdaFiT also maintains the lowest MSE over the entire range (Fig. 5), achieving approximately 3 dB and 5.5 dB gains over the AdaFiT model without the adaptive module and AdaFortiTran, respectively. These results demonstrate that the proposed adaptive feature modulation mechanism effectively improves the adaptability of the model to time-selective and frequency-selective fading.To further evaluate the performance of the AdaFiT model under different channel models, simulation analysis is conducted based on the 3GPP CDL-A channel model. At SNRs of 0–5 dB, AdaFiT outperforms LMMSE by approximately 4–5 dB, while the performance gain reaches approximately 7 dB at SNRs of 20–30 dB (Fig. 6(a)). Compared with AdaFortiTran, AdaFiT achieves an MSE gain of approximately 2 dB under low-SNR conditions, and the gain increases to approximately 5.5 dB under high-SNR conditions. In the maximum Doppler shift experiment, the MSE of AdaFiT remains between approximately –38 dB and –37 dB over the range of 200–800 Hz and is approximately 5 dB lower than that of AdaFortiTran (Fig. 6(b)). Although the MSE of AdaFiT increases when the maximum Doppler shift exceeds 1000 Hz, it still achieves an approximately 6 dB gain over LMMSE at 1400 Hz and remains superior to AdaFortiTran. Over the entire delay spread range, AdaFiT achieves approximately 4–6 dB gain over LMMSE and approximately 4–5.5 dB gain over AdaFortiTran (Fig. 6(c)).  Conclusions  The proposed AdaFiT framework performs channel estimation by jointly exploiting local time-frequency correlations, global dependencies, and an explicit channel-aware adaptation mechanism. Simulation results under the CDL-C and CDL-A channel models show that AdaFiT consistently achieves lower MSE than the LS-based bilinear interpolation method, LMMSE, AdaFortiTran, and the AdaFiT model without the adaptive module under different SNR, maximum Doppler shift, and delay spread conditions. These results verify the effectiveness of the proposed adaptive feature modulation mechanism and demonstrate that AdaFiT maintains high estimation accuracy and stable performance under different channel models, indicating good adaptability to varying channel environments.
Energy Efficiency Analysis of Discrete Phase-Shifted Active RIS Enhanced Communication Systems
SHU Feng, LIN Zhiyuan, ZHENG Weihai, WANG Yan, JIANG Hao, WANG Jiangzhou
Available online  , doi: 10.11999/JEIT260462
Abstract:
  Objective  Active Reconfigurable Intelligent Surface (RIS) can effectively overcome the multiplicative fading problem of passive RIS by integrating radio frequency amplifiers, significantly improving the performance of wireless communication systems. However, it also introduces amplification noise and increases power consumption. In addition, the high-precision digital phase control performed by base stations on RIS poses a challenge due to excessive communication overhead. Adopting low-precision phase shifters is a key approach to promoting the practical application of RIS. Therefore, deriving indicators that affect the energy efficiency (EE) performance of active RIS-assisted systems, as well as analyzing the impact of finite bit quantization errors on EE, have important guiding significance for future RIS system design and practical deployment. In view of this, a discrete active RIS-aided system model is proposed under Rayleigh fading channels, and the EE loss caused by phase quantization error is analyzed. The approximate optimal solution of the power allocation factor and the number of RIS elements that can achieve maximum EE is derived, and the relationship between EE at RIS and user is revealed, providing a theoretical basis for the practical engineering deployment of active RIS.  Methods  Based on the weak law of large numbers and Taylor series expansion, closed-form expressions for EE performance loss at user and corresponding approximate performance loss expression are derived. The impacts of system parameters on EE are investigated by formulating univariate functions. Combined with the Ferrari’s method and the Lambert W function, the approximate optimal solutions of the power allocation factor and the number of RIS elements for EE maximization are derived, respectively. By applying the weak law of large numbers and the Lambert W function, the relationship between EE at RIS and EE at user is revealed.  Results and Discussions  The expression of EE at user is derived to be a function of six factors: the number of quantization bit (\begin{document}$ k $\end{document}), power allocation factor (\begin{document}$ \beta $\end{document}), the number of RIS elements (\begin{document}$ N $\end{document}), the total power sum of base station and active RIS (\begin{document}$ {P}_{\text{t}} $\end{document}), the noise at active RIS (\begin{document}$ \sigma _{\text{r}}^{2} $\end{document}), and the noise at user (\begin{document}$ \sigma _{\text{u}}^{2} $\end{document}). Firstly, the EE performance loss decreases as \begin{document}$ k $\end{document} increases. When \begin{document}$ k $\end{document}=3, the gap between the approximate performance loss and the without performance loss is less than 0.0268 Mbit/J, and the performance loss and the without performance loss is less than 0.0265 Mbit/J (Fig. 3). Therefore, discrete phase shifters with 3 to 4 bits can achieve extremely low performance loss.EE at user exhibits a unimodal characteristic with respect to both \begin{document}$ \beta $\end{document}and \begin{document}$ N $\end{document}. The approximate optimal solution of \begin{document}$ \beta $\end{document} for EE maximization has an error less than 0.01 compared with the exact optimal solution obtained by the Dinkelbach algorithm (Fig. 4), and there is no error between the approximate optimal solution and the exact optimal solution of \begin{document}$ N $\end{document} (Fig. 5), demonstrating the effectiveness of the two approximate models.EE at user first increases and then decreases with the increase of \begin{document}$ {P}_{\text{t}} $\end{document}showing a unimodal variation. The higher the quantization precision, the larger the EE peak value, and the smaller the optimal \begin{document}$ {P}_{\text{t}} $\end{document} required to achieve the EE peak (Fig. 6). EE decreases with the increase of both \begin{document}$ \sigma _{\text{r}}^{2} $\end{document} and \begin{document}$ \sigma _{\text{u}}^{2} $\end{document}. Moreover, EE at user is more susceptible to the amplification noise at RIS, and reducing the noise at RIS can achieve a more significant EE gain (Fig. 7). EE at user first increases and then drops sharply to zero with the increase of EE at active RIS, and the SNR at active RIS is the key factor determining the relationship between these two EE performances (Fig. 8).  Conclusions  The EE performance of active RIS assisted wireless networks with discrete phase shifters in Rayleigh fading channels is studied. Firstly, the expressions for the loss, no loss, and approximate loss of EE at active RIS and user are derived. Simulation results show that 3 to 4-bit discrete phase shifters can fully exploit the gain of RIS. Secondly, functions for each parameter of EE at user are constructed. Based on the Ferrari method and the Lambert W function, the approximate optimal solution of the power allocation factor and the number of RIS elements that can achieve maximum EE is derived. Simulation results showed that the error between the approximate optimal solution and the exact solution is minimal. Finally, the relationship between EE at RIS and EE at user is revealed, and the results show that EE at user first increases and then decreases to zero with the increases of EE at RIS.
A High-Parallelism Simulated Adiabatic Bifurcation Processor for Combinatorial Optimization Problems
LI Renlong, HAO Xin, CHEN Zhuojun, DING Ding
Available online  , doi: 10.11999/JEIT260780
Abstract:
  Objective  Combinatorial optimization problems (COPs) are widely encountered in fields such as network optimization, autonomous driving path planning, VLSI design, and computational biology. The solution space of these problems grows exponentially with problem size, making it impossible for traditional von Neumann architectures (e.g., CPUs) to find high-quality solutions within a feasible time. Quantum-inspired Ising machines have emerged as a promising computing paradigm for accelerating COP solving. However, existing CMOS-based Ising machines still face significant challenges in simultaneously achieving high speed, high energy efficiency, and high accuracy. Discrete-time Ising machines often require thousands of iterations and rely heavily on random number generators, leading to large chip area and long solving times. Continuous-time Ising machines suffer from poor solution quality, frequently getting trapped in local minima. To address these limitations, this paper designs and fabricates a high-parallelism simulated adiabatic bifurcation application-specific processor for combinatorial optimization problems in 65 nm CMOS technology.  Methods  The proposed processor adopts the simulated adiabatic bifurcation (SAB) algorithm, which is inspired by quantum adiabatic optimization. Unlike simulated annealing, SAB does not require Gibbs sampling or random number generators to escape local minima. The algorithm models each spin as a nonlinear oscillator and solves a set of ordinary differential equations to simulate the adiabatic evolution of a classical nonlinear Hamiltonian system exhibiting bifurcation phenomena. The processor builds a hardware architecture that supports fully connected spin topologies and leverages the inherent fully parallel spin update characteristic of the SAB algorithm. To achieve bubble-free iterative computation, a three-stage pipelined spin update strategy is proposed, dividing the update process into coupling coefficient access, momentum update, and position update. To reduce the storage overhead introduced by the fully connected coupling matrix, a folded coupling coefficient storage array is designed, exploiting matrix symmetry to eliminate redundant storage. The momentum update unit and position update unit are implemented using 8-bit fixed-point arithmetic (2 bits for integer, 6 bits for fractional part) to achieve SAB evolution with low hardware overhead. The chip is fabricated in 65 nm CMOS technology, occupying an area of 0.47 × 1.25 mm2 and operating at a 200 MHz clock frequency and 1 V supply voltage.  Results and Discussions  The chip achieves a total power consumption of only 14 mW, with the spin evolution module consuming 58% of the total power. The folded coupling coefficient storage array reduces storage area by 52% compared to full matrix storage, and the rectangular restructuring avoids irregular shapes in physical layout, reducing routing congestion and layout voids (Fig.6). The parallel loading access mechanism allows all coupling coefficients to be read and distributed within a single clock cycle, eliminating the memory access bottleneck inherent in serial reading. For predefined Max-Cut problems configured as 8×8 grid structures (64 nodes) with coupling coefficients quantized to 2-bit precision, the chip converges to the global optimum within only 5 computing cycles, achieving a final Ising energy of –4032 (Fig.9). This energy follows the analytical expression (n4−n2)(n4−n2), confirming that the chip solves Max-Cut problems with the shortest solving time. For larger extended Max-Cut problems (108×108 nodes), the simulated adiabatic bifurcation algorithm achieves 100% accuracy relative to the theoretical ground state (Fig.10). Monte Carlo simulations over 1,000 independent trials on randomly generated Max-Cut problems demonstrate that simulated adiabatic bifurcation achieves an average Hamiltonian of –3617.42, significantly outperforming simulated annealing which achieves only –3352.12 (Fig.11). For 3-SAT problems with 40 clauses and 8 variables, the chip solves instances with clause-to-variable ratios of 3, 4, and 5 in approximately 6 μs, 10 μs, and 16 μs, respectively (Fig.12). These results align perfectly with theoretical phase transition predictions, confirming the processor's effectiveness across varying problem complexities.  Conclusions  This work proposes an simulated adiabatic bifurcation machine that enables fully parallel updates without duplicating spin copies. To improve throughput, a three-stage pipeline strategy is designed that integrates coupling-coefficient access, momentum update, and position update, achieving bubble-free parallel updating and low-latency solving. For sparse coupling coefficients, a folded storage scheme is adopted to significantly reduce memory area overhead. Both momentum and position variables are represented in 8-bit fixed-point format, ensuring sufficient computational accuracy while balancing resource efficiency. Compared with previous fully connected Ising machines, the proposed bifurcation machine achieves 100% solving accuracy, along with higher energy efficiency and lower hardware overhead, demonstrating substantial application prospects in edge-side combinatorial optimization.
A State Prediction Method for Long-Endurance Fixed-Wing UAV Propulsion Systems
LI Sicheng, WANG Lianqing, LI Zhiyong, WANG Guochang, GE Kaihua, CHEN Junfeng, TAN Rongqing
Available online  , doi: 10.11999/JEIT260188
Abstract:
  Objective  Accurate single-step prediction of key propulsion-system states is essential for early fault warning and autonomous health management of long-endurance fixed-wing unmanned aerial vehicles (UAVs). During high-altitude missions lasting more than 24 h, electrical and thermal variables in the propulsion system exhibit strong coupling, multi-time-constant dynamics, and pronounced day-night regime shifts. These characteristics cause short-term disturbances and long-term drifts to coexist, and hinder adaptive feature weighting under time-varying variable sensitivities. General-purpose time-series predictors may therefore fail to meet the accuracy and robustness requirements of multivariate propulsion-state prediction. To address these challenges, a Grouped Squeeze-and-Excitation Multi-scale Temporal Convolutional Network (GEMS-TCN) is developed by enhancing a modern pure-convolution forecasting backbone with multi-scale embedding and grouped channel attention. The aim is to obtain accurate 10 s-ahead single-step forecasts for 16 key propulsion states from 72-dimensional flight telemetry while satisfying the real-time inference requirement of the 1 Hz telemetry cycle.  Methods  Real flight telemetry from a representative long-endurance fixed-wing UAV is used for model construction and evaluation. The data are sampled at 1 Hz for 9 consecutive days, yielding 806,629 time steps and 72 variables. (Fig.2) Sixteen propulsion-related key states, including the bus voltage, control-unit temperature, winding temperature, and power-device temperature of four motors, are selected as prediction targets, and all 72 variables are used as inputs. (Table 1) The raw data are processed through time indexing, interquartile range (IQR)-based anomaly handling, and interpolation, and are then chronologically divided into training, validation, and test sets at a ratio of 8:1:1 to avoid information leakage. (Fig.4) (Fig.5) GEMS-TCN uses a multi-scale embedding layer with parallel one-dimensional convolutions to extract temporal patterns at different receptive-field scales. Stacked GEMS-TCN blocks combine depthwise temporal convolution, grouped convolutional feed-forward networks, and Grouped Squeeze-and-Excitation (GroupSE) modules to recalibrate intra-variable and cross-variable channel responses hierarchically. (Fig.1) The models are trained with Adam and mean squared error (MSE) loss, and are evaluated using mean absolute error (MAE), MSE, and symmetric mean absolute percentage error (SMAPE). PatchTST, ModernTCN, FEDformer, DLinear, and TimesNet are used as comparison models, with ablation and robustness experiments conducted for further verification.  Results and Discussions  On the full 16-dimensional target set, GEMS-TCN achieves a test-set MAE of 0.171 and an MSE of 0.069. (Table 4) Compared with TimesNet, the strongest baseline in overall trend tracking, GEMS-TCN reduces MSE by 28.1% while maintaining comparable MAE and SMAPE, indicating stronger suppression of large prediction deviations. (Table 4) Stable accuracy is obtained in both daytime and nighttime segments, with MAE/MSE values of 0.189/0.085 and 0.150/0.051, respectively, demonstrating robustness to diurnal operating-condition changes. (Table 4) The prediction trajectories show tighter alignment at thrust-transition points, reduced overshoot, and fewer spurious spikes, while low-bias tracking is preserved for slowly varying nighttime temperature profiles. (Fig.7) Ablation results show that multi-scale embedding and GroupSE provide complementary improvements, and their combination achieves the best overall performance among the ablation settings. (Table 6) Under 1% and 5% random dropouts and 60 s continuous missing intervals, the MSE remains below 0.07, indicating tolerance to practical data-loss scenarios. (Table 5) In addition, GEMS-TCN contains 27.88 M parameters and achieves an inference latency of 0.90 ms per sample, which is well below the 1 s sampling interval.  Conclusions  GEMS-TCN provides a practical convolution-based solution for multivariate propulsion-state prediction in long-endurance fixed-wing UAVs. By integrating multi-scale temporal embedding with hierarchical group-wise channel recalibration, the proposed method jointly represents rapid fluctuations and slow evolution, and better captures multi-time-constant dynamics and multivariate coupling in propulsion telemetry. Real-flight experiments demonstrate stable prediction performance across diurnal regimes, state categories, and data-missing scenarios. Ablation results further confirm the effectiveness and complementarity of multi-scale embedding and GroupSE. These findings indicate that structure-aware modeling tailored to propulsion-state evolution can support health monitoring, early fault warning, and autonomous health management of long-endurance fixed-wing UAVs.
Customized Preparation of Single Crystal Tungsten Tips and Surface Reconstruction Mechanism
GUO Jiamei, YIN Shengyi, ZHANG Yongqing, SUN Wanzhong
Available online  , doi: 10.11999/JEIT260328
Abstract:
  Objective  Refractory metal tungsten, particularly single crystal tungsten, serves as a critical material for high-performance field emission cathodes, which are core components in advanced electron microscopes and electron beam lithography systems. The fabrication of single crystal tungsten microtips with precisely controlled geometry and surface cleanliness remains a significant challenge, especially for applications requiring tip radii ranging from sub-100nm for cold field emission to 0.3–1.0 μm for Schottky-type thermal field emission. Currently, the domestic development of high-end electron microscopes in China faces a major bottleneck due to heavy reliance on imported field emission cathodes. Electrochemical corrosion, the mainstream method for tip preparation, suffers from limitations such as difficulty in cutoff timing control, susceptibility to tip bending or passivation, residual surface impurities, and low yield rates. Moreover, the anisotropic nature of single crystal tungsten introduces additional complexity in morphology control. This study aims to establish a controllable fabrication method for single crystal tungsten tips, enabling tailored geometry and surface quality to meet demanding requirements of different field emission applications while significantly improving fabrication yield.  Methods  Single crystal tungsten wires (diameter 0.12 mm, purity 99.95%, (100) orientation) were used as the starting material. The tips were first pre-shaped by electrochemical corrosion in 1 mol/L NaOH solution under a 10 V DC voltage with pulsed control (6 kHz, 100 μs pulse width) for 10 min. Following corrosion, the tips underwent surface cleaning by sequential immersion in ultrapure water and anhydrous ethanol, followed by drying with high-purity nitrogen. Subsequently, the samples were placed in an ultra-high vacuum system (base pressure < 1 × 10–6 Pa) and subjected to high-temperature oxygen treatment. Process parameters were systematically varied, including temperature (15001600 °C), oxygen flow rate (0.015–1.0 sccm), and treatment duration (5–1440 min). During treatment, high-purity oxygen was introduced while maintaining vacuum levels better than 10–4 Pa. Tip morphologies were characterized by scanning electron microscopy, and surface compositions were analyzed by energy dispersive X-ray spectroscopy. Geometric parameters, including half-cone angle and tip radius, were measured following standardized protocols.  Results and Discussions  The combination of electrochemical corrosion and high-temperature oxygen treatment enabled both effective surface purification and controlled morphological reconstruction. Scanning electron microscopy characterization revealed distinct evolutionary pathways depending on oxygen partial pressure (Fig. 3, Fig. 4). Under oxygen-rich conditions (1.0 sccm, 15001600 °C), the tips underwent “blunting reconstruction,” evolving from an initial inverted cone into a characteristic “cylindrical segment + hemispherical cap” structure (Samples 2–5, Table 2). For Sample 3 treated at 1500 °C for 10 min, the tip radius increased to 327 nm with a half-cone angle reduction of 21%; for Sample 4 treated for 15 min, further evolution occurred with tip radius decreasing to 275 nm and cylindrical segment height increasing substantially. With increasing temperature from 1500 °C (Sample 2) to 1600 °C (Sample 5), the half-cone angle change rate increased from 11% to 25%, while tip radius decreased from 380 nm to 172 nm, indicating accelerated kinetics. This anisotropic behavior is attributed to orientation-dependent surface energies of tungsten and preferential oxidation along specific crystallographic planes. Under oxygen-lean conditions (0.015 sccm, 1500 °C, 1440 min), “sharpening reconstruction” occurred, with tip radius decreasing dramatically from 143 nm to 28 nm and half-cone angle reducing from 8° to 1° (Sample 6). In contrast, treatment in vacuum at 1500 °C primarily removed surface contaminants without altering tip geometry (Sample 1). EDS analysis confirmed that treated surfaces were free from impurities other than trace carbon contamination (Table 3), demonstrating the dual functionality of the process in achieving both purification and reconstruction. The findings reveal that oxygen partial pressure serves as the key determinant of reconstruction direction, with oxygen-rich conditions favoring blunting and oxygen-lean conditions promoting sharpening.  Conclusions  A combined process of electrochemical corrosion followed by high-temperature oxygen treatment was successfully developed for the controllable fabrication of single crystal tungsten field emitter tips. The process achieves dual functionality: effective surface purification through high-temperature vacuum treatment, which removes residual surface contaminants from electrochemical corrosion, and controlled morphological reconstruction enabled by the introduction of oxygen under precisely regulated conditions. By systematically adjusting process parameters including treatment temperature, oxygen flow rate, and treatment duration, tip morphology and dimensions can be tailored to meet specific application requirements. The oxygen partial pressure plays a decisive role in determining the reconstruction pathway. Under oxygen-rich conditions (e.g., 1500 °C, 1.0 sccm), blunting reconstruction yields a “cylindrical segment + hemispherical cap” structure ideally suited for Schottky-type thermal field emission cathodes requiring tip radii between 0.3 μm and 1.0 μm. Under oxygen-lean conditions (e.g., 1500 °C, 0.015 sccm), sharpening reconstruction produces nanoscale sharp tips with tip radius as low as 28 nm, suitable for cold field emission cathodes and scanning tunneling microscope probes. This anisotropic reconstruction mechanism is explained by the selective reaction of oxygen atoms with different tungsten crystal planes under high-temperature conditions, where the formation of volatile oxides modifies local surface energy distribution and promotes the exposure of specific crystallographic planes. Control experiments confirmed that such morphological reconstruction does not occur under oxygen-lean or oxygen-free conditions at the same temperatures, further demonstrating the critical role of oxygen in triggering this process. Importantly, this approach provides an effective method to correct imperfections from the initial electrochemical corrosion step, significantly improving fabrication yield and process robustness, thereby offering a viable technical pathway for the domestic fabrication of high-performance field emission cathodes.
A Spatial-Temporal Collaborative Optimization Method for Stable Grab Trajectory Extraction
CHEN Xiaoyu, ZHANG Fengzhuo, CHEN Yang, LIU Wenyuan, KONG Deming
Available online  , doi: 10.11999/JEIT260512
Abstract:
  Objective  In port operation videos, the grab is a continuously moving target, and accurate trajectory extraction is of great importance for operation monitoring, equipment coordination, and collision warning. However, complex backgrounds, scale variations, partial occlusion, and boundary degradation often make target region extraction unstable, leading to centroid deviation, trajectory jitter, missed detections, and trajectory interruption. To address these issues, this paper proposes a spatial-temporal collaborative optimization method for stable and continuous grab trajectory extraction. While maintaining a relatively high inference speed, the proposed method effectively improves both the accuracy and continuity of trajectory extraction, providing a practical solution for stable perception of continuously moving targets in port industrial video scenarios.  Methods  Built on YOLOv8-seg, the proposed framework integrates spatial representation enhancement and temporal consistency constraint. First, CBAM, BiFPN-lite, and shallow feature aggregation are introduced to improve target-background separability, strengthen multi-scale representation, and preserve boundary details. Then, a temporal consistency constraint is imposed on prototype features through global average pooling, a cache-based pairing mechanism, and a weighted Charbonnier loss, thereby suppressing the temporal accumulation of local errors. In addition, a stage-wise training strategy with warmup epochs and a unified loss function is adopted to ensure stable convergence.  Results and Discussions  Experiments are conducted on DAVIS2016, SegTrackV2, and a real portal crane grab dataset to evaluate the proposed method in terms of segmentation performance, trajectory stability, and occlusion robustness. The results show that the proposed method achieves the best segmentation performance on the grab dataset, with J and F scores of 90.05% and 98.56%, respectively. It also shows clear improvement on DAVIS2016 and maintains comparable performance with slight gains on SegTrackV2 (Tables 1 and 2, Fig. 2). In terms of trajectory stability, compared with YOLOv8-seg, the proposed method reduces MAE and RMSE by approximately 55.3% and 52.6%, respectively, while lowering the miss rate to 0.56% (Table 5). It also produces a more concentrated trajectory error distribution and a smaller fluctuation range (Fig. 3). Occlusion robustness experiments further show that, under different occlusion ratios, the proposed method maintains good region integrity and continuous extraction capability, reducing the maximum number of consecutive missed frames from 52 to 47 (Table 6, Figs. 4 and 5). Ablation studies verify the complementarity of spatial representation enhancement and the temporal consistency constraint, while parameter analysis shows that setting the TCONS weight to 0.3 provides the best balance between segmentation quality and trajectory stability (Tables 7 and 8).  Conclusions  This paper proposes a spatial-temporal collaborative optimization method to address the challenge of stable grab trajectory extraction in port operation videos. Experimental results demonstrate that the proposed method achieves favorable segmentation accuracy and trajectory stability on DAVIS2016 and the real grab dataset, maintains comparable segmentation performance on SegTrackV2, shows strong continuous extraction capability under occlusion, and incurs no significant loss in inference speed. Since the current study is limited to fixed crane viewpoints, future work will focus on cross-scene generalization and long-term continuous perception under more complex operating conditions to further enhance the robustness and applicability of the proposed method in real-world environments.
Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding
ZHOU Tong, CHEN Hongzhi, XU Jialong, SUN Peng, JIANG Dajie, LIU Jiankang
Available online  , doi: 10.11999/JEIT260379
Abstract:
Currently, the 3rd Generation Partnership Project (3GPP) Technical Specification Group Service and System Aspects Working Group 4 (TSG SA Working Group 4, SA4) has begun researching Artificial Intelligence (AI)-based audio codecs. Based on the Descript Audio Codec (DAC) being studied by SA4, this paper presents a joint optimization scheme for DAC source-channel coding and modulation with total power constraints, a joint optimization scheme for DAC source-channel coding and modulation with constant modulus constraints, and a DAC joint source-channel coding scheme. Furthermore, simulation evaluation is conducted using the typical physical layer parameter configuration of 3GPP Geostationary Earth Orbit (GEO) voice scenario. Finally, preliminary research suggestions are made, with the hope of providing useful insights for the subsequent research on GEO voice.  Objective  In recent years, Joint Source-Channel Coding (JSCC) has gained widespread attention in academia and industry. In academia, the sources studied for JSCC include video, image, and audio. However, the industry has not focused on these sources as in academia, but rather on JSCC where channel state information (CSI) is considered as the source. The reason the industry has not considered JSCC for video and images is due to two main factors. First, for video and images, 3GPP generally does not conduct independent research, but instead directly reuses source compression standards developed by external organizations such as the Moving Picture Experts Group (MPEG) and the Joint Photographic Experts Group (JPEG). Second, it is constrained by the challenges of the source-channel coding architecture. For example, in joint source-channel coding at the application layer, the channel information obtained at the application layer is often outdated, which limits the potential gains of source-channel coding. However, the situation is quite different for audio. Audio coding and decoding are designed independently by 3GPP, and currently, 3GPP SA4 is researching AI-based speech coding and decoding. Moreover, in GEO speech scenarios, where satellites remain largely stationary relative to ground stations and the channel changes slowly, the physical layer CSI can be transmitted to the application layer, overcoming the issue of outdated channel information. This makes the research on JSCC for audio more promising within 3GPP, compared to JSCC for video or image sources.  Methods  This paper firstly uses the typical physical layer parameter configurations of 3GPP GEO voice and the DAC encoder, currently being studied by SA4, as an evaluation baseline, ensuring the fairness of the proposed scheme’s gain evaluation. Based on this, the paper proposes the integration of DAC with channel coding and modulation to improve audio quality in low-bitrate scenarios. First, the total power constrained (TPC) DAC-JSCCM scheme is considered, which offers the maximum potential gain due to the joint optimization of source coding, channel coding, and modulation. Then, considering the high PAPR (Peak-to-Average Power Ratio) impact on the power amplifier, the paper proposes a constant-envelope constrained (CEC) DAC-JSCCM scheme. Finally, since the modulation symbols included in RTP payload require significant changes to the existing protocol, a DAC-JSCC scheme that is more compatible with current protocols is proposed, where bit sequences are included in RTP payload.  Results and Discussions  In the evaluation of baseline scheme 1, the link-level simulation at the physical layer is first conducted to obtain the relationship between SNR and BLER for an input length of 64 bits using a 1/3 Turbo code. Then, assuming that the error rate of an audio frame is the same as the BLER, the transmission error rate of an RTP packet containing L audio frames is calculated. According to the existing processing logic, if one audio frame in an RTP packet is transmitted with an error, the entire RTP packet will be discarded. In real-time communication, VoIP, or speech enhancement scenarios, a PESQ score of 2.5 is generally considered the minimum usable threshold, and voice quality below this score is regarded as unsuitable for professional or commercial services. Therefore, in GEO voice calls, we choose PESQ = 2.5 as the reference point. When PESQ = 2.5, the TPC DAC-JSCCM and CEC DAC-JSCCM provide a coverage gain of approximately 4 dB compared to baseline scheme 1 (Fig.4). The DAC-JSCC scheme offers a coverage gain of 2.5 dB compared to baseline scheme 1 (Fig.4). The coverage gains of DAC-JSCCM and DAC-JSCC come from the application layer performing joint source-channel coding, which avoids the cliff effect caused by using Turbo coding at the physical layer. When the Complementary Cumulative Distribution Function (CCDF) of Peak-to-Average Power Ratio (PAPR) is \begin{document}$ {10}^{-3} $\end{document}, the TPC DAC-JSCCM is 4 dB higher than the QPSK modulation, with sharper signal peaks, higher requirements for power amplifier linearity, and is more prone to distortion. In contrast, the CEC DAC-JSCCM is very close to the QPSK modulated PAPR.  Conclusions  This paper proposes the Total Power Constrained DAC-JSCCM, Constant Modulus Constrained DAC-JSCCM, and DAC-JSCCM schemes, based on the DAC codec currently being researched by SA4. Simulations are conducted using the typical physical layer configuration of the existing 3GPP GEO voice scenario. The simulation results show that, when the PESQ is 2.5 (the minimum usable threshold), compared to DAC combined with existing physical layer transmission technologies, the proposed DAC-JSCCM provides a coverage gain of 4 dB, while the proposed DAC-JSCC scheme provides a coverage gain of 2.5 dB. SA4 is currently researching more advanced GEO voice codecs beyond DAC, and we will extend our research by combining better GEO voice codecs in the future.
A Frequency Domain Self-Attention Guided MultiscaleInverse Lithography Technology
LUO Binling, WANG Ying, CAI Shuting
Available online  , doi: 10.11999/JEIT251382
Abstract:
  Objective  Optical Proximity Effect (OPE) in lithographic processes causes printed wafer patterns to deviate from target layouts. Therefore, Optical Proximity Correction (OPC) is required for mask optimization before exposure. Traditional rule-based OPC methods show reduced accuracy for complex layouts, whereas model-based OPC methods require high computational cost. Deep learning-based methods have recently been used to accelerate mask generation. However, their limited receptive fields make it difficult to model long-range optical interference, which restricts optimization accuracy. To address these limitations, this work proposes Frequency-Domain Self-Attention-Guided Multiscale Inverse Lithography Technology (FMS-ILT). The method jointly models local geometric details and global optical interference to improve printed image fidelity, edge placement accuracy, and process robustness.  Methods  FMS-ILT uses a residual convolution-based multiscale encoder-decoder architecture. Shallow layers extract fine geometric features, such as edges and corners, whereas deeper layers capture large-scale layout context. Residual blocks and multilevel skip connections are used to preserve high-frequency information and stabilize training. To overcome the limited receptive field of spatial convolutions, a Frequency-domain Self-Attention Mechanism (FSAM) is introduced at the encoder output. Global feature interactions are modeled using the Fourier transform. The resulting attention responses are then mapped back to the spatial domain through the inverse Fourier transform to adaptively reweight feature representations. A two-stage training strategy is adopted. During pretraining, a dual-branch structure jointly learns mask geometry and imaging consistency, providing physically meaningful initialization. During main training, lithography simulation is applied under nominal, maximum, and minimum process corners to refine mask optimization under physical constraints.  Results and Discussions  The comparison results with baseline models are summarized in Tables 2 and 3. FMS-ILT is used as the reference method (Ratio = 1), and all experiments are conducted on the LithoBench dataset. For the overall imaging \begin{document}$ \mathcal{L}2 $\end{document} error, FMS-ILT achieves the lowest value of 19,998, outperforming the baseline models by 2%-107%. For Process Variation Band (PVB), GAN-OPC obtains the best value of 19,156, which is 31% lower than that of FMS-ILT. However, its \begin{document}$ \mathcal{L}2 $\end{document} error and Edge Placement Error (EPE) are 107% and 1 115% higher, respectively, indicating an imbalance between imaging fidelity and edge accuracy. The remaining baseline models show PVB performance comparable to that of FMS-ILT. For EPE, FMS-ILT also shows a clear advantage, achieving an average value of 1.95, which is 47%-1 115% lower than those of the baseline models. These improvements are mainly attributed to the multiscale encoder-decoder fusion mechanism, which integrates local and global features; the combination of attention mechanisms and frequency-domain operations, which guides the model toward critical regions; and the dual-branch pretraining strategy, which introduces physical priors into the network. These modules enable FMS-ILT to achieve balanced performance in imaging fidelity, process stability, and edge accuracy.  Conclusions  This work proposes FMS-ILT for mask optimization in computational lithography. The model uses a residual convolution-based multiscale encoder-decoder architecture to extract rich spatial features. It also incorporates FSAM to jointly model local geometric details and global optical interference. A two-stage training strategy is used. In the pretraining stage, mask generation and target image reconstruction are used as dual-branch tasks to improve the physical consistency between the mask and the printed image. In the main training stage, lithography simulation is introduced to further improve imaging accuracy and process robustness. Experimental results on the public LithoBench dataset show that FMS-ILT achieves strong performance in terms of L2, PVB, and EPE. The method improves printed image quality and provides a feasible and efficient solution for computational lithography.
Lightweight Semantic Communication System Driven by User Personalization in UAV Networks
WEI Yuxuan, CHEN Xiao, CHEN Qiuyu, JIANG Hao, YANG Zhaohui
Available online  , doi: 10.11999/JEIT260370
Abstract:
  Objective  With the rapid development of the low-altitude economy and 6G intelligent networks, Unmanned Aerial Vehicle (UAV) image communication shows strong potential in target reconnaissance, emergency communication, and intelligent inspection. However, conventional pixel-level transmission cannot meet the requirements of efficient, low-latency, and intelligent communication because UAVs are constrained by limited bandwidth, payload capacity, and onboard computational resources. Semantic communication, which transmits only task-relevant information, provides an effective solution for improving communication efficiency in resource-constrained scenarios. However, current studies on UAV image transmission face several challenges. First, fixed network architectures use unified semantic encoding and transmission strategies for all users and cannot adapt to different personalized requirements. Second, new user access usually requires interest pre-training or model fine-tuning, which increases deployment overhead. Third, most models have high computational complexity. To address these issues, this paper proposes the Lightweight Personalized UAV Semantic Communication (LPUSC) system to balance computational cost, transmission bandwidth, and personalized requirements. The system enables personalized transmission through low-overhead semantic index interaction and a lightweight semantic extraction module, without pre-training for new users. A dual-branch end-to-end network is also designed. In this network, the semantic index transmission network works with the semantic image transmission network trained by a weighted hybrid loss function, thereby supporting high-precision and high-quality transmission of personalized semantic images.  Methods  The proposed LPUSC system adopts a dual-branch architecture for accurate task-driven semantic content transmission. In the semantic index interaction branch, the lightweight object detection model YOLO11s is used to perform semantic perception on UAV-captured visual scenes. Complex image information is compressed into low-dimensional semantic index vectors, which reduces transmission redundancy and communication overhead. On this basis, an end-to-end semantic index transmission network is designed to improve the robustness of semantic index transmission under complex wireless channel conditions. Through the semantic index interaction mechanism, the system accurately identifies targets of user interest and provides prior guidance for subsequent semantic content extraction. In the semantic image transmission branch, the lightweight and high-precision MobileSAM model is adopted for semantic region extraction. This branch uses the target bounding boxes returned by the semantic index interaction branch as box-prompt inputs, enabling pixel-accurate segmentation and extraction of specific semantic targets. To further improve semantic image reconstruction quality, a weighted hybrid loss function is designed. This function integrates Mean Squared Error (MSE), L1-norm loss, Structural Similarity Index Measure (SSIM) loss, gradient loss, perceptual loss, and background suppression loss. These losses jointly optimize pixel accuracy, structural preservation, and fine-detail restoration. Through the joint constraints of multiple loss terms, the proposed system improves semantic region reconstruction and achieves high-quality semantic image transmission.  Results and Discussions  Simulation results validate the proposed LPUSC system in semantic extraction and end-to-end transmission. For semantic extraction, three schemes are compared: YOLO11s-seg, YOLO11s + Segment Anything Model (SAM), and YOLO11s + MobileSAM (Fig. 4). The results show that the detection-segmentation decoupled architecture achieves better semantic boundary localization accuracy. Combined with the quantitative analysis in Table 1, the YOLO11s + MobileSAM scheme reduces resource use while maintaining high extraction accuracy. This confirms its suitability for resource-constrained UAV platforms. For end-to-end transmission, the semantic index vector transmission results (Fig. 5) show that the Bit Error Rate (BER) decreases monotonically as the Signal-to-Noise Ratio (SNR) increases in all three channel environments. The rural environment achieves the best performance, followed by the suburban and urban environments. These differences are mainly caused by variations in scatterer density and link blockage across environments. The proposed transmission network maintains stable BER under different Doppler frequencies, demonstrating its robustness under dynamic channel conditions. For semantic image transmission, the proposed weighted hybrid loss function shows good training stability (Fig. 6), and LPUSC consistently outperforms the Deep Joint Source-Channel Coding (DeepJSCC) and JPEG + Low-Density Parity-Check (LDPC) baselines across the full SNR range (Fig. 7). Specifically, LPUSC achieves SSIM and Peak Signal-to-Noise Ratio (PSNR) gains of 1.3% and 4.8% over DeepJSCC, respectively, and gains of 43% and 79.5% over JPEG + LDPC, respectively. These results indicate that the proposed personalized semantic image transmission network achieves high-quality reconstruction and remains robust to channel variations.  Conclusions  To improve the efficiency and flexibility of UAV image communication, this paper proposes LPUSC, a lightweight personalized semantic communication system. The system uses a dual-branch transmission architecture that integrates lightweight, high-precision object detection and semantic segmentation models. It enables personalized content transmission without interest pre-training. This design satisfies personalized user requirements while maintaining low computational and communication overhead. Simulation results show that the LPUSC system achieves stable and reliable semantic index interaction and outperforms the DeepJSCC and JPEG + LDPC baselines in semantic region reconstruction. The proposed system provides a useful reference for efficient UAV image semantic communication in 6G low-altitude intelligent networks.
Defending against Deepfakes by Attribute-Aware Attack
GAO Fan, YAN Weidan, SHAO Wenze, ZHANG Dengyin
Available online  , doi: 10.11999/JEIT260043
Abstract:
  Objective  Deepfakes can cause serious personal and property damage when misused. To prevent forged images from spreading, existing methods often use adversarial examples to protect facial images from deepfake manipulation. However, traditional gradient-based attacks show limited generalization and low generation efficiency in black-box attack scenarios. Their performance is also weaker than that of current methods based on Generative Adversarial Networks (GANs), which are used to train cross-model adversarial examples. Although GAN-based methods support fast inference, their lack of perceptual constraints often makes the generated adversarial perturbations visually noticeable. The rapid development of deepfake models also raises higher requirements for the generalization ability of adversarial examples. Therefore, imperceptible and generalizable adversarial attack methods are needed for proactive deepfake defense.  Methods  To further improve the transferability and imperceptibility of adversarial examples generated by existing methods, this paper proposes an attribute-aware adversarial example generation method for deepfake defense. The proposed method generates imperceptible perturbations and improves cross-model generalization through a frequency-domain identity fusion mechanism. Specifically, it focuses on the foreground regions of facial images, uses attribute-aware salient segmentation masks to separate facial and hairstyle regions, and combines these masks with adaptive spatial-frequency attention-based perturbation generators to generate region-specific adversarial perturbations. This strategy improves the imperceptibility of adversarial examples and reduces the additional computational cost caused by global processing. From the perspective of data augmentation, this paper further uses phase swapping in the frequency domain to fuse identity-related features from reference face images. This design reduces perturbation overfitting and improves generalization performance.  Results and Discussions  The proposed method is trained and tested on the CelebA-HQ dataset using proxy models. Compared with existing proactive defense methods, the experimental results show that the proposed method generates adversarial examples with strong imperceptibility and cross-model defense capability. It achieves a high defense success rate against various proxy models. The average Peak Signal-to-Noise Ratio (PSNR) of forged outputs under adversarial perturbations is reduced to 16.79 dB, representing an improvement of approximately 1.87% over the second-best method. Defense performance against HiSD is improved by approximately 7.5% compared with the second-best method. Defense performance against AttGAN is approximately 12.7% higher than that of the second-best GAN-based defense method. Moreover, the Learned Perceptual Image Patch Similarity (LPIPS) metric shows that the adversarial perturbations have high imperceptibility.  Conclusions  This study proposes a facial attribute-aware attack method for deepfake defense. The method incorporates a frequency-domain identity fusion mechanism to increase the diversity of adversarial feature inputs. Adaptive spatial-frequency attention-based perturbation generators are also designed to extract local facial information and dynamically adjust adversarial features. These designs allow the method to preserve perturbation components that are both imperceptible and attack-effective, leading to strong cross-model generalization. Future work will focus on proactive deepfake defense methods with improved imperceptibility and generalization, especially in cross-model transfer attack scenarios.
Box Particle Filter δ-GLMB Algorithm for Multiple Maneuvering Group Targets Tracking
GAN Linhai, WANG Gang, LI Zhihui, SUN Wen, WANG Baotang
Available online  , doi: 10.11999/JEIT251273
Abstract:
  Objective  Targets that move in a coordinated manner or show similar motion patterns are commonly referred to as group targets. Dense group targets contain many closely spaced individuals and often suffer from poor measurement resolvability, severe measurement overlap, and frequent target disappearance and reappearance. These factors make it difficult to establish stable tracks for individual targets within the group. Such groups are therefore usually treated as a whole to jointly estimate the kinematic state of the centroid and the extended shape. To improve tracking accuracy and computational efficiency for multiple maneuvering group targets under nonlinear measurements, an Interacting Multiple Model Gamma Box Particle δ-Generalized Labeled Multi-Bernoulli (IMM-GBP-δ-GLMB) algorithm is proposed. Tracking efficiency under nonlinear measurements is improved using the Box Particle Filter (BPF). The likelihood function of the GBP algorithm is improved, and the IMM algorithm is introduced to enhance tracking of the extended shape and centroid kinematic state of group targets. Finally, the method is integrated with the GLMB filter to track an unknown number of multiple maneuvering group targets.  Methods  Existing algorithms mainly describe the area-based overlap between the predicted extended state of group targets and the measurement distribution, but they do not fully capture shape similarity. To address this limitation, the likelihood function of the BPF is modified. Geometric parameters, including the semi-major axis, semi-minor axis, and inclination angle, are incorporated into the likelihood function. This improves the modeling of similarity between the predicted extended state and the measurement distribution. The modification is particularly useful for maneuvering group targets, because the inclination angle of the extended shape changes frequently during maneuvering. Based on IMM modeling of group motion, a model index is added to the centroid kinematic state of each box particle. The model index and centroid kinematic state are jointly estimated in each iteration, allowing mode transitions of individual box particles to be tracked and further improving tracking accuracy. The improved IMM-GBP filter is then embedded into the labeled random finite set framework, and the IMM-GBP-δ-GLMB algorithm is derived for effective tracking of multiple maneuvering group targets.  Results and Discussions  Simulation experiments are conducted to compare the proposed IMM-GBP-δ-GLMB algorithm with the IMM Sequential Monte Carlo δ-GLMB (IMM-SMC-δ-GLMB) filter. The proposed algorithm maintains comparable estimation accuracy for the centroid state, extended state, measurement rate, and number of targets, while improving computational efficiency. In the given simulation scenario, the proposed algorithm achieves a 3.8-fold improvement in timeliness, with an approximately 8.5% reduction in tracking accuracy. In scenarios with two and three group targets, the average tracking time growth rate of the proposed algorithm is 96% of that of the IMM-SMC-δ-GLMB filter. This result indicates good temporal robustness as the number of group targets increases. Therefore, the proposed algorithm has strong practical value.  Conclusions  This paper addresses the tracking of multiple maneuvering group targets under nonlinear measurement conditions by proposing the IMM-GBP-δ-GLMB algorithm. The main contributions are as follows: (1) The likelihood function of the BPF is improved to strengthen the measurement of similarity between the target extended shape and the measurement distribution, improving the tracking accuracy of group target states. (2) A motion model label is assigned to each box particle, and transitions in the target motion state are tracked during filtering. This allows the filter to achieve higher tracking accuracy with fewer box particles and improves computational efficiency. (3) The IMM-GBP method is integrated into the δ-GLMB framework to obtain the final IMM-GBP-δ-GLMB filter, which realizes effective tracking of multiple maneuvering group targets.
Research on Covert Communication Transmission Scheme Combining Relay Selection and Mode Selection over Nakagami-m Fading Channels
HUANG Haiyan, HUANG Yi, ZHANG Ning, LIANG Linlin, ZHANG Xuejun
Available online  , doi: 10.11999/JEIT260287
Abstract:
  Objective  Covert communication enhances the security of wireless communication systems by concealing both transmitted information and communication activities from unauthorized detection. However, practical wireless channels exhibit random and uncertain propagation conditions. The Nakagami-m fading channel, which can characterize a wide range of channel conditions, provides a realistic framework for evaluating the performance of covert communication. Relay-assisted transmission has attracted considerable attention because it improves transmission reliability over fading channels. Moreover, relay selection and transmission mode selection substantially affect system performance. Therefore, investigating their combined effect on covert communication over Nakagami-m fading channels is of both theoretical and practical significance for the design of next-generation secure wireless communication systems.  Methods  This paper proposes a covert communication system incorporating relay selection and transmission mode selection. The source node transmits covert information to the destination node through multiple relays, while a warden monitors transmissions from both the source and relay nodes. A friendly jammer transmits interference signals to degrade the warden’s detection capability. Four transmission schemes are considered: optimal relay selection with fixed Half-Duplex (HD) or Full-Duplex (FD) operation, optimal relay selection with random transmission mode selection, random relay selection with optimal transmission mode selection, and joint optimal relay and transmission mode selection. Closed-form expressions for the warden’s detection error probability under both HD and FD optimal relay selection are derived over Nakagami-m fading channels. Closed-form expressions for the transmission outage probability, asymptotic transmission outage probability, and covert rate are also derived for all transmission schemes. The theoretical analysis is validated through MATLAB simulations.  Results and Discussions  Simulation results demonstrate that an optimal detection threshold exists that minimizes the detection error probability (Fig. 2). As the detection threshold or jamming power increases, the warden’s ability to detect covert communication decreases, causing the detection error probability to approach one (Figs. 2 and 3). Under the same target transmission rate and high Signal-to-Noise Ratio (SNR) conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability, thereby providing the highest transmission reliability (Figs. 4 and 5). At a target transmission rate of \begin{document}$ \text{6.5 bit/(s}\cdot \text{Hz)} $\end{document}, the transmission outage probability of the joint relay and transmission mode selection scheme is 6.9% lower than that of the FD transmission scheme (Fig. 4). As the transmit power and the number of relays increase, the covert rate gradually approaches a constant value. Among all transmission schemes, the joint relay and transmission mode selection scheme consistently achieves the highest covert rate (Figs. 6 and 7).  Conclusions  This paper proposes a covert communication system based on relay selection and transmission mode selection over Nakagami-m fading channels. Closed-form expressions for the warden’s detection error probability and the system’s transmission outage probability are derived under different relay selection and transmission mode selection strategies. The asymptotic transmission outage probability and covert rate are then analyzed. Simulation results show that increasing the detection threshold or jamming power weakens the warden’s ability to detect covert communication, causing the detection error probability to approach one. Under identical target transmission rates and high SNR conditions, the joint relay and transmission mode selection scheme achieves the lowest transmission outage probability. These results indicate that appropriate relay selection and transmission mode selection not only reduce the warden’s detection capability and protect covert communication, but also improve both transmission reliability and covertness. Future work will consider practical factors, including imperfect channel state information, residual self-interference, and incomplete knowledge of the warden’s channel.
Design of Lightweight Gated Recurrent Unit Network Model Based on Memristor
HUA Honghu, XU Jia, ZHANG Bohao, WANG Wei, LI Zhiwei, LIU Haijun
Available online  , doi: 10.11999/JEIT260152
Abstract:
  Objective  With the slowdown of Complementary Metal-Oxide-Semiconductor (CMOS) technology scaling and the inherent memory-computation separation of von Neumann architectures, conventional computing systems face increasing challenges in processing large-scale sequential data. Memristors provide a promising solution because of their high integration density, fast switching speed, and synaptic plasticity. Memristor crossbar arrays naturally support Vector-Matrix Multiplication (VMM) in the analog domain, enabling energy-efficient in-memory computing. As a representative recurrent neural network, the Gated Recurrent Unit (GRU) has achieved excellent performance in sequential tasks such as trajectory prediction and urban sound classification. However, conventional hardware implementations of GRU networks require frequent data transfer between memory and processing units, resulting in high energy consumption and limited throughput. Although memristor-based GRU implementations improve computational efficiency, their large parameter size and high weight precision require substantial hardware resources and reduce deployment reliability on resource-constrained memristor arrays. In addition, device non-idealities, such as conductance fluctuations, further reduce inference accuracy. Existing memristor-based GRU methods generally treat weights and activations using the same quantization strategy without considering their different hardware implementation characteristics, and they provide limited robustness against device variations. This paper addresses these issues through a hardware-algorithm co-design strategy.  Methods  This paper proposes a lightweight memristor-based GRU network model. A 1T1R (one-transistor-one-resistor) memristor crossbar array is adopted for weight mapping and analog Multiply-Accumulate (MAC) operations. Signed weights are represented by differential pairs of positive and negative conductance matrices because memristor conductance values are inherently non-negative. A linear transformation is used to map trained network weights to memristor conductance values. To account for the different hardware implementation paths of weights and activations, a device-aware fusion quantization method based on performance analysis is proposed. Symmetric quantization is applied to weights stored in the memristor array because the zero-centered quantization range eliminates zero-point storage and simplifies write-driver circuit design. In contrast, asymmetric quantization is applied to activation values computed in peripheral circuits, thereby preserving the dynamic range and reducing quantization error. To improve robustness against memristor conductance fluctuations, weight noise training is incorporated into Quantization-Aware Training (QAT). Gaussian noise with an intensity determined by the device variation parameter is injected into quantized weights during each forward pass. This strategy acts as a regularizer that guides the model toward flatter loss minima and improves tolerance to weight perturbations. During backpropagation, the straight-through estimator updates the full-precision floating-point weights, whereas noise is dynamically resampled in every forward pass.  Results and Discussions  On the public UrbanSound8K dataset, the proposed full-precision lightweight memristor-based GRU network model achieves a classification accuracy of 93.94%. After applying the device-aware fusion quantization method, the 6-bit quantized model achieves 92.68% accuracy, corresponding to only a 1.26% decrease while reducing weight precision by 81.25% (Table 1). The proposed model outperforms Dilated Convolution (78.00%), LM-MFCC+GRU (92.00%), TFFS-DNN (88.74%), TFCNN (93.10%), and CL-Transformer (92.95%) under their full-precision settings (Table 2). Under noisy input conditions with Signal-to-Noise Ratios (SNRs) ranging from −10 dB to 10 dB, the 6-bit quantized model exhibits robustness comparable to or better than that of the full-precision model, demonstrating the effectiveness of the proposed device-aware fusion quantization strategy (Table 3). From the perspectives of storage, hardware resources, and device feasibility, 6-bit quantization reduces weight storage from 5.6 MB to 1.05 MB, corresponding to a compression ratio of 81.2%, while requiring only 2.8 million memristor cells under the 1T1R mapping scheme. Weight noise training also substantially improves robustness against device non-idealities. When the conductance variation reaches 14%, the classification accuracy increases from 82.97% to 91.14%. At the maximum simulated variation of 28%, the accuracy increases from 54.23% to 87.01% (Fig. 7), demonstrating improved tolerance to memristor device variations. On a self-constructed true-false trajectory dataset, the lightweight memristor-based GRU network model achieves 97.35% accuracy at full precision and 96.51% after 6-bit quantization, with only a 0.84% decrease, outperforming the Dilated Convolution baseline (Table 4). To further verify its applicability to different sequential tasks, the lightweight memristor-based GRU network model is evaluated on lithium-ion battery State-of-Charge (SOC) estimation using a public dataset. The 6-bit quantized model achieves Root Mean Square Errors (RMSEs) of 1.48%, 0.79%, and 0.74% at 0 °C, 25 °C, and 45 °C, respectively, outperforming the existing memristor-based GRU implementation. The proposed model also achieves lower RMSEs than the comparison method at all evaluated quantization precisions of 6 bits and above (Table 5).  Conclusions  This paper presents a lightweight memristor-based GRU network model for hardware deployment. By combining device-aware fusion quantization with weight noise training integrated into Quantization-Aware Training (QAT), the model achieves substantial memory compression while maintaining high classification accuracy and improving robustness to memristor device non-idealities. Experimental results on multiple datasets and sequential tasks demonstrate that the 6-bit quantized model preserves competitive accuracy and stable performance, providing an effective solution for deploying GRU networks on resource-constrained memristor-based edge computing platforms.
A Survey of Cooperative Mission Planning for Imaging Satellites Observing Moving Targets
XU Zhuo, FAN Shenghua, YUE Haitao, QU Tao, WANG Dingwen, SUN Shilei
Available online  , doi: 10.11999/JEIT260133
Abstract:
  Significance   Cooperative mission planning for imaging satellites observing moving targets is a key technique that supports the transition of space-based Earth observation systems from static regional coverage to a dynamic closed-loop paradigm consisting of wide-area search, dynamic tracking, and feedback-guided supplementary search. It plays an important role in emergency response, maritime monitoring, wide-area situational awareness, and persistent observation of high-value moving targets. The primary challenge arises from the conflict between uncertainty in future target states and the reliance of conventional mission planning models on deterministic inputs. Unlike static targets, moving targets have neither fixed locations nor fixed visibility windows. Their future states are generally represented by probability distributions, confidence regions, or grid-based target existence probabilities. Effective mission planning therefore requires not only accurate target motion prediction but also systematic integration of uncertainty into planning objectives, constraints, and replanning triggers. A comprehensive review from an uncertainty-driven perspective is therefore needed.  Progress   This survey reviews cooperative mission planning for imaging satellites observing moving targets from an uncertainty-driven perspective. Typical moving targets are classified into maritime moving targets, highly time-sensitive aerospace targets, and ground moving targets according to their operating environments, dynamic characteristics, and observation requirements. Although these target categories differ in maneuverability, prior constraints, and observation windows, they share a common planning challenge: coupling uncertain target motion with deterministic satellite observation actions under stringent platform and resource constraints. Methods for target motion prediction and spatiotemporal uncertainty representation are first reviewed. Physics-based methods characterize target state evolution using kinematic constraints, dynamic models, covariance propagation, reachable sets, and Markov state transition models. Data-driven methods learn motion patterns from historical trajectories, Automatic Identification System (AIS) data, remote sensing observations, meteorological information, and geographic constraints. From the perspective of mission planning, the utility of these methods depends on whether outputs such as covariance, target existence probability, confidence regions, and information gain can be directly incorporated into planning models. Observation task modeling, cooperative planning architectures, optimization algorithms, and closed-loop replanning mechanisms are then analyzed. Deterministic task models simplify uncertainty into trajectory points, visibility windows, or fixed geographic regions, while probabilistic task models incorporate target existence probability, belief states, and information gain into objective functions, constraints, and state transition models. Centralized, distributed, and hybrid planning architectures are compared with respect to global optimization capability, onboard autonomy, communication overhead, and response timeliness. Exact optimization methods, heuristic methods, metaheuristic algorithms, Deep Reinforcement Learning (DRL), and Large Language Model (LLM)-assisted solution strategies and algorithm design are also reviewed. Finally, state-triggered replanning, Receding Horizon Optimization (RHO), and Model Predictive Control (MPC) are summarized as representative approaches for closed-loop dynamic scheduling.  Conclusions  The reviewed studies indicate that cooperative mission planning is evolving from open-loop static scheduling to closed-loop dynamic planning. Nevertheless, several challenges remain. First, uncertainty information generated during target prediction is not fully exploited in planning decisions. Rich probabilistic information is frequently reduced to deterministic time windows, discrete trajectory points, or geometric regions, thereby limiting risk-aware task allocation. Second, distributed cooperation lacks reliable belief-state consistency. Differences in local observations may lead satellites to maintain inconsistent estimates of the same target state, resulting in redundant observations, task conflicts, and inefficient resource utilization. Third, dynamic replanning lacks unified benefit-cost criteria for determining replanning triggers. Excessively frequent replanning increases attitude maneuver time, energy consumption, and onboard storage resource use, whereas delayed replanning may fail to respond to actual target maneuvers. Fourth, LLMs have demonstrated potential for task requirement parsing, constraint modeling, heuristic generation, and algorithm design for satellite scheduling, but their application to cooperative mission planning for moving targets remains limited.  Prospects   Future research should focus on developing a more robust closed-loop planning framework. Prediction uncertainty should be incorporated directly into planning models through chance-constrained planning, belief-state planning, or Partially Observable Markov Decision Processes (POMDPs), enabling covariance, target existence probability, and information entropy to be integrated into planning objectives, constraints, and replanning triggers. Bayesian updating or sequential Bayesian filtering should use both successful detections and missed detections to continuously refine the prediction layer. Distributed cooperation requires lightweight state synchronization and belief fusion supported by compact state-sharing descriptors and event-triggered communication. Replanning decisions should be guided by information gain and benefit-cost evaluation. In addition, LLMs should be developed as verifiable auxiliary tools rather than direct replacements for optimization solvers. They can assist with task requirement structuring, constraint modeling, heuristic generation, and algorithm component design, whereas feasibility verification, solution refinement, and performance evaluation should remain the responsibility of formal verification methods, conventional optimization algorithms, and simulation environments. These research directions are expected to improve the robustness and uncertainty awareness of mission planning for satellite observation of moving targets.
Heterogeneous Task Cooperative Scheduling Architecture for Networked Radar in Saturation Attack Air Defense Early Warning
YE Juhang, FANG Yuyuan, WEI Shaopeng, DUAN Jia, ZHANG Lei
Available online  , doi: 10.11999/JEIT260373
Abstract:
  Objective  To address the severe challenges posed by Unmanned Aerial Vehicle (UAV) swarms and intelligent loitering munitions, which generate massive, sudden, and heterogeneous early warning tasks during saturation attacks, a scalable networked-radar cooperative scheduling architecture is proposed. Existing architectures suffer from rigid dynamic coordination and insufficient capability for heterogeneous task scheduling. The non-convex cooperative scheduling problem is therefore decoupled into a multi-stage decision process consisting of multidimensional dynamic resource coordination and adaptive heterogeneous task scheduling. By incorporating a dispatch mechanism and Hierarchical Reinforcement Learning (HRL), a hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling is developed. The architecture is implemented through an execution-redispatch cognitive closed loop, together with a Target Dispatch Algorithm (TDA) and a hierarchical command-and-scheduling method, to address multi-radar, multi-target, and multi-task scheduling in saturation attack scenarios.  Methods  The cooperative scheduling problem is first decoupled into dispatch-oriented network-level multidimensional resource coordination and hierarchical-command-based node-level adaptive heterogeneous task scheduling. An environment perception layer establishes a dynamic uncertainty model based on the Bayesian Cramér-Rao Lower Bound (BCRLB) and radar detection probability to jointly characterize target threat levels and radar operating states. The network coordination layer then adopts an adaptive weighted dispatch model based on comprehensive combat effectiveness to achieve dynamic target allocation while constructing a scalable execution-redispatch cognitive closed loop. To solve the resulting large-scale, nonlinear, multi-constraint generalized bipartite graph matching problem, the proposed TDA, an MMAS-based algorithm incorporating feasibility-rule-based constraint handling, is developed as a constructive solution. At the node level, Hierarchical Q-Learning (HQL) is implemented through a serial dual-Q-table implementation for distributed heterogeneous task scheduling. Using task proportion control as the hierarchical subgoal, the proposed method transforms upper-level operational intent into lower-level beam dwell scheduling, enabling adaptive, long-term, and interpretable execution of heterogeneous tasks with complex dependency relationships.  Results and Discussions  A simulated point-defense scenario against UAV swarm and loitering munition saturation attacks is established using three networked homogeneous S-band medium-range phased-array radars to counter 200 high-speed maneuvering targets. For network-level coordination, the proposed TDA replaces conventional penalty functions with a hierarchical solution framework based on feasibility-rule constraint handling. By exploiting prior model information, TDA achieves higher solution quality and faster convergence than the Max-Min Ant System (MMAS), Artificial Bee Colony (ABC), and Genetic Algorithm (GA) (Fig. 3). Although computational complexity increases, the execution time remains well within the dispatch cycle, improving solution quality with only millisecond-level computational overhead while satisfying real-time operational requirements (Fig. 4). For node-level scheduling, HQL employs hierarchical macro- and micro-level decisions to ensure policy consistency. Supported by an internal dense transfer-reward mechanism, HQL achieves higher learning efficiency and better long-term policy quality than Q-Learning (QL) and the Priority-Based Method (PBM) (Fig. 5). The hierarchical serial dual-Q-table framework maintains balanced task proportions, maximizes comprehensive combat effectiveness, and improves resource utilization (Fig. 6). Furthermore, comparisons of the target track-loss rate and mean tracking error show that HQL achieves the lowest mean tracking error by prioritizing high-quality tracking tasks, despite a moderately higher target track-loss rate, demonstrating superior long-term scheduling capability (Fig. 7).  Conclusions  The proposed hybrid architecture integrating network-level centralized dynamic target allocation with node-level distributed heterogeneous task scheduling effectively addresses the limitations of rigid dynamic coordination and insufficient heterogeneous task scheduling capability in existing networked radar systems. The proposed framework enables real-time cooperative scheduling of search, confirmation, and tracking tasks during large-scale, high-speed saturation attacks, thereby improving the operational capability of the air defense early warning system. Simulation results demonstrate improvements in applicable processing scale, environmental adaptability, long-term scheduling capability, scalability, and interpretability. Future work will explore online learning to reduce the discrepancy between offline training and online deployment and will further extend the architecture by integrating weapon-target assignment to support unified early warning and fire-control systems.
Analysis of Age upon Decisions and Distortion at Decisions in IoT Status Update Systems with Batch Arrivals
LIU Lei, JIN Wenkai, ZHANG Qingqing, LI Yuzhou, JIANG Fan
Available online  , doi: 10.11999/JEIT260359
Abstract:
  Objective  The rapid development of the Internet of Things (IoT) makes the timely transmission and processing of status updates essential for modern systems, where information freshness at decision epochs plays a critical role. In many IoT applications, such as smart grid fault detection and Industrial Internet of Things (IIoT) cluster monitoring, status updates typically arrive in batches rather than individually. However, most existing studies on Age of Information (AoI) assume single-update arrivals and therefore cannot accurately characterize the queueing dynamics caused by batch arrivals. Besides information freshness, distortion at decision epochs is another key factor because it directly affects decision quality. A fundamental tradeoff therefore exists between information freshness and distortion. Waiting for more complete status update information allows more completed status updates to be incorporated into joint estimation, but increases queueing and transmission delays, thereby reducing information freshness. In contrast, triggering decisions earlier reduces delay but increases distortion because fewer completed status updates are available for joint estimation. To address this problem, this paper investigates the tradeoff between information freshness and distortion in an IoT status update system with batch arrivals by adopting Age upon Decisions (AuD) and Distortion at Decisions (DaD) as performance metrics. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. Furthermore, for the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. The results provide theoretical insight and practical guidance for the design of IoT status update systems with batch arrivals.  Methods  Information freshness and distortion at decision epochs are analyzed for an IoT status update system with batch arrivals. AuD and DaD are adopted to quantify information freshness and distortion, respectively. Based on queueing theory, analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. A typical case with geometrically distributed batch sizes is then investigated. An alternating iterative optimization algorithm is further developed to jointly optimize the batch arrival rate, average batch size, and decision threshold to minimize the weighted sum of the average AuD and average DaD.  Results and Discussions  Simulation results validate the theoretical analysis. The average AuD exhibits a nonmonotonic trend as the batch arrival rate increases, first decreasing and then increasing. In addition, the Batch-size Coefficient Of Variation (BCOV) has a significant effect on the average AuD, with a smaller BCOV providing better information freshness performance. Under high-load conditions, queue backlogs become more severe, and stochastic fluctuations in batch arrivals have a greater effect on the queueing process. This increases service-time variability and amplifies the effect of BCOV on the average AuD (Fig. 2). As the average batch size increases, the system queue length and queueing delay increase, leading to a larger average AuD. At the same time, the decision control unit can utilize more completed status updates for joint estimation, thereby reducing the average DaD (Fig. 3). Moreover, the average DaD decreases as the decision threshold increases because more completed status updates are incorporated into the joint estimation process, improving estimation accuracy. A larger BCOV also increases the number of completed status updates available for joint estimation and therefore further reduces the average DaD (Fig. 4). The optimization results show that the solutions obtained by the proposed algorithm lie on the Pareto frontier, demonstrating its effectiveness. By comparison, fixed batch arrival rates and decision thresholds produce performance that is considerably farther from the Pareto frontier, demonstrating the advantage of jointly optimizing system parameters (Fig. 5).  Conclusions  This paper investigates an IoT status update system with batch arrivals by adopting AuD and DaD to quantify information freshness and distortion, respectively. Analytical expressions for the average AuD and average DaD are derived under a general batch-size distribution. For the typical case of geometrically distributed batch sizes, an alternating iterative optimization algorithm is developed to jointly optimize the batch arrival rate, average batch size, and decision threshold, thereby minimizing the weighted sum of the average AuD and average DaD. Simulation results verify the theoretical analysis and reveal the effects of the batch arrival rate, average batch size, and decision threshold on the average AuD and average DaD. The results also demonstrate that the proposed low-complexity algorithm effectively identifies Pareto-optimal solutions for the AuD-DaD tradeoff. This study considers only the batch arrival characteristics of status updates. Future work will incorporate batch service mechanisms to further examine their effects on the tradeoff between AuD and DaD. Flexible decision mechanisms can also be developed to achieve adaptive AuD-DaD tradeoffs according to the heterogeneous requirements for information freshness and distortion across applications with different batch characteristics.
Kolmogorov-Arnold Nonlinear Enhancement Method for Aerial-Ground Person Re-IDentification
CHEN Yijun, ZENG Xianxian, LIU Shun, WANG Leijun
Available online  , doi: 10.11999/JEIT260430
Abstract:
  Objective  Aerial-Ground Person Re-IDentification (AG-PReID) aims to match the same person across Unmanned Aerial Vehicle (UAV) and ground-camera views. Compared with conventional same-platform person re-identification, this task faces larger cross-view appearance variation and more severe cross-domain distribution shifts. Under these conditions, identity-consistent cues are often weakened by strong viewpoint asymmetry and cross-domain appearance distortion. Existing methods mainly focus on feature extraction and cross-view representation alignment. However, the classification supervision branch still relies heavily on linear feature transformation, which limits its ability to model complex nonlinear discriminative relationships in high-dimensional feature spaces. A stronger nonlinear supervision mapping is therefore needed to better exploit high-order feature interactions and local discriminative variations. To address this issue, this paper proposes a Kolmogorov-Arnold Nonlinear Enhancement Module (KANEM). KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier. It uses learnable nonlinear mappings to adaptively enhance features for more discriminative cross-view representation learning.  Methods  The backbone follows the View-Decoupled Transformer (VDT), which introduces an additional view token and performs layer-wise view decoupling. This design separates view-related factors from identity features and reduces representation bias between aerial and ground domains. Based on this framework, KANEM replaces the conventional fully connected feature transformation between backbone features and the linear classifier, thereby providing adaptive nonlinear mappings for feature enhancement. Specifically, KANEM consists of a base activation branch and a spline branch, which are stacked into cascaded function-mapping layers. This design enables more flexible nonlinear modeling than conventional linear or MultiLayer Perceptron (MLP)-based transformations. It allows the model to capture local nonlinear variations and complex correlations among feature dimensions. To improve discriminability and further separate identity and view information, the network is jointly optimized using identity classification loss, view classification loss, triplet loss, and orthogonality loss. KANEM is used only during training and is removed during inference, so no extra inference cost is introduced.  Results and Discussions  Comprehensive evaluations are conducted on the CARGO and AG-ReID datasets. The results show that the proposed method consistently performs better than the baseline model and existing state-of-the-art methods. On CARGO, the proposed method achieves 70.19%/63.16%/51.34% in Rank-1 accuracy, Mean Average Precision (mAP), and Mean Inverse Negative Penalty (mINP), respectively, under the overall ALL retrieval protocol. It also achieves 58.75%/53.27%/41.11% under the most challenging aerial-ground (A↔G) cross-view retrieval protocol (Table 1). On AG-ReID, the proposed method achieves the best performance under both retrieval protocols. It reaches 84.41%/76.21%/53.05% in Rank-1/mAP/mINP for aerial-to-ground (A→G) retrieval and 86.69%/77.99%/52.28% for ground-to-aerial (G→A) retrieval (Table 2). Ablation studies on CARGO further verify the effectiveness of KANEM. They show that KANEM achieves better overall performance than conventional linear transformation and MLP-based alternatives. This result indicates that the proposed nonlinear enhancement strategy is more suitable for supervision mapping in CARGO (Tables 3 and 4). In addition, integrating KANEM into other person re-identification tasks further demonstrates its potential generalization ability across different scenarios (Table 5). Parameter analysis shows that setting λ to 0.001 enables the model to better balance the complexity difference between view classification and identity classification (Fig. 2(a)). When G and P are set to 5 and 3, respectively, the model effectively fits nonlinear variations in the feature space while preserving the smoothness and continuity of spline functions. This setting achieves effective nonlinear feature enhancement (Fig. 2(b)(d)). The two-dimensional t-distributed Stochastic Neighbor Embedding (t-SNE) visualization shows that the enhanced features have higher intra-class compactness and better inter-class separability (Fig. 3). The top-5 retrieval comparisons further provide qualitative evidence that the proposed method improves ranking quality and retrieval robustness under all four retrieval protocols on CARGO. It promotes correct matches to higher positions and returns more relevant samples among the top-ranked results (Fig. 4).  Conclusions  This paper presents KANEM for AG-PReID. The proposed module is motivated by the large discrepancy between UAV and ground-camera views and by the limited capacity of linear feature transformation in the classification branch to capture complex nonlinear discriminative relationships. By replacing the conventional fully connected feature transformation between backbone output features and the linear classifier, KANEM provides a more flexible nonlinear supervision mechanism for cross-view representation learning. Through adaptive nonlinear enhancement, it better models complex feature interactions in high-dimensional spaces and strengthens the representation of cross-view consistency and fine-grained discriminative cues. Experimental results on CARGO and AG-ReID demonstrate the effectiveness of the proposed method, particularly in challenging scenarios with large view discrepancies. Future work will further refine the nonlinear mapping mechanism of KANEM and explore its use in more complex cross-view settings to improve model discriminability and generalization performance.
A Task Prediction-augmented Hierarchical Offloading Method for Space-Air-Ground Integrated Networks
ZHANG Linghao, XU Bo, SUN Jinlong, LAI Haiguang, ZHAO Haitao
Available online  , doi: 10.11999/JEIT260217
Abstract:
  Objective  Space-Air-Ground Integrated Networks (SAGIN) have become key infrastructure for future 6G communications. They support wide-area coverage and flexible deployment through the coordinated operation of Low Earth Orbit (LEO) satellites, Unmanned Aerial Vehicles (UAVs), and Ground Users (GUs). With the rapid growth of Internet of Things (IoT), Internet of Vehicles (IoV), and smart city applications, terminal devices generate increasingly diverse computation-intensive tasks. These tasks impose high requirements on real-time computing and resource scheduling. Mobile Edge Computing (MEC) has been integrated into SAGIN architectures to provide near-user computing services by using UAVs and satellites as edge nodes, thereby reducing task completion latency. However, efficient task offloading remains challenging when average task completion latency and UAV flight energy consumption must be jointly reduced. This difficulty is caused by the strong coupling among UAV trajectory planning, task offloading, and computational resource allocation. It is further intensified by the dynamic and partially observable nature of SAGIN environments. Existing Multi-Agent Reinforcement Learning (MARL) methods mainly rely on reactive decisions based on instantaneous observations. They lack awareness of future task workload changes, which leads to decision lag and limited adaptability under bursty traffic. To address these issues, a task prediction-augmented MARL method is proposed to support forward-looking decisions in dynamic SAGIN environments.  Methods  A three-layer SAGIN-MEC architecture is considered, including one LEO satellite, multiple UAVs, and GUs. Tasks can be processed locally, offloaded to UAVs through Ground-to-Air (G2A) links, or further relayed to the LEO satellite through Air-to-Satellite (A2S) links under a partial offloading mechanism. The joint optimization of UAV trajectory, user association, offloading ratios, and computational resource allocation is formulated as a Mixed-Integer NonLinear Programming (MINLP) problem. The objective is to minimize the weighted sum of average task completion latency and UAV flight energy consumption. Owing to the nonconvexity and high dimensionality of this problem, it is reformulated as a DECentralized Partially Observable Markov Decision Process (DEC-POMDP). A Prediction-Augmented Multi-Agent Proximal Policy Optimization (PA-MAPPO) algorithm is then developed. A lightweight Exponential Smoothing-Autoregressive (ES-AR) prediction module is used to generate multi-step workload forecasts, which are incorporated into the state space of each agent. The algorithm adopts a bilevel structure. In the outer layer, Centralized Training and Decentralized Execution (CTDE)-based PA-MAPPO generates UAV trajectory actions. In the inner layer, Block Coordinate Descent (BCD)-based convex optimization solves the resource allocation and offloading subproblems, and closed-form resource allocation solutions are obtained through Lagrangian analysis. Generalized Advantage Estimation (GAE) and the PPO-Clip objective are used to improve training stability and convergence.  Results and Discussions  Simulations are conducted with one LEO satellite, five UAVs, and 50 GUs in a 1×1 km2 area. PA-MAPPO is compared with MAPPO without prediction and Prediction-Augmented Multi-Agent Deep Deterministic Policy Gradient (PA-MADDPG). The training curves show that PA-MAPPO converges within 500~700 episodes, with the highest average reward and the smallest variance, indicating better stability (Fig. 3). As the number of GUs increases from 20 to 80, PA-MAPPO consistently achieves the lowest system cost. Compared with MAPPO and PA-MADDPG, it reduces the average cost by approximately 12.4% and 18.7%, respectively (Fig. 4). Experiments with different UAV numbers show a U-shaped cost curve for all algorithms. The best configuration is obtained when U=5, where PA-MAPPO achieves the minimum cost (Fig. 5). Sensitivity analysis of the latency-energy tradeoff weight ω confirms that PA-MAPPO remains robust under different optimization preferences (Fig. 6). The prediction horizon H has a nonmonotonic effect on performance. When H=5, PA-MAPPO obtains the best result and reduces the cost by approximately 14.9% compared with the no-prediction case. Longer horizons degrade performance because prediction errors accumulate (Fig. 7).  Conclusions  The PA-MAPPO algorithm is proposed to solve the joint optimization of UAV trajectory planning, user association, task offloading, and computational resource allocation in dynamic SAGIN environments. By integrating a lightweight ES-AR task workload prediction module into the MARL process, PA-MAPPO enables UAV agents to account for future task dynamics. This design reduces the decision lag caused by purely reactive methods. The inner BCD-based convex optimization converges to a Karush-Kuhn-Tucker (KKT)-stationary point, while the outer CTDE-based PPO mechanism improves training stability and scalability. Simulation results show that PA-MAPPO outperforms baseline methods in average task completion latency, UAV flight energy consumption, and overall system cost. It also maintains strong scalability and robustness under different system configurations. Future work will study online prediction and decision co-optimization in multi-satellite cooperative scenarios and examine the effect of dynamic network topology changes on algorithm performance.
A Dual-polarized Magnetoelectric Dipole Antenna Array with Differential Feeding
TANG Li, WANG Zhihui, ZHAO Luyu
Available online  , doi: 10.11999/JEIT260505
Abstract:
  Objective  This study addresses key challenges in Fifth-Generation (5G) millimeter-wave terminal antennas by designing a compact, high-performance dual-polarized array. Existing designs often face trade-offs among bandwidth, beam-scanning range, and integration complexity. To address these limitations, this paper proposes a differentially fed magnetoelectric dipole array. A stacked stripline-slot-stripline balun is used to enable efficient single-ended-to-differential conversion, and the array design is optimized. The objective is to realize an integrated solution with wideband operation, low cross-polarization, wide-angle beam scanning, and high integration density for practical 5G millimeter-wave applications.  Methods  A structured design method is adopted. First, a stacked differential balun based on a stripline-slot-stripline configuration is developed to achieve efficient single-ended-to-differential conversion. A single-polarized magnetoelectric dipole antenna element is then designed and integrated with the balun, and its performance is characterized. The design is further extended by orthogonally integrating two elements to form a dual-polarized unit, which is used to construct a 1×4 linear array. Iterative full-wave electromagnetic simulation and optimization are conducted to balance wideband impedance matching, high port isolation, stable wide-angle beam scanning, grating-lobe suppression, and mutual-coupling reduction.  Results and Discussions  The optimized 1×4 dual-polarized differentially fed magnetoelectric dipole antenna array uses an element spacing of 4.6 mm, corresponding to 0.4 free-space wavelength at 26 GHz. This spacing achieves a favorable balance between grating-lobe suppression and inter-element mutual-coupling reduction. The measured –10 dB reflection coefficient bandwidths are 25~29.4 GHz for the +45° polarization port and 25~27.7 GHz for the –45° polarization port (Fig. 20). The slight matching difference is attributed to the incomplete structural symmetry of the baluns under the two polarization modes (Fig. 13). At 26 GHz, both polarization modes provide a peak gain of 10.7~11 dBi and support ±60° wide-angle beam scanning, with main-lobe gain attenuation no greater than 3 dB (Fig. 21). The measured radiation performance agrees well with the simulated results. Minor deviations are mainly caused by the high dimensional sensitivity of millimeter-wave structures and small errors in fabrication and test assembly. The array also maintains stable low cross-polarization and high port isolation across the operating band. These results are achieved through equal-length feed lines, symmetric layout, ground-pad shielding, and metallized-via electromagnetic isolation (Fig. 16), which suppress mutual coupling and parasitic radiation and ensure consistent dual-polarized radiation performance.  Conclusions  This paper presents a dual-polarized magnetoelectric dipole antenna array with differential feeding for 5G millimeter-wave applications. By using a stacked stripline-slot-stripline balun and optimizing the radiating structure and array layout, the design achieves wide bandwidth, high gain, low cross-polarization, and wide-angle beam scanning. The differential balun enables efficient single-ended-to-differential conversion with good amplitude and phase balance across the target band. The implemented 1×4 array, with an optimized element spacing of 4.6 mm, achieves a simulated peak gain of 11 dBi at 26 GHz and supports ±60° beam scanning, with gain variation below 3 dB. The overall design verifies the feasibility of a differentially fed magnetoelectric dipole architecture for compact, high-performance 5G millimeter-wave terminal antenna modules. Future work may focus on larger array configurations and further integration with BeamForming Integrated Circuits (BFICs).
An Incremental Density Clustering Method with Time-Difference Prior for Mobile Multi-Station Radar Signal Sorting
CHEN Jinli, FAN Yu, WANG Yanjie, ZHANG Jindong
Available online  , doi: 10.11999/JEIT260151
Abstract:
  Objective  In modern electronic warfare, radar signal sorting is a pivotal technology for electronic reconnaissance. Its primary objective is to deinterleave and categorize pulses from multiple radar emitters within dense and overlapping pulse streams. However, in practical reconnaissance missions, particularly those involving mobile platforms such as aircraft, the observation stations are in constant motion, while the target radar emitters typically remain stationary. This dynamic geometric relationship between the observation stations and the emitters causes the Time Difference of Arrival (TDOA) of intercepted signals to exhibit non-stationary characteristics that evolve over time. Conventional sorting algorithms lack the mechanism to perceive this dynamic feature drift. Consequently, continuous TDOA trajectories generated by the same emitter are prone to being incorrectly partitioned into multiple clusters during data processing, resulting in cluster proliferation and degraded sorting performance. To address these issues, an incremental density-based clustering method incorporating TDOA prior information is proposed. The proposed method effectively exploits the temporal evolution characteristics of TDOA, achieving high stability and high sorting accuracy in mobile multi-station scenarios.  Methods  First, by analyzing the temporal evolution characteristics of TDOA in mobile multi-station scenarios, the TDOA evolution process is approximated as a linear function of time, and a linear prior model is established to characterize its dynamic behavior. An online micro-cluster construction strategy is then adopted for multi-station TDOA samples, and macro-clusters are formed according to the spatiotemporal intersection relationships among micro-clusters. For each macro-cluster, independent Kalman filters are constructed for different TDOA dimensions. The TDOA value and its rate of change are defined as state variables, and recursive state prediction is performed to provide dynamic feature references for newly arriving samples. During sample association, a spatiotemporal joint scoring function is designed, in which prediction residuals generated by the Kalman filters are incorporated as dynamic constraints. Consequently, the matching criterion evolves from a conventional static density-based rule into a joint spatiotemporal consistency constraint, enabling more accurate assignment of newly arriving TDOA samples. To suppress cluster proliferation caused by TDOA feature evolution, a concept drift detection mechanism based on macro-cluster center evolution is further introduced. Once concept drift is detected, posterior state vectors and covariance matrices estimated by the Kalman filters are employed to construct a dual Mahalanobis-distance criterion that jointly evaluates state-distribution overlap and predicted-position overlap. Under a 95% confidence threshold, mis-split radar clusters are adaptively merged and assigned to the same radiation source, thereby effectively suppressing cluster proliferation induced by concept drift.  Results and Discussions  In the simulation experiments, radar pulse data distributions are generated according to the radar parameters (Table 1). The multi-station TDOA corresponding to the same radiation source exhibits a near-linear evolution trend with respect to the Time of Arrival (TOA) at the observation station(Fig. 5). To comprehensively evaluate the performance of the proposed method in complex environments, a simulation analysis on radar cluster proliferation suppression is first conducted (Fig. 6). In this scenario, radar emitters E4 and E9 are selected from the nine simulated radiation sources as representative cases for analysis. Compared with DBSCAN, Incremental Density-Based Clustering (ICDC), cloud model-based sorting, and PointNet++ sorting algorithms, the proposed method effectively suppresses radar cluster proliferation and demonstrates superior cluster stability during the sorting process. Furthermore, to further demonstrate the superiority of the proposed approach, its sorting performance was comprehensively compared with several representative algorithms, including the conventional histogram method, grid-based clustering, DBSCAN, ICDC algorithm, cloud model-based sorting, and PointNet++ sorting algorithms. The sorting accuracy of various algorithms under different TOA measurement errors (Fig. 7) indicates that the proposed method achieves a sorting accuracy exceeding 96% when the TOA measurement error ranges from 50 to 300 ns, demonstrating remarkable robustness to measurement noise. Under different pulse interference rates (Fig. 8), the sorting accuracy of the proposed method remains above 96%, exhibiting excellent interference suppression capability. To further assess the stability of different algorithms under mobile observation platforms, simulations were performed at different observation station velocities (Fig. 9). The proposed method consistently maintains a high sorting accuracy across all tested velocities. Even at relatively high observation station speeds, the sorting accuracy remains at approximately 94%, indicating that the proposed approach retains favorable stability and robustness under dynamic observation conditions. In addition, simulations on radar cluster proliferation and missed-cluster probabilities are conducted (Fig. 10), and the computational complexities of different algorithms are analyzed (Table 2). The results demonstrate that the proposed method provides a favorable balance among sorting accuracy, cluster proliferation suppression, missed-cluster control, and computational cost.  Conclusions  To address the sorting performance degradation caused by the temporal evolution of TDOA features in mobile multi-station scenarios, an incremental density-based clustering method incorporating TDOA prior information is proposed. By incorporating observation station motion information to construct a dynamic TDOA state model and integrating Kalman Filters to achieve recursive prediction of TDOA evolution, the proposed method effectively suppresses radar cluster proliferation and mis-splitting caused by concept drift. In addition, the spatiotemporal joint criterion and the adaptive cluster merging mechanism further enhance the robustness of the algorithm under complex and dynamic environments. Simulation results verify the effectiveness and stability of the proposed method in mobile multi-station cooperative reconnaissance scenarios. Future work will focus on real-time multi-parameter fusion-based sorting methods in complex electromagnetic environments. Furthermore, adaptive adjustment strategies for the micro-cluster spatial intersection threshold, process noise covariance matrix, and measurement noise variance will also be investigated.
Resilience-Aware Cooperative Mission Planning Algorithm for Multi-UAV Systems in Complex Dynamic Environments
ZHAO xuejian, XIE lulu, WANG enliang
Available online  , doi: 10.11999/JEIT260138
Abstract:
  Objective  This paper addresses the strongly coupled problem of task allocation and route planning in cooperative mission planning for multiple UAV systems under complex dynamic environments, where dynamic task arrivals, UAV failures, no fly zone constraints, and link quality degradation coexist.   Methods  A resilience aware hybrid swarm optimization algorithm, termed RAHSO, is proposed. The method first builds an integrated task and route planning model that incorporates task value, route cost, energy consumption, interference penalty, time window constraints, capability constraints, conflict resolution, and link quality. It then combines clustering and genetic initialization, DBO and PSO hybrid search, GA and VNS local refinement, and Tarjan based deadlock detection and repair to obtain high quality baseline solutions. Furthermore, an event driven PPO online replanning module is introduced to rapidly adjust affected task subsets when emergent tasks, UAV failures, or topology changes occur.   Results and Discussions  Comparative and ablation experiments are conducted under static, scale-expansion, dynamic-event, and interruption scenarios. The results demonstrate that the proposed method achieves consistent advantages in task completion rate, total utility, average energy consumption, recovery time, and resilience index over representative baselines while maintaining acceptable replanning latency.   Conclusions  The proposed method provides an effective solution for resilient cooperative mission and route planning of multiple UAV systems in complex dynamic environments.
Survey of Satellite Covert Communications: Status, Key Technologies, and Future Challenges
DENG Hao, SUN Weiyuan, ZHU Zhengyu, PAN Gaofeng, SUN Gangcan
Available online  , doi: 10.11999/JEIT260177
Abstract:
  Objective  This survey comprehensively integrates the theoretical foundations, key technologies, and future challenges in the field of satellite covert communications. Based on the classic Alice-Bob-Willie model, it analyzes the impact of satellite channel characteristics on covert communication capacity, laying the theoretical groundwork. It summarizes the covert communication network model under the space-based, air-based, and ground-based three-layer architecture (Fig. 1), and systematically reviews core technologies and their optimization methods, including signal camouflage coding, beamforming, spectrum diversity, quantum encryption, and AI-assisted techniques. The main security threats faced by satellite covert communications are outlined, and multi-layered defense strategies, such as physical layer security and intelligent collaborative protection, are summarized. This provides theoretical and technical support for promoting highly secure and intelligent development in this field.  Significance   The research significance of this survey lies in its systematic integration of the theoretical framework and technological systems in the field of satellite covert communications. Addressing the threats of detection, interference, and eavesdropping in the vast, open satellite environment, it summarizes representative space-air-ground architectures reported in the literature. These architectures overcome the limitation of traditional encryption technologies that only protect information content, supporting low probability of detection by reducing statistical distinguishability at the physical layer. By elucidating the constraints of satellite channels on covert capacity through the refined square root law, and reviewing enhancement strategies reported in prior work centered on core technologies such as dynamic encoding, beamforming, and spectrum diversity, it provides theoretical and technical foundations for constructing highly survivable space-air-ground integrated secure communication networks. This holds significant strategic value for national defense, emergency communications, and the security assurance of 6G integrated space-terrestrial networks.  Progress   Existing studies reveal the dual impact of Doppler spread on satellite covert capacity: while increasing the missed detection probability, it simultaneously causes signal distortion, necessitating reliance on adaptive coding for compensation. The research further quantifies the detection characteristic differences among terrestrial, aerial, and orbital wardens (Willie) (Table 1), providing a theoretical basis for hierarchical defense design. In terms of covertness enhancement techniques, existing schemes propose multi-level strategies. AI-driven dynamic camouflage combined with sparse coding integrates background noise, inter-satellite links, and dynamic beamforming, improving covert throughput. At the network level, cooperative UAV-assisted transmission and dynamic spectrum coordination are summarized as typical network-level enhancement schemes (Fig. 3). Based on representative studies, a hierarchical defense framework for satellite covert communications is summarized. This combines reconfigurable intelligent surface control, Stackelberg game-theoretic incentives for jamming cooperation, and XOR network coding, and utilizes federated learning to achieve cross-domain threat signature sharing. These systematic advances provide innovative solutions for the covertness and security of satellite communications.  Conclusions  This paper systematically investigates the foundations and advancements of satellite covert communication, highlighting the integration of multi-layer satellite constellations, dynamic aerial relays, and quantum-encrypted, software-defined networks to establish resilient global covert channels. By adapting the Alice-Bob-Willie model to real-world satellite channel imperfections, it guides covert throughput and security optimization. The infusion of AI into coding and waveform design enables adaptive, environment-aware concealment strategies. Future research should focus on robust covert links in dynamic LEO environments, scalable constellation management, and the deep integration of AI and quantum technologies for 6G NTN systems, as the convergence of programmable satellites, intelligent surfaces, and advanced machine learning is expected to further influence secure space communications.  Prospects   Future research challenges and development trajectories focus on four critical domains: robust transmission under non-ideal channels, AI-enabled intelligent decision-making, 6G NTN integrated networking, and quantum-communication integration (Fig. 4). High-precision Doppler compensation models must be developed to mitigate rapid channel variations in LEO satellites. Concurrently, robust transmission mechanisms should be developed under non-ideal CSI conditions, potentially leveraging deep reinforcement learning for real-time resource optimization. Research should prioritize synergistic advancement of AI and quantum technologies. This entails integrating cross-layer Quantum Key Distribution (QKD) designs with covert transmission protocols, while utilizing Software-Defined Satellite (SDS) capabilities for dynamic strategy deployment. Key opportunities include exploiting Reconfigurable Intelligent Surfaces (RIS) for enhanced spatial-domain signal control and implementing blockchain solutions to address trust constraints in multi-node cooperative networks. Essential objectives encompass deploying efficient lightweight onboard algorithms and establishing optimized international coordination frameworks for spectrum and orbital resource allocation.
Cross-Domain Collaborative Enhancement for Tiny Object Detection in Remote Sensing Images
ZHANG Tianyang, ZHANG Xiangrong, WANG Guanchun, TANG Xu
Available online  , doi: 10.11999/JEIT260317
Abstract:
  Objective  The rapid development of deep learning has significantly advanced object detection in remote sensing images (RSIs).Due to constraints imposed by imaging conditions and object size, a large number of tiny objects (with a pixel area of less than 16×16) are widely distributed in RSIs.However, compared with satisfactory detection performance on normal-scaled objects, current object detection methods exhibit a significant performance gap for tiny objects.This is primarily attributed to the two critical limitations: insufficient positive label assignment and weak feature representations.To address these issues, a novel Cross-Domain Collaborative Enhancement Detector (CDCEDet) is proposed.The CDCEDet jointly optimizes the label assignment in the spatial domain and strengthens the feature representations in the frequency domain for tiny objects, thereby establishing an accurate and robust framework for remote sensing tiny object detection.  Methods  The overall framework of the proposed CDCEDet is depicted in Fig.2 and comprises three key components.First, a Scale-Adaptive Anchor Generator (SAAG) is devised to dynamically generate anchors matched with the scales of ground-truth (GT) objects, effectively mitigating the scale mismatch issue neglected in previous works.Compared to conventional uniform-scale anchor generators, the proposed SAAG significantly increases positive samples for tiny objects, even under the IoU-threshold-based label assignment.Second, a Quantile-based Adaptive Label Assignment (QALA) mechanism is proposed to replace the fixed IoU threshold label assignment.The QALA exploits the IoU distribution between each GT object and its matched anchors to dynamically generate an adaptive threshold tailored to each GT object, thereby further boosting positive samples for tiny objects.Third, a Frequency-Adaptive Fusion (FAF) module is constructed to enhance feature representations from a frequency-domain perspective.On the one hand, adaptive high-pass filters are employed to highlight high-frequency information, which mitigates the information loss during channel compression.On the other hand, adaptive low-pass filters are applied to smooth and up-sample high-level features, thereby alleviating the feature inconsistency within the up-sampled objects.  Results and Discussions  Extensive experiments are conducted on two publicly available remote sensing tiny object detection datasets, AI-TODv2 and AI-TOD-R, by comparing the proposed method with several latest approaches, including RFLA, DCNet, and DFCL.On the AI-TODv2 dataset (Table 1), the proposed CDCEDet achieves performance gains of 1.8% and 0.7% in terms of AP50 and AP50-95, respectively, compared to the latest method.On the AI-TOD-R dataset (Table 2), AP50 and AP50-95 are improved by 2.6% and 0.7%, respectively.These results demonstrate that the proposed CDCEDet possesses superior performance and robust generalization capabilities for tiny object detection in RSIs.Ablation studies and configuration analyses of the CDCEDet core modules (Table 3, Table 4, Table 5, and Table 6) further verify the effectiveness of each proposed component and their complementary nature.Qualitative evaluations on both datasets (Fig.3) confirm that the proposed method accurately detects tiny objects in both sparse and dense distributions scenarios.Fig.4 confirms that the proposed SAAG generates scale-matched anchors for each object and assigns more positive samples to tiny objects than widely used uniformly distributed anchor generator.Furthermore, in comparative visualizations with RFLA and DCNet (Fig.5), CDCEDet exhibits superior detection accuracy and demonstrates a more effective reduction in missed detections.  Conclusions  This paper proposes a novel CDCEDet to address the insufficient label assignment and weak feature representations in remote sensing tiny object detection.Specifically, a SAAG is proposed to generate scale-matched anchors tailored to each GT object, significantly increasing the number of positive samples for tiny objects.Additionally, a QALA mechanism is devised that exploits the IoU distribution between GT objects and their matched anchors to dynamically adjust the assignment threshold, thereby effectively mitigating the scale bias induced by the fixed threshold.Finally, a FAF module is constructed from a frequency-domain perspective, which adopts the adaptive high-pass filters and low-pass filters to strengthen the feature representations for tiny objects.Extensive experiments on two tiny object detection benchmarks demonstrate the superior performance and robust generalization of the proposed CDCEDet.Future research will focus on enhancing model efficiency and real-time performance to facilitate the practical deployment of the proposed approach in remote sensing scenarios.
Infrared Small Target Detection Enhanced by Multi-dimensional Fusion Attention
LI Weixing, WANG Shuai, CHEN Huaiyu, SHENG Weidong
Available online  , doi: 10.11999/JEIT260040
Abstract:
  Objective  Infrared imaging boasts advantages such as long operating range, wide coverage, strong concealment, and all-weather visibility, making it widely applicable in aerospace target surveillance, maritime emergency rescue, forest fire monitoring, and earth remote sensing. Infrared payloads are typically mounted on platforms like satellites and aircraft, capturing images over long distances where targets appear small in the imagery and lack distinct texture and morphological features. Due to weak thermal radiation signals, targets are prone to being submerged in strong clutter. Current infrared small target detection networks face several key challenges. First, target segmentation networks expand the receptive field through continuous down-sampling, which can cause the features of infrared small targets to be easily lost in the deeper layers of neural networks. Second, features in the deeper neural network layers tend to diffuse. Therefore, under the constraints of long-range imaging and complex backgrounds, achieving high-precision and highly reliable infrared small target detection remains a research hotspot in the field of infrared imaging.  Methods  A channel-spatial Multi-Dimensional Fusion Attention Mechanism (MFAM) is proposed to address the challenges of small target size and feature diffusion in deep networks. Multi-dimensional features across channels, height, and width are captured, as well as the dependencies across these dimensions. By integrating feature fusion and cross-dimensional synergistic interaction, feature diffusion of small targets in deep networks is mitigated and the robustness of infrared small-target detection is enhanced. The channel attention and spatial attention modules are applied in parallel directly to the input feature maps, performing feature extraction and local information interaction along the channel and spatial dimensions, respectively. In the channel domain, features are compressed and processed through a Multi-layer Perceptron (MLP) to capture inter-channel relationships. In the spatial domain, Global Average Pooling (GAP) and Global Maximum Pooling (GMP) are employed to encode width and height dimensions. Finally, feature fusion is achieved via a Sigmoid activation function. Compared with traditional hybrid attention mechanisms, the channel attention module and spatial attention module are applied in parallel directly to the input feature maps. This approach enables refined encoding of the channel, height, and width dimensions under a global receptive field. Meanwhile, shallow features with spatial details and deep semantic features rich in contextual information are captured. The proposed MFAM features a plug-and-play design, allowing flexible integration into various baseline models such as ResNet and DNA-Net without introducing complex additional structures, which demonstrates excellent compatibility and generalization capability.  Results and Discussions  The publicly available dataset NUDT-SIRST is utilized to conduct a performance analysis of the proposed algorithm. By incorporating the proposed MFAM into the baseline DNA-Net, the Intersection over Union (IoU), detection rate (Pd), and false alarm rate (Fa) are 87.34%, 98.72% and 3.22×10–6, respectively. Compared to CBAM, BAM, GAM, CA, and TA, MFAM improves IoU by 0.4%, 1.68%, 2.46%, 2.11%, and 1.61%, respectively(Table.1). CBAM places the spatial attention module in series of the channel attention, which may lead to information loss during feature propagation. CA and BAM only employ GAP for information encoding, overlooking the role of max pooling in deep feature extraction. TA realizes cross-dimensional dependencies through rotation operations but fails to achieve simultaneous three-dimensional interaction. MFAM simultaneously integrates channel and spatial information, enabling sufficient cross-domain interaction, and combines global average and max pooling to enhance context and detailed feature extraction capabilities, thereby achieving more refined infrared small target detection. MFAM is also embedded into ALC-Net and AMFU-Net, resulting in IoU improvements of 0.43% and 0.28% compared with CBAM(Table.2). Through ablation experiments, detection performance is compared under different attention combination strategies. Compared to the serial fusion approach, the proposed parallel structure achieves improvements of 0.37% and 0.39% in IoU and Pd, respectively, while reducing Fa by 1.41×10–6. The advantage stems from the parallel structure applying channel and spatial attention directly to the input features, mitigating the diffusion of target features in deep networks. To validate the inference performance of the proposed algorithm on edge-side processors, a verification system is designed by using FPGA and NVIDIA Jetson AGX Xavier. Practical testing confirms that the proposed algorithm can be deployed and inferred on edge-side GPUs, with an average single-frame inference latency of 46.7 ms (Fig. 8) for 256×256 input images.  Conclusions  A Multi-dimensional Fusion Attention Module (MFAM) is constructed in this paper. The MFAM module achieves effective aggregation of channel and spatial salient features and enables cross-dimensional adaptive interaction, enhancing the preservation of target features in deep networks and ensuring robust output. The MFAM module exhibits favorable plug-and-play characteristics. Experiments demonstrate that the proposed algorithm performs better in metrics of IoU, Pd, and Fa. A lightweight intelligent processing unit based on FPGA+GPU is developed, which successfully achieves deployment of the algorithm on edge devices and enables high real-time inference, demonstrating promising engineering applicability and future application prospects.
Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism
ZHANG Xin-di, CHEN Hui, ZHANG Hong-yun, LIAN Feng, ZHANG Guang-hua, YIN Zhi-peng
Available online  , doi: 10.11999/JEIT260045
Abstract:
  Objective  Multi-target tracking requires sensor management to adapt the observation process to clutter, missed detections, target-number variations, and changes in target motion. Conventional methods often search over a finite set of sensor actions, resulting in increasing computational cost and limited control resolution. Moreover, rewards formed by combining several single-target quantities may not adequately represent the joint multi-target posterior. A continuous-action sensor management method integrating the Twin-Delayed Deep Deterministic Policy Gradient algorithm with the Labeled Multi-Bernoulli filter is therefore developed to optimize mobile-sensor heading according to the multi-target belief state.  Methods  The LMB posterior, including target existence probabilities and state densities, is used to construct the belief state. At each filtering step, the mobile sensor selects a continuous heading angle that affects the sensor-target geometry, detection probability, and LMB update. Predicted target states and candidate headings are used to generate ideal measurements and obtain pseudo-updated LMB densities. The Cauchy-Schwarz divergence between the predicted and pseudo-updated densities is used to construct the information-gain reward. TD3 employs two critics, target policy smoothing, and delayed actor updates to reduce value-estimation bias. Random control, policy-gradient control, an information-driven discrete method, and DDPG-based continuous control are used for comparison.  Results and Discussions  DDPG-LMB and TD3-LMB produce more continuous steering changes than the discrete-action methods (Fig. 2). TD3-LMB obtains the highest or near-highest detection probabilities for most targets (Fig. 3) and gives relatively large Cauchy-Schwarz divergence values during most time steps, while random control remains relatively low (Fig. 4). TD3-LMB also achieves the lowest overall OSPA in the tested scenario, with DDPG-LMB generally outperforming the discrete baselines (Fig. 5). These results show that continuous heading control improves observation quality and overall LMB tracking performance.  Conclusions  A TD3-based continuous-action sensor management framework for the LMB filter is presented. Candidate heading actions are evaluated through pseudo-updated LMB densities and Cauchy-Schwarz divergence, linking action selection directly to the joint multi-target posterior. The simulation results show more continuous sensor motion, higher detection probabilities for most targets, larger information gain, and lower OSPA in the tested scenario. Future work will consider higher-dimensional actions and cooperative multi-sensor management.
A Behavioral Economics-Based Game Model for Side-Channel Security Attack and Defense Strategies
CAI Juesong, YAN Yingjian, WANG Jindong
Available online  , doi: 10.11999/JEIT260121
Abstract:
  Objective  The field of side-channel security currently lacks a systematic, quantifiable, and reproducible model for guiding the selection of attack and defense strategies, particularly in real-world engineering contexts where resource constraints necessitate informed cost-benefit trade-offs. The absence of such a framework impedes the practical realization of the “appropriate security” principle, often leading to either over-protection or under-protection of cryptographic modules. Traditional approaches to evaluating attack and defense costs rely heavily on subjective expert judgments, which are inherently arbitrary, difficult to replicate, and lack a structured multi-dimensional assessment. To bridge this critical gap, this research proposes an interdisciplinary model that integrates game theory, behavioral economics, and the Analytic Hierarchy Process (AHP). The primary objective is to establish a holistic decision-support system that not only quantifies the multi-faceted costs of various side-channel strategies but also incorporates the psychological dimensions of decision-making under risk, thereby enabling dynamic and economically rational security strategy selection tailored to specific asset values and security levels.  Methods  This study constructs a multi-layered modeling framework based on a static non-cooperative game with incomplete information. First, the side-channel analyst and the defense designer are formally defined as rational players, each possessing a finite set of strategies: the attacker may choose from non-modeling attacks such as DPA/CPA, modeling-based attacks like template attacks, or emerging deep learning-based side-channel analysis; the defender may adopt countermeasures including time hiding, amplitude hiding, or masking/blinding techniques. To systematically quantify the often-overlooked cost dimension, an AHP-based structured cost model is introduced. Through pairwise comparison matrices, the model decomposes costs into multiple criteria—such as time, data storage, computational resources, expertise, and hardware overhead—and assigns objective weights to each criterion, thereby replacing subjective cost estimates with a reproducible, hierarchical evaluation system. Furthermore, to reflect real-world decision-making behavior, key concepts from behavioral economics are integrated: Prospect Theory models how gains and losses are perceived relative to a reference point, while risk aversion coefficients capture players’ tolerance for uncertainty. These behavioral parameters are explicitly linked to the security level of the cryptographic module, allowing the model to adapt to different operational contexts. The resulting behavioral-augmented Bayesian game is then solved using the concept of Bayes-Nash Equilibrium, wherein each player’s optimal mixed strategy is derived based on their private type (behavioral profile) and beliefs about the opponent. To ensure engineering relevance, the As Low As Reasonably Practicable principle is incorporated as a constraint, enforcing that any selected defense strategy must be justifiable in terms of risk reduction versus cost incurred. Numerical solutions are obtained via a customized sequential quadratic programming algorithm implemented in Python.  Results and Discussions  A comprehensive experimental evaluation was conducted to validate the proposed model’s consistency, sensitivity, and practical utility. The AHP-based cost quantification demonstrated strong internal consistency, with all consistency ratios below the 0.1 threshold, confirming the reliability of the judgment matrices. The derived weight distributions revealed intuitive priorities: attackers placed greater emphasis on technical barriers and computational cost, whereas defenders prioritized design complexity and performance overhead. The behavioral adjustment layer successfully modulated perceived costs according to security levels: under low-security conditions (high risk aversion), costs were perceptually inflated, leading to conservative strategy choices; under high-security conditions, decision-making aligned more closely with objectively quantified costs. Equilibrium analysis across varying asset values and security levels yielded interpretable and rational strategy profiles. For low-value assets, both players exhibited a strong tendency toward low-cost or “no action” strategies, adhering to the lower bound of the ALARP region. As asset value increased, a clear threshold effect was observed, triggering a shift toward high-cost, high-efficacy strategies such as deep learning-based attacks and masking-based defenses. Sensitivity analysis further confirmed that defense strategy probabilities increased monotonically with asset value, validating the model’s ability to capture the non-linear relationship between protection intensity and asset criticality. These findings underscore the model’s capacity to support context-aware, adaptive security decision-making that balances risk, cost, and psychological factors.  Conclusions  This research presents a novel, behaviorally informed game-theoretic model for side-channel security strategy selection, addressing a significant void in existing literature regarding structured cost-benefit assessment. By integrating AHP-based objective cost quantification, behaviorally adjusted subjective valuations, and ALARP-driven engineering constraints, the proposed framework offers a multi-dimensional, reproducible, and context-sensitive tool for analyzing attack-defense interactions. The model advances the field by explicitly linking security levels to behavioral parameters, enabling dynamic strategy adaptation in response to both asset value and decision-makers’ risk perceptions. Although the current implementation relies partially on expert-defined parameters and operates within a static game setting, it establishes a critical foundation for transitioning from heuristic-based security decisions to quantitatively grounded, interdisciplinary analysis. Future work will focus on parameter calibration using real-world attack/defense datasets, extension to multi-stage dynamic games to capture strategic evolution over time, and empirical validation in industrial cryptographic evaluation scenarios. This study contributes the first systematic methodology for cost-aware, behaviorally realistic strategy optimization in side-channel security, offering both theoretical insights and practical guidance toward achieving “appropriate security” in cryptographic engineering.
Research Status and Prospects of Mid-Wavelength Infrared Superlattice Detector Technology
LIU Ming, ZHAO Yaqi, GUAN Xiaoning, ZHANG Fan, LU Pengfei
Available online  , doi: 10.11999/JEIT260083
Abstract:
  Significance   Mid-Wavelength Infrared (MWIR) detectors are widely used in civilian and military applications because of their high sensitivity and excellent temperature discrimination. Type-II SuperLattice (T2SL) materials, especially the InAs/GaSb and InAs/InAsSb systems, have become promising candidates for third-generation infrared photodetectors. This review systematically analyzes the research status and future trends of MWIR T2SL detector technology. It focuses on key photoelectric parameters, including Quantum Efficiency (QE), dark current density, and Specific Detectivity (D*). This work provides a reference for material selection and performance optimization in this rapidly developing field.  Progress   Considerable progress has been made in dark current suppression and photoresponse enhancement for MWIR T2SL detectors. For dark current suppression, advanced barrier structures, such as nBn, XBn, and M-structures, are designed through band-structure engineering. These structures effectively block majority-carrier transport while allowing efficient collection of photogenerated carriers. For instance, an nBn device with an AlAsSb/InAsSb superlattice barrier shows a dark current density of 2.01×10–5 A/cm2 at 150 K (Fig. 2(c)). Strain compensation and optimized epitaxial growth further reduce bulk dark current. One device achieves a dark current density of 4.5×10–7 A/cm2 at 140 K (Fig. 4(f)). Device process optimization, including two-step etching and Zn-diffusion-based planar junction formation, also reduces surface leakage current (Fig. 5, Fig. 6). For photoresponse enhancement, the main strategies include micro/nano-optical structure integration, epitaxial growth optimization, and device process improvement. Monolithically integrated metalenses increase th,e peak responsivity to 9.01 A/W at 300 K (Fig. 7(d)). Guided-mode resonance architectures enable a room-temperature External Quantum Efficiency (EQE) of approximately 60% (Fig. 8(c)). Epitaxial optimization, including stepped absorption layers and interfacial graded doping, increases the QE to 59.4% at 150 K (Fig. 10(c)). Device process optimization, such as substrate removal and Anti-Reflection (AR) coating deposition, also improves QE. An average QE of 63.7% is reported in the 3.7~4.8 μm range (Fig.13(c)). Comparative analysis shows that InAs/GaSb detectors are mainly reported at 77~150 K, whereas InAs/InAsSb detectors show stronger potential for higher-temperature operation, especially near 150 K (Fig. 15, Fig. 16). Overall, at 150K, dark current densities are generally suppressed below 10–4 A/cm2, and peak QEs approach 70%.  Conclusions  T2SL materials, with tunable band structures and low Auger recombination rates, have become a core material platform for high-performance MWIR detection. Current studies have addressed key challenges in dark current suppression and photoresponse enhancement. Through advanced barrier design and device process optimization, dark current densities have been suppressed to the 10–6 A/cm2 level at approximately 150 K. Through optical and epitaxial engineering, QEs have been increased to approximately 60% or higher. The InAs/InAsSb material system is particularly promising for High-Operating-Temperature (HOT) applications.  Prospects  Future development will focus on four main directions. First, the HOT limit should be further increased, with the goal of maintaining diffusion-limited performance at 180 K or higher. Second, large-format Focal Plane Arrays (FPAs) should be developed based on highly uniform material growth through mature Molecular Beam Epitaxy (MBE), aiming for pixel operability higher than 99%. Third, multicolor and multispectral detection should be expanded by precisely tuning superlattice periods, enabling integrated dual-band or multiband MWIR detection with reduced crosstalk. Fourth, new device architectures and coupled physical mechanisms should be explored to extend detector performance and application boundaries.
Optimizing SATisfiability-Based Automatic Test Pattern Generation Systems: Unified Fault Set Construction,Modeling, and Solving
YAN Dapeng, HE Qirun, GUO Jing, WANG Boning, CAI Zhikuang
Available online  , doi: 10.11999/JEIT260025
Abstract:
  Objective  Boolean SATisfiability-Based Automatic Test Pattern Generation (SAT-Based ATPG) is widely used to generate tests for hard-to-detect single stuck-at faults and to prove fault untestability in combinational logic. When SAT-Based ATPG is applied to large netlists with dense fanout and reconvergence, its runtime and memory consumption are often dominated by three interacting issues. Representative fault lists produced by conventional dominance- or equivalence-based fault collapsing can remain large, increasing the number of SAT calls and enlarging the incremental context that must be maintained across faults. Meanwhile, SAT modeling may introduce redundant Conjunctive Normal Form (CNF) overhead, especially when an explicit faulty-circuit copy is constructed or when propagation constraints are encoded globally without locality control. In addition, fanout-reconvergence structures amplify assignment correlations along sensitized paths, and such correlations are often exposed only after repeated decisions and backtracking when only standard unit propagation is used. The unified optimization objective is therefore to reduce overall CNF size and solving cost while preserving completeness, so that a practical SAT-Based ATPG system remains efficient and stable across circuits of different scales.  Methods  A three-part framework is developed and implemented in an incremental SAT-Based ATPG flow, and the overall workflow is illustrated (Fig. 1). First, a checkpoint-driven dynamic fault-set construction method is proposed. Checkpoints are collected during netlist-to-directed-acyclic-graph conversion, including all primary inputs and all fanout branches, and XOR/XNOR outputs are additionally recorded as supplementary checkpoints to avoid over-collapsing XOR-related fault behavior. Representative faults are initialized on checkpoints by compact rules that combine dominance-oriented fault collapsing with equivalence-aware refinement, and solver-guided repair is performed when an untestable representative fault indicates potential masking under structural constraints. The procedure is summarized in Algorithm 1. Second, a SAT modeling method based on fault sensitization constraints is adopted to avoid explicit faulty-circuit duplication. Fault activation, propagation, and observability are represented by additional fault sensitization constraints over the original circuit variables, and auxiliary variables are introduced only when local bookkeeping is required. Constraint localization is restricted to the fault fanout cone, and cone-boundary and internal vertices are identified through a graph-traversal procedure (Fig. 2). Third, a dynamic implication learning mechanism oriented to fanout-reconvergence pairs is integrated into the incremental solving loop. Reconvergence pairs within the fault fanout cone are monitored under partial assignments, and structure-induced implications are injected either as implied assignments when a reconvergent output becomes functionally determined or as short conflict clauses when a branch-value combination becomes inconsistent with the fault sensitization constraints. The dynamic implication learning procedure is summarized in Algorithm 2.  Results and Discussions  The unified system is evaluated on ISCAS’85 and ISCAS’89 benchmark circuits, with TG-PRO used as the baseline implementation under the same SAT solver and termination settings. The checkpoint-driven dynamic fault-set construction method substantially reduces the representative fault space entering ATPG. Relative to the uncollapsed fault space, the average representative-fault ratio decreases from 51.38% to 42.41%, corresponding to an average fault-space reduction of 57.59%. The best-case ratio reaches 33.19% on large circuits with heavy reconvergence, which indicates that checkpoint-centered representative-fault allocation effectively suppresses redundancy without enlarging the untestable fault set (Table 1). The reduced fault-set size is reflected in preprocessing efficiency, and the total runtime for fault-set construction is consistently reduced, with an average reduction of 8.37% across the evaluated circuits (Fig. 3). For SAT model construction, the fault-sensitization-constraint encoding reduces CNF overhead relative to the baseline model construction. Across the benchmark set, the numbers of CNF clauses and CNF variables are reduced by 11.44% and 3.50%, respectively, which shows that avoiding explicit faulty-circuit duplication and localizing auxiliary constraints to the fault fanout cone effectively lowers memory demand (Table 2). The reduced CNF size and strengthened locality of constraints are further reflected in end-to-end runtime, and the total runtime of SAT modeling and solving is reduced across the evaluated benchmarks (Fig. 4). Dynamic implication learning further improves solving efficiency in reconvergence-heavy structures. Compared with static implication learning, CNF construction time increases by 3.0% on average because of the additional monitoring and injection operations, yet the overall runtime decreases by 4.42% on average, which indicates a favorable cost-benefit trade-off. The overhead attributed to dynamic implication learning accounts for 2.51% of the total runtime aggregated across circuits, which confirms that the injected implications and pruning clauses provide measurable solving benefits at limited extra cost (Table 3).  Conclusions  A unified optimization framework for SAT-Based ATPG is developed by combining checkpoint-driven dynamic fault-set construction, localized fault sensitization constraints for CNF modeling, and fanout-reconvergence-oriented dynamic implication learning. Representative faults are compressed through solver-guided repair of dominance and equivalence relations to avoid masking, CNF growth is controlled through duplication-free modeling localized to the fault fanout cone, and reconvergence correlations are exploited through incremental implication injection to strengthen propagation and enable early conflict pruning. Experimental results on standard benchmark circuits show consistent reductions in representative fault scale, CNF size, and total runtime, providing a practical approach for scaling SAT-Based ATPG to larger designs with complex fanout and reconvergence.
A Review of Advances and Challenges in Intelligent Disaster Assessment of High-Value Objects in Remote Sensing Imagery
ZHANG Yidan, FENG Yingchao, WANG Tianqi, LIU Yu, WANG Mengyu, HOU Zhongyan
Available online  , doi: 10.11999/JEIT251297
Abstract:
  Significance   The rapid and accurate disaster assessment of objects in the aftermath of disasters is critical for enabling effective emergency response and post-disaster recovery operations. Intelligent remote sensing technologies, empowered by deep learning, provide scalable, objective, and efficient means of assessing disaster across large and complex environments, such as densely populated urban centers, transportation hubs, and critical energy infrastructure. By leveraging high-resolution satellite and aerial imagery, these technologies can provide timely situational awareness for decision-makers, supporting prioritization in rescue and recovery tasks. Despite significant advancements in methods and applications, there remains a lack of comprehensive synthesis in the field, leading to fragmented practices and inconsistent benchmarks across studies. This study addresses this critical gap by providing a structured and systematic review that consolidates technical foundations, commonly used datasets, evaluation metrics, and methodological advances in deep learning-based remote sensing disaster assessment. The overarching goal is to accelerate the adoption and operationalization of these technologies in real-world disaster scenarios, ultimately contributing to the construction of resilient cities and infrastructures under increasing environmental and geopolitical risks.  Progress   In recent years, deep learning-based remote sensing disaster assessment methods have developed rapidly, demonstrating notable improvements in classification accuracy, processing automation, and scalability. Significant advances include bi-temporal change detection methods that can precisely localize disaster by capturing differences before and after disaster events, and multi-temporal sequence modeling approaches that extract evolving disaster patterns and degradation trends across time-series data. Furthermore, multi-modal data fusion strategies that combine optical, Synthetic Aperture Radar(SAR), and Light Detection and Ranging (LiDAR)data have enhanced the ability to analyze complex disaster characteristics under varying observation conditions. Advanced techniques to address data scarcity, such as transfer learning and self-supervised learning, have further extended the applicability of disaster assessment methods in data-constrained environments. These advances collectively contribute to improving the responsiveness and effectiveness of disaster assessment systems in supporting emergency decision-making and resource allocation during critical disaster events.  Conclusions  This study systematically categorizes and evaluates the technical landscape of deep learning-based remote sensing disaster assessment for objects, highlighting the strengths and limitations of bi-temporal, multi-temporal, multi-modal, and data-scarce scenario methods. While existing approaches demonstrate considerable potential in addressing various challenges in post-disaster assessment, challenges remain in ensuring robustness across diverse environments and operational conditions. The review underscores the need for standardized disaster classification criteria and comprehensive evaluation frameworks that consider both physical disaster and functional impacts to facilitate practical and consistent disaster assessments in real-world deployments.  Prospects   Future research in remote sensing-based disaster assessment should focus on developing multi-level collaborative frameworks for evaluating diverse objects across spatial and functional scales, enabling holistic disaster impact assessment that captures both direct and cascading effects. Complex scenarios such as airports, industrial zones, and ports contain static structures, moving objects, and interdependent functional units, thus requiring hierarchical modeling and multi-object reasoning. Physics-driven and hybrid approaches integrating structural mechanics, material degradation, and expert knowledge can further improve interpretability and generalization. Meanwhile, lightweight model design and edge deployment are important for real-time assessment on drones and satellites in emergency situations. Standardized evaluation metrics that combine physical disaster mechanisms with functional impact analysis will also be essential for practical deployment. Together, these directions will help transform intelligent remote sensing technologies into actionable tools for disaster response and recovery.
UWF-YOLO: A Lightweight Framework for Underwater Object Detection via Redundant Information Optimization
HOU Guojia, MA Jiaqi, WANG Yuechuan, HUANG Baoxiang, LI Kunqian
Available online  , doi: 10.11999/JEIT251129
Abstract:
  Objective  The rapid development of underwater imaging technology has increased the significance of underwater object detection for resource exploration and environmental monitoring. Complex underwater environments often degrade image quality through color casts, haze-like effects, and non-uniform illumination. These factors reduce the performance of existing vision-based object detection algorithms, particularly for small objects, and often lead to missed detections and false positives. In addition, current deep learning–based underwater detection models face difficulty balancing detection accuracy and lightweight design under limited computational resources. Therefore, efficient underwater object detection methods are required for water-related vision tasks. Such methods support marine resource exploration, ecological monitoring, underwater robotics, and perception systems for autonomous underwater vehicles.  Methods  A lightweight framework based on redundant information optimization is proposed for underwater object detection. Specifically, a lightweight underwater object detection network, termed UWF-YOLO, is designed based on redundant information optimization. First, the C2f module is reconstructed using the FasterNet Block to optimize both the backbone and neck networks. A feature channel selection mechanism is integrated to reduce redundant feature representations. Furthermore, redundant convolutional features in the conventional YOLO neck limit adaptation to underwater environments. Therefore, Ghost Convolution is introduced to generate Ghost feature maps and improve the multi-scale feature fusion capability of the neck network. Next, parameter sharing is achieved by replacing the original detection head with a redundant optimization group detection head (RRG-Head) based on group convolution, which reduces computational cost. Finally, a structured channel pruning strategy is applied to identify inter-layer dependencies in the computational graph and bind pruning units. Combined with LAMP weight magnitude score normalization to evaluate channel importance, low-contributing groups are pruned and subsequently fine-tuned to compress the network size. In addition, existing underwater detection datasets usually contain monotonous scenes, and the objects are typically small and densely distributed. To address this limitation, an underwater object detection dataset with complex scenes, termed CSUOD, is constructed by collecting real-world underwater images from various websites and platforms. Manual annotation and resolution normalization are then performed to ensure dataset consistency. CSUOD is designed for challenging underwater environments characterized by color casts, haze-like effects, and non-uniform illumination. A total of 1 135 images containing six object categories are manually selected and annotated.  Results and Discussions  Extensive experiments are conducted on three public underwater object detection datasets, namely DUO, RUOD, and TrashCan, and several widely used detection methods are compared. The proposed model is evaluated against mainstream detectors, including YOLOv5s, YOLOv7-tiny, YOLOv8s, YOLOv9-tiny, and Deformable DETR. In terms of computational complexity, the proposed method reduces FLOPs, model size, and parameters by 60.4%, 77.3%, and 78.4%, respectively, compared with the baseline model. Furthermore, the proposed method outperforms YOLOv9-tiny with comparable parameters by 0.3%, 2.3%, and 3.4% in mAP on the three datasets. Additional comparative experiments on the constructed CSUOD dataset also demonstrate improved performance and stable detection capability in complex underwater environments. Qualitative visualization results further demonstrate the robustness and detection stability of the model under various underwater degradations, including haze-like effects and non-uniform illumination.  Conclusions  Quantitative and qualitative experiments on multiple datasets validate the effectiveness and robustness of the proposed method. The proposed framework achieves superior detection performance in complex underwater environments and reduces missed detections and false positives caused by background interference. Experimental results indicate that the proposed UWF-YOLO achieves significant model lightweighting while maintaining detection accuracy comparable to benchmark models. This balance between detection accuracy and low computational cost makes the framework suitable for underwater devices with limited resources. The proposed method also shows strong potential for practical applications such as marine ecological monitoring, underwater resource exploration, and perception systems for autonomous underwater vehicles. It provides a reliable technical foundation for real-time applications, supports integration into embedded platforms, and enables real-time perception and decision-making under different underwater conditions. In addition, the constructed CSUOD dataset helps address the limitations of existing underwater detection datasets and supports further research in underwater object detection. Future work will extend this framework to multi-modal perception systems and larger-scale datasets, enabling adaptive models for dynamic underwater scenarios and supporting broader applications in intelligent ocean observation and autonomous navigation.
An Overview of Key Technologies for 6G-Enabled Communication-Computing Integration and Energy-Efficiency Optimization
LIU Guangyi, CAI Qing, WANG Xinyao, CHEN Tianjiao, JIN Jing, XUE Yahui, WANG Ailing, WANG Hanning
Available online  , doi: 10.11999/JEIT260399
Abstract:
  Significance   Constrained by size, power consumption, and cost, emerging intelligent terminals often face excessive energy consumption and limited battery life. These limitations have become major bottlenecks to large-scale deployment. Compared with Fifth-Generation (5G) wireless networks, Sixth-Generation (6G) wireless networks are expected to enhance the Radio Access Network (RAN) architecture and move computing capability toward the RAN side. High-energy and compute-intensive Artificial Intelligence (AI) tasks that are originally executed by end devices can therefore be processed by the network. Through End-Edge Collaboration, emerging intelligent terminals can be upgraded toward lightweight design, low cost, and long battery life, thereby supporting the large-scale deployment of ubiquitous intelligence in 6G networks.  Progress   Current progress in Terminal Energy Consumption Optimization through 6G End-Edge Collaboration is reviewed, with emphasis on local execution, full offloading, and partial offloading. In local execution, User Equipment (UE) processes all tasks locally, which leads to high computing energy consumption. In full offloading, all tasks are transferred to the RAN. This reduces terminal-side computing energy consumption but can increase transmission energy consumption, especially under poor channel conditions. Partial offloading combines the benefits of both modes and optimizes energy consumption according to real-time network conditions. For partial offloading, four representative optimization techniques are summarized. (1) Feature Extraction and Filtering. Semantic encoding and information extraction are performed at the UE, and only task-relevant data are transmitted to the RAN. This reduces redundant data transmission and lowers transmission energy consumption. (2) Split Offloading. A large Deep Neural Network (DNN) is divided into layers according to its structure. Simpler shallow layers are processed at the UE, whereas more complex deep layers are offloaded to the RAN. This method balances terminal-side and RAN-side computational loads through End-Edge Collaborative Inference. (3) Model Lightweighting. Model complexity is reduced through pruning, quantization, and knowledge distillation, which lowers computational overhead while maintaining task performance. (4) Incremental Inference. Only changed data or updated features are processed, while historical computations are reused. This reduces redundant computation. Together, these techniques improve terminal performance and energy efficiency within the 6G End-Edge Collaboration framework.  Conclusions  This paper systematically reviews Terminal Energy Consumption Optimization for 6G End-Edge Collaboration. It summarizes the functional evolution of enhanced RAN, constructs an End-Edge Collaborative service framework for Communication-Computing Integration, and establishes a theoretical model of terminal computing energy consumption and transmission energy consumption. The composition and influencing factors of energy consumption under different offloading modes are clarified. Key energy optimization technologies, including Feature Extraction and Filtering, Split Offloading, Model Lightweighting, and Incremental Inference, are then discussed. To address energy consumption fluctuations caused by dynamic wireless channels, the paper proposes energy optimization mechanisms based on Adaptive Semantic Compression, Dynamic Split Offloading, Adaptive Model Pruning, and Incremental Inference. These mechanisms maintain a dynamic balance between energy optimization and task performance. Using embodied intelligent robot video understanding as a typical application scenario, a test platform is developed to verify the effectiveness of the proposed mechanisms. Current challenges and future research directions are also analyzed.  Prospects   Although End-Edge Collaborative energy-saving technologies have achieved initial progress, practical deployment still faces challenges in real network environments, dynamic wireless channels, and large-scale user access. Future research should examine the trade-off between optimization overhead and system robustness. It should also study dynamic communication-computing resource substitution modeling in stochastic resource environments, multi-user collaboration strategies, and global energy-efficiency optimization. As the technology matures, standardization and engineering implementation of End-Edge Collaborative energy-saving frameworks will become critical to the large-scale adoption of 6G applications. Future studies should further integrate algorithm design with network architecture, enabling practical deployment of low-power and high-efficiency intelligent communication systems.
Gating Adaptive Repeat Query Framework for Reliable Collaborative Inference with Edge Heterogeneous LLMs
WANG Tengsheng, YU Tao, LI Jihong, ZHENG Guhan, ZHANG Shunqing
Available online  , doi: 10.11999/JEIT260218
Abstract:
  Objective  Reliable inference at the network edge is indispensable for 6G-enabled ubiquitous AI, yet the deployment of large language models (LLMs) in such environments remains a cornerstone challenge. In resource-constrained edge settings, single-LLM inference often proves unreliable due to knowledge limitations and inherent biases, severely hampering real-world deployment. Collaborative inference leveraging multiple heterogeneous LLMs emerges as a promising remedy to boost robustness, but it introduces nontrivial hurdles under stringent latency and energy budgets, especially when wireless channel conditions and query content vary unpredictably. These challenges include the need for dynamic sequential decision-making for LLM selection and resource allocation, the fundamental paradigm mismatch between bit-level reliability protocols and semantic-level error correction, and the lack of adaptive mechanisms to align and fuse disparate LLM outputs effectively. To fill these critical gaps, this paper presents a novel framework that fundamentally reinterprets collaborative inference as a semantic-driven, closed-loop process, thereby transitioning from conventional bit-retransmission to semantic-retransmission and offering a practical path toward reliable 6G edge intelligence.  Methods  In response to these critical challenges, we propose the Gating Adaptive Repeat Query (G-ARQ) framework. Its core innovation is the Semantic-Space Alignment and Error-Guided Retransmission (SEMAR) mechanism. SEMAR first aligns the token-level probability distributions from heterogeneous LLMs into a unified semantic space using relative representation, enabling comparable outputs. It then models the collaborative process probabilistically, explicitly capturing error dependencies among models, and uses an Expectation-Maximization (EM) algorithm to infer a latent error direction, which guides the selection of the next LLM for query retransmission, steering it towards outputs orthogonal to previous errors. To jointly optimize the LLM gating and uplink power allocation under communication constraints without requiring explicit system dynamics—often unavailable in practice—we design a black-box trajectory optimizer. This optimizer formulates the sequential decision problem as sampling from a target distribution that encodes dynamic feasibility, optimality, and constraints. It employs a diffusion-based sampling process with a model-guided prior and Monte Carlo estimation to generate near-optimal policy trajectories that satisfy hard latency and energy limits.  Results and Discussions  To evaluate the practical viability of G-ARQ under realistic edge conditions, simulations are conducted in a scenario with five base stations hosting five heterogeneous 7B-parameter LLMs (Mistral-7B, Vicuna-7B, Nous-Capybala-7B, Gemma-7B, and Llama-2-7B). The user equipment (UE) performs a question-answering task evaluated on a mixed SQuAD and TriviaQA dataset. Component-level evaluations, each designed to isolate the contribution of a single innovation, validate the effectiveness of every key element. The error-guided gating of SEMAR, compared to a Top-k gating baseline, improves accuracy by 0.23 % on average, and its dynamic weight ensemble contributes an additional 0.7 % gain (Fig. 3). The black-box trajectory optimizer, which operates without any explicit channel model, achieves accuracy close to that of the unconstrained model-greedy strategy while ensuring strict latency constraints (Fig. 4). The convergence of the optimizer is verified by tracking the evolution of \begin{document}$ J(\boldsymbol{S}) $\end{document} and selection probability over diffusion steps (Fig. 5). System-level performance under varying latency and energy constraints demonstrates that G-ARQ consistently surpasses two baselines: one combining model-greedy selection with Proximal Policy Optimization (PPO) for power optimization, and another combining model-greedy selection with Simulated Annealing for power optimization, both using average output weights. The accuracy improvement is most significant under the most stringent resource limits, reaching up to 2.2 % for \begin{document}$ {K}_{\max }=1 $\end{document} and 1.9 % for \begin{document}$ {K}_{\max }=2 $\end{document} (Fig. 6, Fig. 7). The framework successfully establishes a Pareto boundary that characterizes the inherent trade-off between inference accuracy and communication latency, providing a valuable design guideline for resource-constrained edge systems and offering actionable insights for real-world deployment. The GARQ-S variant is noted to outperform GARQ-E by avoiding the integration of outputs from previously erroneous models.  Conclusions  This paper proposed G-ARQ framework, an innovative closed-loop framework that transforms collaborative edge inference into a semantics-guided retransmission process. By introducing SEMAR for error-based alignment and selection of heterogeneous LLMs, and employing a black-box trajectory optimizer for joint model selection and power allocation, the framework achieves up to a 2.2% accuracy improvement under strict resource constraints. The results validate G-ARQ as an effective and practical approach toward reliable and efficient 6G edge intelligence.
A Two-layer Closed-loop Cooperative Resource Allocation Framework for Improving QoS of MEC Network Slicing
XU Juntao, FAN Xinggang, XU Changfu, SHEN Minyang, LIANG Yuzhu, WANG Tian
Available online  , doi: 10.11999/JEIT260156
Abstract:
  Objective  Driven by 5G/6G networks, Multi-access Edge Computing (MEC) environments face major challenges in ensuring Quality of Service (QoS) for heterogeneous network slices while improving resource utilization. Existing resource allocation methods often lack dynamic adaptability in heterogeneous settings. They also fail to jointly optimize caching, bandwidth, and computing resources, which reduces resource utilization and service success rates. This paper addresses these limitations by proposing a robust framework for joint multi-dimensional resource optimization under strict QoS constraints. The framework provides a tailored solution for heterogeneous MEC environments.  Methods  The network slicing resource allocation problem is first formally defined. Its NP-hardness is proved by reducing the NP-hard multidimensional 0-1 knapsack problem to this problem. To solve this complex optimization problem, a cache-aware two-layer closed-loop cooperative framework, termed QCache, is proposed. The framework uses a synergistic “generate-evaluate-feedback” loop. In the upper Global Exploration Layer, a hybrid heuristic algorithm is designed by combining a population evolution strategy, including selection, crossover, and mutation, with a particle update mechanism guided by historical individual and global best positions. This layer broadly explores the solution space and generates high-quality candidate resource allocation schemes under complex constraints. Adaptive parameter adjustment and elite retention strategies are used to avoid local optima. In the lower Multi-Dimensional Weight Evaluation Layer, a quantitative assessment model is constructed. This model converts low-latency and high-bandwidth service demands into explicit QoS constraints by normalizing key performance indicators, including delay and rate, and by dynamically assigning weights through the entropy weight method. The weighted score reflects different slice priorities. The evaluated score (SCORE) from this layer is fed back to the upper layer as the fitness value, guiding the iterative evolution of candidate solutions until convergence.  Results and Discussions  Extensive simulations are conducted to validate the effectiveness of QCache against several baseline methods, including Genetic Algorithm (GA), PSO-Leader, GraphSAGE, and the No-Consideration-of-Service-Quality (NCSQ) scheme. Under identical resource and user demand scenarios, the overall comparison (Fig. 3) shows that QCache achieves the highest resource utilization rate of 80.30% and the best average service score of 1.109. Compared with the baseline methods, QCache improves resource utilization by 2.29% to 24.50% and increases user service scores by 4.13% to 59.34%. Experiments with varying total cache resources (Fig. 4) show that QCache maintains superior performance across different cache states. It improves resource utilization by 2.43% to 27.53% and service scores by 10.81% to 119.56%, confirming its cache-aware adaptability. Tests with increasing user numbers (Fig. 5) show that QCache scales effectively, achieving up to 85.83% resource utilization and a service score of 3.06. These results demonstrate its ability to handle dense access scenarios. Experiments with time-varying user demands (Fig. 6) further confirm the dynamic robustness of the framework. In these tests, QCache achieves average improvements of 9.20% in resource utilization and 23.45% in service score over the baselines.  Conclusions  This paper studies the NP-hard resource allocation problem in dynamic MEC environments with heterogeneous network slices. The proposed cache-aware two-layer closed-loop cooperative framework, QCache, jointly optimizes caching, bandwidth, and computing resources under explicit QoS constraints. The upper-layer hybrid heuristic provides strong global search capability. The lower-layer multi-dimensional weight model supports accurate QoS quantification and dynamic feedback. Comprehensive experimental results show that QCache outperforms existing methods in both resource utilization efficiency and user QoS satisfaction. Future work will explore reinforcement learning and traffic prediction mechanisms to further improve the response of the framework to bursty traffic and anomalous demands. This may support more intelligent and autonomous MEC network slice resource management.
Queue Stability Constrained Robust Secure Beamforming for Low-Altitude UAV-ISAC Systems
OUYANG Jian, REN Wei, XU Ba, LIU Xiaoyu, JIANG Wanmu
Available online  , doi: 10.11999/JEIT260275
Abstract:
  Objective  To address the challenges of antenna array angle errors caused by UAV jitter, transmission instability induced by random data arrivals, and secure transmission guarantee in multi-eavesdropper scenarios for low-altitude UAV-ISAC systems, this paper proposes a robust secure beamforming algorithm based on queue stability constraints. The proposed algorithms aims to minimize long-term transmit power consumption while maintaining data queue stability while enhancing beamforming robustness against UAV jitter.  Methods  To guarantee the stability of the wireless transmission in UAV-ISAC systems, this paper formulates an optimization problem aimed at minimizing the long-term average transmit power, subject to constraints on data queue stability, secure communication rate, sensing performance, and maximum transmit power. Since this long-term optimization problem is intractable, the Lyapunov optimization framework is employed to transform it into a sequence of short-term subproblems. To handle the antenna array angle errors caused by UAV jitter within each short slot, we jointly adopt the second-order Taylor series expansion and the S-Procedure method to approximate the short-term subproblem into a convex form. Consequently, a robust secure beamforming algorithm based on penalty successive convex approximation optimization is proposed.  Results and Discussions  Simulation results demonstrate the impact of the number of antennas, secrecy rate threshold, beam gain threshold, UAV jitter error, and Lyapunov weight factor on the system transmit power. As illustrated by the beam gain pattern in Fig. 3, the communication beamformer facilitates cooperative sensing toward the sensing area, while the sensing beamformer enhances system security by directing interference toward potential eavesdroppers. This validates the effectiveness of the proposed algorithm in simultaneously improving sensing and secure communication performance. Furthermore, leveraging the dual-function characteristics of the communication and sensing beamformers, the proposed integrated sensing and communication scheme achieves significantly higher resource utilization efficiency than the communication-only and sensing-only schemes, as shown in Fig. 5. Additionally, Fig. 6 indicates that, compared with the non-robust scheme, the proposed robust scheme strictly satisfies security requirements under various angle errors. Finally, Fig. 8 shows that the proposed queue-aware scheme can effectively suppress the transmit power fluctuations caused by random data arrivals, exhibiting superior stability compared to the queue-free baseline.  Conclusions  This paper investigates a robust secure beamforming method for UAV-ISAC systems subject to queue stability constraints. First, based on the Lyapunov optimization framework, the long-term stochastic optimization problem is transformed into a sequence of short-term subproblems. Second, to address the issue of UAV jitter, the second-order Taylor series expansion and the S-Procedure method are jointly employed to approximate the non-convex constraints into tractable convex forms. Finally, a robust secure BF optimization algorithm based on penalty successive convex approximation is proposed to efficiently solve the deterministic short-term subproblems. Simulation results demonstrate that the proposed scheme can effectively tackle the challenges posed by random data arrivals, UAV jitter, and eavesdropping threats, thereby ensuring the stability and security of downlink data transmission in low-altitude UAV-ISAC systems.
VT2R: Video and Text-driven Method for Generating Large-scale Millimeter-wave Radar Data
DENG Kaikai, LING Yue, XING Ling, WU Honghai, ZHAO Dong, MA Huahong
Available online  , doi: 10.11999/JEIT260240
Abstract:
  Objective  The lack of large-scale training data impedes progress in developing robust and generalized deep learning models. However, existing millimeter-wave radar data generation methods are ineffective due to a lack of sufficient data sources. To address this gap, this paper proposes a video and text-driven radar data generation method, VT2R, which utilizes video or text data to generate large-scale, realistic radar data, solving the key problem of constructing the mapping relationship between video and text and radar data.  Methods  The proposed method consists of three main components: video feature encoding network, text feature encoding network, radar feature encoding network and data fitting and decoding network. Video feature encoding networks and text feature encoding networks extract temporally consistent visual representations and alignable semantic features, respectively, while the radar encoding network learns the structure and dynamic information of point clouds through hierarchical spatiotemporal modeling. In the data fitting and decoding network based on Variational AutoEncoder (VAE), multi-modal features are mapped to a unified latent distribution space and decoded into radar data through reparameterized sampling. During training, reconstruction loss, Kullback-Leibler (KL) divergence loss, and cross-modal similarity loss are jointly optimized.  Results and Discussions  This paper constructs the first radar point cloud dataset for reclining gesture recognition (Figs. 6 and 7), covering 5 gesture categories, 32 participants, and a total of 14,400 samples. Experimental results based on this dataset show that VT2R achieves a recognition accuracy of 89.2% when trained using only generated radar data, a 33.88% improvement over the representative RFGen (Figs. 9 and 10). When combined with a small amount of real radar data for joint training, the accuracy further improves to 97.62%, a 21.48% improvement over RFGen (Figs. 9 and 11). Furthermore, VT2R still achieves average recognition accuracies of 89.35% and 97.21% under different scenarios and factors (Figs. 16-18). In addition, this paper also verifies the accuracy of VT2R under different postures, achieving average accuracies of 89.98% and 97.55% in the first and third settings, respectively (Fig. 19), which is basically consistent with the result obtained when lying down, demonstrating its robustness under cross-posture conditions.  Conclusions  This paper proposes a radar data generation system, VT2R, which addresses the severe lack of realistic radar training data when users are performing gestures in a lying position. Through a video feature encoding network built on a vision-language pre-trained model, a text encoding network incorporating cue templates, a hierarchical radar encoding network for sparse point clouds, and a VAE-based data fitting and decoding network, these components collaboratively generate large-scale, realistic radar data. It also supports augmented reconstruction based on limited real radar data, providing rich data support for radar perception tasks. Future work will focus on solving multi-modal data generation for more complex gesture scenes, providing better data support for emerging large-scale models.
Intelligent Resource Allocation Algorithm Based on Outdated CSI for Multi-Node URLLC
ZHAO Yizhen, GAO Wei, HU Yulin, ZHU Yao
Available online  , doi: 10.11999/JEIT260216
Abstract:
  Objective  Ultra-Reliable and Low-Latency Communications (URLLC) is widely used in Industrial Internet of Things (IIoT) systems. However, in mobile industrial scenarios such as transportation and inspection, instantaneous Channel State Information (CSI) is difficult to obtain because of feedback overhead. Resource allocation decisions therefore need to be made using outdated CSI. This mismatch restricts system energy efficiency. Traditional convex optimization methods have difficulty addressing this problem. Classical Deep Reinforcement Learning (DRL) algorithms also have limited convergence stability and policy performance under the stringent latency and reliability constraints of URLLC. To address these challenges, this paper considers a multi-node URLLC system under outdated CSI in dynamic scenarios. An energy-efficiency maximization problem is formulated under the Finite BlockLength (FBL) regime, with communication latency and reliability constraints. An efficient and stable algorithm is then designed for joint power and blocklength allocation.  Methods  A Successive Convex Approximation (SCA)-assisted DRL framework is proposed to maximize energy efficiency under outdated CSI. First, an SCA-based algorithm is developed to obtain a pre-allocation solution for transmit power and blocklength. This solution is feasible and physically interpretable, but relatively conservative. Based on this baseline, a Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is used for incremental refinement through interaction with the dynamic environment. This process reduces the conservatism of SCA. The SCA solution is used as prior knowledge in the state representation. Node location information is also incorporated into the state space. These designs narrow the policy search space and enable the DRL agent to better capture large-scale channel characteristics and system dynamics under outdated CSI. Learning efficiency and stability are therefore improved.  Results and Discussions  The proposed algorithm is evaluated through simulations and compared with three benchmark algorithms: an SCA-based optimization algorithm, a TD3 algorithm without SCA guidance, and a TD3 algorithm without node location information. The results show that the proposed method outperforms all benchmarks in convergence stability and system energy efficiency. In the training phase (Fig. 3), the average reward of the proposed algorithm increases steadily and converges stably. By contrast, removing node location information leads to lower rewards and stronger fluctuations. Removing SCA guidance causes the algorithm to converge to a much lower reward level. These results confirm the roles of SCA-based prior guidance and location-aware state representation in improving training stability. In the actual operation stage (Fig. 4), the proposed algorithm achieves high and stable energy efficiency and outperforms all comparison algorithms. Under outdated CSI, DRL-based methods can obtain higher energy efficiency than conservative optimization methods when transmission succeeds. However, removing node location information reduces energy efficiency, and removing SCA guidance increases transmission failures. These results verify the effectiveness of both designs in improving energy efficiency and maintaining policy feasibility. The effects of key system parameters are also examined. For basic resource parameters, a moderate increase in the blocklength budget (Fig. 5) or power budget (Fig. 6) improves system energy efficiency. For reliability constraints (Fig. 7), the reliability requirement should be set according to service requirements to avoid resource waste. Finally, the average energy efficiency under different numbers of nodes and different numbers of neurons in the TD3 network is analyzed (Fig. 8). The results provide guidance for algorithm configuration and network-scale design.  Conclusions  This paper addresses energy-efficient resource allocation for multi-node URLLC systems with outdated CSI by integrating SCA and DRL. In the proposed framework, a TD3-based DRL algorithm is guided by an SCA reference solution, and node location information is incorporated into the state representation. This optimization-learning dual-driven framework combines the interpretability and feasibility of model-based optimization with the adaptivity of data-driven learning. Simulation results show that the proposed method achieves higher energy efficiency than SCA-based optimization and conventional TD3 while satisfying URLLC latency and reliability constraints. The SCA reference solution improves policy stability and effectiveness under outdated CSI. Node location information further supports efficient decision-making. This work focuses on a single-cell multi-node scenario under Time Division Multiple Access (TDMA). Practical issues such as multi-cell interference, cooperative scheduling among multiple base stations, and more complex mobility patterns are not considered. Future work will extend the proposed framework to multi-cell and multi-agent scenarios and test its applicability under more severe CSI imperfections.
Review of Non-invasive Brain-Computer Interfaces for Continuous Motor Control
XU Minpeng, JIA Leyi, ZHOU Xiaoyu, CHEN Enze, WANG Junyang, XIAO Xiaolin, MING Dong
Available online  , doi: 10.11999/JEIT260011
Abstract:
  Significance   Continuous motor control is a core capability of Brain-Computer Interface (BCI) systems for natural and efficient interaction with external devices. Compared with discrete command-based control, continuous control supports real-time and smooth regulation of motion parameters such as position, velocity, and trajectory. This capability is required for applications in assistive mobility, neurorehabilitation, robotic manipulation, and immersive human-machine interaction. Although invasive BCIs have achieved high-performance continuous control through high-quality neural recordings, their dependence on surgical implantation limits long-term use and large-scale deployment. A systematic review of non-invasive continuous motor control BCI technologies is therefore needed to clarify research progress, methodological features, and remaining challenges.  Progress   Advances in non-invasive continuous motor control BCIs are reviewed from four closely related aspects: control paradigms, decoding algorithms, applications, and performance evaluation. At the paradigm level, motor imagery, steady-state visual evoked potentials, P300, and hybrid paradigms have been studied to support continuous control through sustained intention modulation, dynamic stimulus encoding, and hierarchical or shared-control strategies. For decoding algorithms, two main frameworks are identified: motion parameter mapping and motion parameter regression. Motion parameter mapping generates continuous output by temporally integrating discrete classification results or mapping them to velocity or state variables, whereas motion parameter regression directly establishes relationships between Electroencephalogram (EEG) features and continuous kinematic parameters. Recent studies increasingly incorporate nonlinear models and deep learning methods to improve robustness under the non-stationary nature of EEG signals. At the application level, non-invasive continuous control has progressed from two-dimensional cursor tasks to more practical scenarios, including wheelchair navigation, robotic arm manipulation, unmanned systems, and virtual or augmented reality environments. Existing studies also assess continuous control performance using both objective and subjective indicators, including trajectory error, task success rate, information transfer rate, workload, and user experience, reflecting varied experimental designs and control aims.  Conclusions  Existing studies show that non-invasive BCIs can support continuous motor control. However, current research remains at a stage in which multiple methods coexist without a unified framework. At the paradigm level, available approaches differ in their ability to elicit and sustain continuous motor intention reliably. For decoding algorithms, both motion parameter mapping and motion parameter regression are limited by the non-stationary nature of EEG signals, which affects robustness, generalization, and long-term stability. At the application level, many studies remain restricted to specific tasks and controlled environments, and the transfer of continuous control strategies to complex real-world scenarios still requires further validation. Moreover, the lack of standardized evaluation protocols hinders direct comparison and systematic optimization across studies.  Prospects   Future research should improve the stability and reliability of continuous control paradigms, enhance decoding robustness under realistic EEG conditions, and strengthen the match between control strategies and application requirements. Unified evaluation frameworks that integrate objective and subjective indicators should also be established to support methodological convergence and fair comparison. With continued progress, non-invasive continuous motor control BCIs are expected to play a growing role in assistive technologies, rehabilitation systems, and advanced human-machine interaction.
Research on Inverse QR Decomposition Optimization for Sparse Adaptive System Identification Algorithms
PENG Yi, ZHANG Pengfei, WANG Xiaoyong, GAO Junqi, LI Changlong, ZHANG Zhiyuan, SUN Tianxiang
Available online  , doi: 10.11999/JEIT250562
Abstract:
  Objective  Traditional sparse-regularized Recursive Least Squares (RLS) algorithms, namely L1/L0-norm Recursive Least Squares (L1/L0-RLS), have theoretical advantages in sparse parameter-space estimation and are widely used in system identification and channel equalization. However, under limited numerical precision, iterative covariance matrix computation may cause rounding errors to accumulate. This can lead to divergence and instability in the least-squares solution.  Methods  To address this problem, an improved algorithm based on the Inverse QR Decomposition (IQRD) framework is proposed. The framework suppresses rounding-error accumulation in traditional regularized RLS algorithms. It also removes the back-substitution step for weight coefficients required in conventional QR decomposition. These features improve numerical robustness and system identification efficiency in finite-precision environments. Specifically, L1-IQRD-RLS and L0-IQRD-RLS algorithms are constructed under an L1/L0-constrained IQRD architecture. A general recursive expression for the weight coefficients is derived. An automatic parameter selection mechanism is also incorporated into the algorithm framework to solve the dynamic optimization problem of the sparse regularization parameter.  Results and Discussions  Monte Carlo simulations are conducted to evaluate the sparse constraints and robustness of the proposed algorithms. The results show that L1-IQRD-RLS and L0-IQRD-RLS maintain long-term numerical stability in an 11-decimal-place fixed-point computing environment. Compared with traditional algorithms, the proposed algorithms show clear advantages in system sparsity representation, parameter estimation variance, and covariance matrix condition number. Measured-data verification further confirms that the improved algorithms maintain numerical stability under limited-precision conditions and are more robust than traditional methods. The measured-data results also show that the regularized RLS algorithms optimized by the IQRD framework have advantages in system sparsity representation, parameter estimation, and numerical stability. Their iterative convergence success rate is higher than that of traditional methods.  Conclusions  This paper addresses sparse system identification in adaptive filtering. Traditional sparse-regularized RLS algorithms still face numerical stability problems under limited numerical precision. To solve this problem, an IQRD framework is constructed to reduce the numerical ill-conditioning caused by accumulated rounding errors in sparse-regularized RLS algorithms. The proposed method improves numerical robustness in low-precision environments. In addition, an automatic parameter selection mechanism is incorporated into the algorithm framework. This reduces repeated parameter tuning and supports stable performance optimization under sparse constraints. In practical electromagnetic signal processing, system identification and beamforming are limited by the finite precision of hardware implementation and often exhibit inherent system sparsity. The proposed algorithm provides a targeted solution. Its finite-word-length robustness suppresses numerical divergence during adaptive weight updates and supports stable implementation on fixed-point processors. The sparse constraints also match the physical characteristics of sparse systems and improve estimation accuracy. This study provides a practical algorithm for high-performance and high-stability sparse-constrained systems on precision-limited hardware platforms.
Index Modulation Design with Sparse Spatial Constellation and Dynamic Multi-RIS-Block Selection for RIS-MIMO Systems
HUANG Fuchun, ZHU Han, TANG Xiaoqing, YANG Fan, HUANG Jie
Available online  , doi: 10.11999/JEIT251289
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS)-assisted Multiple-Input Multiple-Output (MIMO) Index Modulation (IM) systems face two main challenges: the difficult deployment of a single large-scale RIS panel and the high design complexity of efficient transmit spatial signal vectors. To address these issues, a joint design that combines sparse spatial constellation and dynamic multi-RIS-block selection is proposed. The design improves spectral efficiency, Bit Error Rate (BER) performance, and deployment flexibility.  Methods  Inspired by the Extended Space Index Modulation (ESIM) paradigm, a sparse spatial constellation with two active antennas (SCTA) is proposed, forming the SCTA-RIS-SM system. In this design, Pulse Amplitude Modulation (PAM) and Secondary PAM (SPAM) constellations are combined to construct the spatial constellation vector [x1,x2]T, which is modulated onto two active antennas. This design maximizes the Minimum Euclidean Distance (MED) between transmit vectors and improves the anti-interference capability of the system. To address the deployment difficulty of a single large-scale RIS panel, an enhanced SCTA-MBRIS-SM system is further proposed. The system uses a distributed array of small RIS blocks and dynamically selects a subset of blocks for cooperative reflection. Different RIS block selection combinations are used as a new IM dimension. Spectral efficiency and average BER are then analyzed theoretically. Monte Carlo simulations are conducted to compare the proposed systems with several existing schemes.  Results and Discussions  The simulation results show that the proposed SCTA-RIS-SM system achieves clear Signal-to-Noise Ratio (SNR) gains over RIS-SIM, RIS-SM, and DHRIS-SM systems at the same spectral efficiency, such as 10-12 bits/(s·Hz). For instance, when BER = 10–3, SCTA-RIS-SM outperforms RIS-SIM by approximately 1.5-2.5 dB and DHRIS-SM by more than 6 dB. By using additional IM from RIS block selection, SCTA-MBRIS-SM further improves BER performance and spectral efficiency compared with SCTA-RIS-SM, without increasing the number of Radio Frequency (RF) chains. With the same total number of reflecting elements, the proposed multi-RIS-block scheme achieves an SNR gain of up to 5 dB over RIS-SIM when BER = 10−3. The theoretical BER curves agree well with the simulation results in the high-SNR region, confirming the validity of the analytical derivations. The results also indicate that the performance advantage is maintained as the number of transmit antennas increases. In addition, the proposed design is compatible with channel coding.  Conclusions  This paper addresses the challenges of large-scale RIS deployment and high-complexity spatial signal design in RIS-assisted MIMO systems. The proposed SCTA design improves system reliability by optimizing the Euclidean distance distribution in the signal space. Dynamic multi-RIS-block selection transforms hardware deployment constraints into a new dimension for improving spectral efficiency, providing a feasible path for practical large-scale RIS applications. Simulation results confirm that joint optimization of transmit spatial vectors and RIS reflection degrees of freedom is an effective strategy for improving system performance. Future work will focus on robust design under imperfect channel state information, construction of higher-dimensional sparse constellations, extension to extremely large-scale MIMO scenarios, and multi-user communications.
One-step Reconstruction Diffusion Model-based Poisoning Attack on QoS-aware Cloud API Recommender Systems
TAN Zeyu, WANG Haoyuan, QI Mingyang, SUN Mengmeng, SHEN Limin, CHEN Zhen
Available online  , doi: 10.11999/JEIT260115
Abstract:
  Objective  In cloud computing, Cloud Application Programming Interfaces (cloud APIs) serve as key carriers for data output, capability reuse, and service delivery. They have become core elements in service-oriented software development and operation. With the rapid growth of cloud APIs, users often find it difficult to select suitable services from many functionally similar candidates. Quality of Service (QoS) is therefore used to differentiate cloud APIs by non-functional attributes. QoS-Aware cloud API Recommender System (QARS) plays an increasingly important role in guiding users toward suitable cloud APIs. However, existing studies mainly focus on improving recommendation accuracy and often ignore security risks caused by the economic value of cloud APIs and the openness of network environments. These risks are particularly evident in poisoning attacks. By injecting fake users, attackers can manipulate recommendation results and reduce the fairness and credibility of QARS. To address this threat from an attack-informed defense perspective, this paper analyzes the attack mechanisms of diffusion model-based poisoning methods and supports the design of targeted defense strategies.  Methods  The poisoning attack process and fake user profiles are first formally defined. Attack scale is then defined to flexibly simulate poisoning attacks under different settings. To analyze the attack principle of diffusion model-based methods, a One-step reconstruction Diffusion Model (ODM) is adopted, and a Preference guided one-step reconstruction Diffusion model-based Poisoning Attack framework (PDPA) is proposed. According to the collaborative principle that similar users tend to have similar preferences for cloud APIs, fake users generated by an attack method should have QoS values and cloud API invocation distributions similar to those of real users. This similarity allows fake users to exert collaborative influence and interfere with user preference modeling in QARS. PDPA is therefore designed to generate fake users that closely match real users. First, ODM separately models the QoS data and invocation distributions of real users. Unlike standard diffusion models, ODM avoids error accumulation caused by noise-dependent iterative denoising. It can generate fake-user invocation behavior similar to real-user behavior, which helps fake users exert effective collaborative influence. Then, to improve attack effectiveness, PDPA systematically selects fake users with invocation preferences for the target cloud API and assigns the maximum QoS value to the target item. This strategy strengthens the attack while reducing the disturbance caused by adding the target cloud API to fake-user invocation behavior, thereby improving stealthiness.  Results and Discussions  Experiments are conducted on the real-world WS-DREAM response-time QoS dataset. First, six recommendation methods, namely LR, MLP, DeepFM, AFM, DCN, and XSimGCL, are used as target recommender systems. Six baseline attack methods are used to simulate poisoning attacks. The results in Table 3 reveal the vulnerability of QARS to poisoning attacks. All attack methods reduce recommendation accuracy. PDPA achieves the best attack effectiveness in most experimental settings because it sufficiently models user invocation preferences, enabling fake users to exert stronger collaborative influence on QARS. Second, fake users generated by ODM and those generated by the standard diffusion model are compared in terms of F1 score and latent-space distribution. The results in Figure 2 show that ODM outperforms the standard diffusion model in stealthiness and produces a latent-space distribution closer to that of real users. Third, ablation studies are conducted for each module of PDPA. The results in Tables 4 and 5 verify that each module is necessary for attack effectiveness and fake-user stealthiness. Finally, Mean Absolute Error (MAE) and F1 score are compared under different attack scales to evaluate the effect of attack scale on attack effectiveness and stealthiness. The results in Figure 3 and Table 6 show that increasing the attack scale improves attack effectiveness but also increases the number of detected fake users.  Conclusions  This paper investigates the threat of poisoning attacks against QARS by analyzing the attack process and key attack parameters. The proposed PDPA simulates poisoning attacks on QARS and reveals their vulnerability. The results show the potential of diffusion models for poisoning attacks and verify the necessity of separately modeling QoS data and cloud API invocations. PDPA also clarifies how diffusion models generate fake users, providing a basis for future targeted countermeasures.
Multi-task Lightning Nowcasting with Spatio-temporal Focal Perception and Synergistic Weighted Loss
TANG Zhihao, HAN Yuanpeng, ZHANG Hui, SONG Lin, ZHANG Qilin, LIU Yi
Available online  , doi: 10.11999/JEIT260234
Abstract:
  Objective  Lightning nowcasting is essential for early warning systems and for protecting critical infrastructure, including aviation, power grids, and transportation systems. Traditional numerical weather prediction models depend strongly on parameterization schemes and require high computational costs, which limits their use in rapid-update nowcasting. Although deep learning methods have advanced, they still have difficulty handling extreme data sparsity, suffer from serial-computation bottlenecks in recurrent architectures, and mainly focus on binary occurrence prediction rather than the joint optimization of lightning-frequency prediction and regional localization. Moreover, conventional loss functions are easily dominated by extensive non-lightning areas, which biases predictions toward zero or causes excessive false alarms. To address these limitations, a Spatio-Temporal Focal perception and synergistic weighted loss Network (STF-Net) is proposed as a multi-task lightning nowcasting model that jointly predicts lightning frequency and occurrence regions. It integrates three key components: a Lightning Adaptive Attention Module (LAAM) for explicit spatio-temporal dependency modeling, a Spatio-Temporally Weighted Hybrid Loss for data sparsity and imbalance, and a spatio-temporal dual-branch Generative Adversarial Network (GAN) to improve prediction fidelity and temporal coherence.  Methods  STF-Net is built on the SimVP video prediction architecture and adopts an encoder-translator-decoder paradigm. LAAM uses a three-dimensional decoupled attention mechanism along the height, width, and channel dimensions, enabling adaptive focus on convectively sensitive regions while maintaining computational efficiency. The Spatio-Temporally Weighted Hybrid Loss combines Temporally Weighted Mean Squared Error (TW-MSE) for frequency regression and Dual-Weighted Cross-Entropy loss (DWCE) for regional localization. Time-increasing weights are incorporated to improve medium- to long-term forecast robustness (Fig. 5). DWCE integrates static class weights with dynamic grid weights, thereby balancing global class proportions and local lightning-frequency heterogeneity. A spatio-temporal dual-branch GAN, consisting of a spatial PatchGAN discriminator and a temporal three-dimensional convolutional discriminator, is used to improve the textural fidelity and temporal coherence of predicted lightning-frequency fields. The model uses six consecutive historical lightning-frequency frames at 256×256 resolution and 10-min intervals to predict the next six frames, corresponding to a 1-h forecast window. Experiments are conducted on a high-resolution Very Low Frequency Lightning Location Network (VLF-LLN) dataset containing 11,748 images that cover different seasonal and weather conditions. The dataset is split at a ratio of 7:3 for training and testing.  Results and Discussions  Comprehensive evaluation metrics are used, including Mean Squared Error (MSE) and Mean Absolute Error (MAE), which are computed only on lightning pixels; Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) for image fidelity; and Probability Of Detection (POD), False Alarm Rate (FAR), and Critical Success Index (CSI) for regional detection. STF-Net achieves a CSI of 0.663 within the 1-h forecast window, representing a 14.5% improvement over the SimVP baseline (0.579). It also reduces FAR from 0.351 to 0.216, corresponding to a relative reduction of 38.5% (Table 1). Ablation studies validate each component. Adding GAN improves CSI to 0.624 and reduces MSE to 0.109. Further incorporation of LAAM increases CSI to 0.629 and yields the highest POD of 0.894. The complete STF-Net with the hybrid loss achieves the best performance, with a CSI of 0.663 and an MSE of 0.105 (Table 1, Fig. 5). The joint prediction of frequency and region is supported by simultaneous improvements in regression metrics (MSE and MAE) and detection metrics (CSI and FAR). Time-step analysis shows that LAAM reduces long-term performance degradation, with STF-Net maintaining the highest CSI compared with SimVP+GAN and SimVP (Fig. 6). Comparative experiments with ConvLSTM and PredRNN further demonstrate the superiority of STF-Net across all lead times. STF-Net consistently achieves higher CSI and lower FAR, and its advantage becomes more evident as the forecast horizon increases (Fig. 7). Its consistent gains in PSNR and SSIM further indicate that the spatio-temporal GAN helps generate coherent, detail-rich predictions. Visualization results show that STF-Net produces structurally clear and continuous lightning-activity regions centered on high-frequency areas. It accurately tracks dynamic evolution patterns, including movement, merging, and splitting, while generating minimal noise in non-lightning regions. These results demonstrate effective collaborative prediction of both frequency magnitude and spatial distribution.  Conclusions  STF-Net is presented as a deep learning model for joint lightning-frequency prediction and regional localization. It explicitly models long-range spatio-temporal dependencies, focuses on critical convective zones, addresses extreme data sparsity and class imbalance, and jointly optimizes frequency regression and regional localization while suppressing false alarms. Its spatio-temporal dual-branch GAN further improves the spatial structural consistency and temporal coherence of predictions. Experimental results show that STF-Net outperforms baseline and state-of-the-art models, achieving a CSI of 0.663, an FAR of 0.216, and the best MSE and MAE values within a 1-h forecast window. The model effectively reduces long-term performance degradation, captures the evolution trends of lightning regions, and generates physically plausible predictions with minimal background noise. This study provides an efficient end-to-end solution for operational lightning nowcasting systems and offers guidance for model design in sparse meteorological spatio-temporal sequence prediction.
Intelligent Privacy-Aware Computation Offloading Method against Multi-server Joint Inference Attacks
MIN Minghui, LIU Mingcheng, ZHANG Peng, DUAN Jincheng, LI Shiyin, ZHANG Hongliang
Available online  , doi: 10.11999/JEIT260249
Abstract:
  Objective  With the rapid development of the low-altitude economy, services such as intelligent transportation, smart healthcare, and low-altitude logistics have become increasingly common. Their efficient operation depends on the real-time processing of massive sensing data. Mobile Edge Computing (MEC) improves task execution efficiency and reduces device computational burdens by offloading tasks to nearby servers. However, user privacy and security risks have become increasingly severe. In dynamic scenarios where multiple MEC servers jointly process tasks, information sharing can enable multi-server joint inference attacks and greatly increase the risk of user location privacy leakage. Although existing studies have used Differential Privacy (DP) to protect user location privacy, current DP-based solutions remain limited. These methods inject noise into offloading decisions, but unconstrained noise may reduce task allocation accuracy. In addition, user mobility causes continuous changes in channel states during dynamic computation offloading. Privacy leakage risks and attacker behaviors are also uncertain. Traditional optimization methods based on static system models are therefore unsuitable for such dynamic environments. To address these challenges, this paper proposes an Asynchronous Advantage Actor-Critic (A3C)-based Intelligent Privacy-Aware Computation Offloading (AIPCO) scheme against multi-server joint inference attacks. The proposed scheme protects user location privacy while maximizing the overall utility of the MEC system.  Methods  This paper proposes a DP-based task offloading rate perturbation mechanism. By adding controlled noise, the mechanism increases the randomness of user task offloading toward multiple MEC servers. A truncated Laplace mechanism is used to constrain the perturbed offloading rates within valid boundaries. This design satisfies the mathematical guarantees of DP and reduces the accuracy of multi-server joint inference attacks on sensitive user locations. Privacy entropy is then introduced to dynamically evaluate the real-time privacy protection level. Finally, the AIPCO scheme is constructed. Through a multi-threaded asynchronous training mechanism, the scheme interacts with the environment through iterative trial and error and efficiently learns the optimal real-time offloading policy online. The proposed scheme dynamically protects user privacy, reduces computational cost, and maximizes comprehensive system utility.  Results and Discussions  The AIPCO scheme jointly optimizes user privacy and task offloading cost by incorporating multidimensional performance variables into the reinforcement learning reward function. A comprehensive performance analysis (Fig. 4) shows that, when the number of continuous learning iterations reaches 200, the privacy protection level of AIPCO increases by 2.52%, 3.56%, and 22.90% compared with RCLM, JODRL, and DODA-DT, respectively. This advantage is mainly attributed to the DP-based task offloading rate perturbation method, which uses the truncated Laplace mechanism to increase data randomness while strictly constraining the perturbation range. By contrast, RCLM perturbs the task offloading rate through range-limited DP without using the truncated Laplace mechanism. JODRL increases randomness only through network policy optimization, resulting in a lower privacy protection level. DODA-DT focuses on balancing energy consumption and system latency without optimizing user privacy. For the privacy weight parameter (Fig. 5), increasing $\omega$ improves privacy protection. For example, the privacy protection level increases by 5.64% when $\omega$ rises from 0.2 to 0.7, with a clear performance gain at 0.7. As the system agent reduces its focus on computational cost, user utility remains optimal despite increased cost. When the physical distance between users and the MEC server is adjusted (Fig. 6), AIPCO shows stronger privacy protection in long-distance scenarios. A greater distance reduces the number of tasks offloaded to the server. Therefore, attackers obtain less information, and privacy protection improves. Although computational cost increases with distance, AIPCO consistently outperforms competing schemes. These results confirm that AIPCO achieves optimal MEC system utility while protecting user privacy.  Conclusions  To mitigate multi-server joint inference attacks caused by information sharing among collaborative MEC servers, this paper proposes an AIPCO method. A DP-based task offloading rate perturbation scheme is designed to increase randomness, and a truncated Laplace mechanism is used to constrain perturbed rates within reasonable boundaries. The scheme is proven to satisfy strict DP mathematical guarantees, and privacy entropy is introduced to quantitatively evaluate the privacy protection level. In addition, the AIPCO scheme uses a multi-threaded asynchronous training mode, enabling the agent to efficiently learn the optimal perturbed offloading policy in a continuous space and maximize overall system utility. Simulation results show that the proposed scheme outperforms the baselines in both dynamic and average performance. It achieves optimal system utility while protecting user privacy.
Bearing Fault Diagnosis of Roadheader via Cross-modal Kernel Fusion-sphere Space Learning
SU Shuzhi, GUI Yang, MA Tianbing, ZHU Yanmin, WU Kanghui
Available online  , doi: 10.11999/JEIT260494
Abstract:
  Objective  Traditional roadheader bearing fault diagnosis methods often struggle with high-dimensional and nonlinear multi-sensor data. They also fail to effectively perceive cross-modal, multi-scale fault information or integrate local and global structural features. To address these limitations, this paper proposes a Cross-modal Kernel Fusion-sphere Space Learning (CKFSL) method. By perceiving cross-modal multi-scale fault information, CKFSL extracts highly discriminative features from roadheader bearing cross-modal fault samples and improves diagnostic accuracy.  Methods  CKFSL first maps roadheader bearing cross-modal fault samples into a high-dimensional kernel space through implicit transformation. Dual extremal point anchoring and polar neighbor allocation mechanisms are then used to capture fault sample clusters with similar isomorphic information, forming kernel fusion-spheres. An adaptive binary partitioning strategy is designed according to the geometric span of internal fault samples. This strategy tightens isomorphic boundaries, constructs micro-neighbor kernel fusion-spheres, and achieves highly isomorphic manifold aggregation at the microscopic scale. A micro-neighbor kernel fusion-sphere space is further formed to re-evaluate local isomorphism (Fig. 1). To characterize wide-area topological correlations, a wide-area topological isomorphism constraint is proposed. This constraint constructs a wide-area dynamic isomorphism graph among micro-neighbor kernel fusion-spheres (Fig. 1). Finally, an objective optimization function is formulated within the space learning framework. It integrates local manifold isomorphism and wide-area topological correlations of roadheader bearing cross-modal fault samples, as shown in the CKFSL diagnostic flowchart (Fig. 2). The analytical solution for spatial projection is theoretically derived to obtain discriminative cross-modal kernel fusion-sphere space isomorphic features from roadheader bearing cross-modal fault samples.  Results and Discussions  CKFSL is first validated on the self-built AUST roadheader bearing cross-modal fault dataset, with the experimental platform shown in Fig. 3. The average recognition rates obtained with increasing numbers of training fault samples are shown in Fig. 4. On the AUST dataset, CKFSL achieves a recognition rate of 99.49% with only 70 training fault samples and reaches 100% as the number of training fault samples increases. Table 1 summarizes the standard deviations under different training fault sample sizes. The results show that CKFSL has the lowest standard deviation and stronger robustness than the other seven comparison algorithms. Three-dimensional fault feature distributions are shown in Fig. 5. The results confirm that CKFSL effectively separates highly overlapping fault samples into different clusters and reduces the boundary confusion observed in the comparison algorithms. To verify generalization capability, CKFSL is further evaluated on the public Paderborn dataset, with the experimental setup shown in Fig. 6. As shown in Fig. 7 and Fig. 8, CKFSL achieves a 100% average recognition rate across four complex fault categories. It also outperforms the comparison algorithms, which have difficulty exceeding an 85% recognition rate for the F4 fault category.  Conclusions  CKFSL effectively addresses the inability of traditional roadheader bearing fault diagnosis methods to perceive complex multi-scale fault information. By using the wide-area dynamic isomorphism graph learned in the micro-neighbor kernel fusion-sphere space, CKFSL integrates local manifold isomorphism with wide-area topological correlations of roadheader bearing cross-modal fault samples. This process enables CKFSL to extract highly discriminative cross-modal kernel fusion-sphere space isomorphic features. It improves the accuracy of roadheader bearing fault diagnosis and supports the reliability and continuous operation of roadheaders.
Physiological Signal-driven QoE Optimization for Wireless Virtual Reality Transmission
WU Chang, PENG Mingyu, CHEN Yuang, CHEN Yiyuan, GUO Fengqian, QIN Xiaowei, LU Hancheng
Available online  , doi: 10.11999/JEIT260067
Abstract:
  Objective  Virtual Reality (VR) has become a transformative medium for immersive digital experiences because it can deliver high-resolution 360° video with ultra-low Motion-To-Photon (MTP) latency. However, its dependence on wireless transmission creates major challenges. Uncompressed data rates above 1 Gbit/(s·Hz) and latency thresholds below 20 ms place stringent demands on network infrastructure. In mobile scenarios, channel fluctuation and user mobility often compromise service continuity and cause abrupt resolution changes. Traditional Quality of Service (QoS) metrics, such as bandwidth, jitter, and packet loss, provide useful network-level information but cannot adequately reflect subjective user satisfaction. Existing Quality of Experience (QoE) models and Adaptive BitRate (ABR) algorithms often use symmetric metrics, such as Mean Opinion Score (MOS), and overlook the fact that users perceive quality deterioration and quality improvement differently. Sudden resolution downgrading has a stronger negative effect on immersion than the positive effect caused by resolution upgrading. This perceptual asymmetry is consistent with behavioral psychology but remains insufficiently addressed in current transmission schemes. In addition, the separation between Radio Access Network (RAN) resource provisioning and application-layer bitrate adaptation often causes mismatched optimization, video-quality oscillation, and resource underuse. To address these issues, this study establishes a quantitative link between physiological responses and resolution changes. It further develops a physiological signal-driven QoE framework integrated with Deep Reinforcement Learning (DRL) to support adaptive transmission, maximize immersion, and reduce the adverse effects of resolution fluctuation in resource-constrained wireless networks.  Methods  A two-stage method is adopted, including physiological signal analysis and joint optimization framework design. A controlled VR experiment is conducted to quantify the perceptual effect of resolution changes. Nineteen healthy subjects participate in a viewing task using an eye-tracking VR headset, a 32-channel wireless ElectroEncephaloGraphy (EEG) system, ElectroCardioGraphy (ECG) recording, and Galvanic Skin Response (GSR) sensors. The subjects view natural-scene videos in which the resolution levels, including 8k, 4k, 1080P, 720P, and 480P, switch randomly every 8 s. The collected EEG signals are preprocessed by independent component analysis and band-pass filtering. Event-Related Potential (ERP) components are analyzed, with emphasis on the N200 component in the temporal and occipital regions, which reflects visual processing and attention allocation. A Linear Discriminant Analysis (LDA) classifier is used to distinguish different response types. The analysis focuses on the asymmetry between resolution upgrading and downgrading, and on sensitivity to the magnitude of resolution jumps. Based on these physiological findings, a QoE model is formulated by adding penalty terms for resolution degradation and large-amplitude resolution switching. These penalties are weighted more strongly than upgrade rewards to represent user aversion to quality drops. The model is then integrated into an edge-computing environment through a dual-timescale DRL framework. The framework separates control into two cooperative agents: the Scheduling and Utility (SU) agent and the Resolution Scaling (RS) agent. The SU agent operates at the millisecond timescale and performs real-time wireless resource allocation. It uses a Gated Recurrent Unit (GRU) to extract temporal features from Channel State Information (CSI) and transmission history. It then dynamically allocates bandwidth to improve frame delivery success and maintain fairness under VR frame-deadline constraints. The RS agent operates at the frame timescale and determines the resolution of subsequent video frames. Its decision-making is guided by the physiological signal-driven reward function, which penalizes actions that may trigger negative physiological responses, such as sharp resolution drops, unless channel deterioration makes them necessary. Proximal Policy Optimization (PPO) is selected for both agents because of its stable learning behavior in continuous and discrete action spaces. Simulations are conducted using a 3GPP-based wireless channel module with user mobility, shadow fading, and path loss to create a dynamic network environment.  Results and Discussions  The physiological experiment and network simulations validate the proposed framework. In the physiological analysis, a clear N200 response is observed approximately 200 ms after resolution changes. The N200 amplitude is significantly larger during resolution downgrading than during resolution upgrading (p < 0.001), indicating that users are more sensitive to quality deterioration. Large resolution jumps, such as changes from 8k to 1080P, also induce stronger neural responses and more concentrated occipital energy than minor adjustments. The LDA classifier achieves an average Area Under the Curve (AUC) of 74.12% across 19 subjects, confirming that neural responses contain discriminative information about the direction of resolution change. The GSR results support these findings. A dual-branch GSR feature extraction and classification model reaches an average AUC of 78.10% in distinguishing upward and downward switching events. By contrast, ECG signals do not show a stable effect under the current experimental setting and analysis granularity. Therefore, the subsequent QoE model is mainly constructed from EEG and GSR findings. In the network performance evaluation, the proposed physiological signal-driven DRL framework is compared with several baselines, including Proportional-Fair (PF) scheduling, equal resource allocation, and traditional congestion control represented by SCReAM. The training curves show that the dual-agent system converges and learns to coordinate capacity provisioning with resolution decisions. The SU agent smooths short-term channel fluctuation and provides a stable capacity basis, which enables the RS agent to make more reliable resolution decisions. Quantitative results show that the proposed scheme improves the average video resolution by up to 88.7% compared with the equal-resource baseline. More critically, the resolution switching frequency is reduced by up to 81.0%. This reduction is essential because frequent switching, especially downward switching, causes user discomfort, as demonstrated by the physiological analysis. By prioritizing long-term resolution stability and penalizing abrupt drops through the physiological signal-driven reward function, the proposed system reduces the “ping-pong” effect commonly observed in traditional ABR algorithms. Compared with schemes using different penalty weights, the proposed method achieves a better balance. It avoids overly conservative behavior under large penalties, which lowers the average resolution, and unstable visual quality under small penalties, which increases resolution fluctuation. The joint optimization also allocates resources preferentially to users with urgent frame deadlines or higher risks of perceptible quality degradation, while maintaining a frame delivery success rate above 99%.  Conclusions  This paper addresses the conflict between wireless-channel instability and the human need for visually consistent VR streaming. By adopting a physiological signal-driven approach, the asymmetric effect of resolution changes on user experience is quantified, which challenges the symmetric assumptions used in traditional QoE models. Integrating this physiological evidence into a dual-timescale DRL framework enables the RAN to go beyond throughput-oriented optimization. Wireless resource allocation supports stable application-layer adaptation, while application-layer demands guide resource scheduling. The proposed solution improves immersive experience by increasing average resolution and reducing the physiologically disruptive effects of sudden quality degradation. The reduction in resolution switching frequency by more than 80% shows that the system can shield users from network variability. This study also indicates the value of edge intelligence in making resource-allocation decisions based on human perception rather than network statistics alone. Future work should extend the QoE model by considering multisensory factors, such as MTP latency, cybersickness, spatial distortion, stalling, and audiovisual synchronization. Individual differences in physiological sensitivity should also be addressed through personalized modeling. For real-world deployment, privacy protection is essential. Federated learning and local edge updates may allow biometric data to be processed locally while supporting global policy optimization. This work provides a human-centric basis for immersive networking and shifts the focus from QoS to physiologically validated QoE.
A General Evaluation Framework for Mission Planning Algorithms for Remote Sensing Satellite Constellations
LI Jinfei, YU Xiaogang, TIAN Jing, HE Haochen, XING Xiangwei, ZHANG Xiaohan
Available online  , doi: 10.11999/JEIT260335
Abstract:
  Objective  The rapid growth in remote sensing satellite constellations has shifted mission planning from single-satellite static scheduling to large-scale dynamic coordination across heterogeneous constellations. However, evaluation methods have not kept pace with algorithm development. Existing studies often rely on private datasets, simplified metrics centered on Completion Rate, and idealized simulations that ignore realistic constraints, such as attitude maneuvers, illumination conditions, and dynamic task insertion. These limitations prevent fair cross-paper comparison and slow engineering application. To address this gap, this paper proposes the Remote Sensing Constellation Mission Planning Benchmark (RSCMP-Bench), a general, open, and reproducible evaluation framework. It is designed as a unified benchmark for the community, similar to ImageNet in computer vision and General Language Understanding Evaluation (GLUE) in Natural Language Processing (NLP).  Methods  RSCMP-Bench consists of three components. First, the multi-scenario standard task library contains 300 standardized scenarios at three difficulty levels: Low, Medium, and High, with 100 scenarios per level. Satellite numbers range from 30 to 200, and task demands range from 56 to 560. All scenarios are generated from public Two-Line Element (TLE) data and explicitly model realistic constraints. Optical satellites require a minimum solar elevation angle, and Synthetic Aperture Radar (SAR) satellites require incidence angles within specified ranges. General constraints, such as per-orbit maximum on-time, minimum single-operation on-time, attitude maneuver time, and valid execution windows, are also modeled. The scenarios include point tasks, area tasks, static tasks, and dynamically inserted tasks. Second, the multi-dimensional effectiveness evaluation system includes a Basic Performance layer and a Dynamic Adaptability layer. The Basic Performance layer uses Completion Rate, Weighted Completion Rate, Average Response Delay, and Time Utilization. The Dynamic Adaptability layer uses multi-stage rolling evaluation with random dynamic task insertion. The Dynamic Adaptability Score measures the post-insertion Completion Rate relative to the baseline, and Dynamic Response Efficiency measures the performance gain per unit replanning time. A composite RSCMP-Bench Score is also provided. Third, the simulation and evaluation platform uses a client-server architecture. It integrates a Simplified General Perturbations 4 (SGP4) propagator, algorithm adapters, two-stage constraint verification, an intelligent scenario generator, and visualization tools. The platform has been deployed at https://www.tianzhibei.com and has supported a national competition with more than 80 research teams.  Results and Discussions  Baseline experiments comparing Random Scheduler and Priority Greedy validate the feasibility, reproducibility, and discriminative capacity of RSCMP-Bench. Random Scheduler yields very low Completion Rates of 7.3%, 3.8%, and 1.9% on the Low, Medium, and High levels, respectively. These results confirm the extreme sparsity of the feasible solution space. Priority Greedy achieves higher Completion Rates but still degrades as scenario difficulty increases, decreasing from 76.1% at the Low level to 63.7% at the Medium level and 49.2% at the High level. These findings indicate that high-difficulty scenarios remain challenging even for reasonable heuristic methods. They also show considerable room for more advanced algorithms. The dynamic adaptability protocol quantifies algorithm robustness under unexpected dynamic task insertion, which is not captured by static evaluations. The two-stage constraint verification module rejects infeasible plans and generates detailed error reports to support debugging.  Conclusions   RSCMP-Bench provides a unified, fair, and reproducible benchmark for remote sensing constellation mission planning. By combining a public library of 300 standardized scenarios, a multi-dimensional effectiveness evaluation system based on Basic Performance and Dynamic Adaptability, and a simulation and evaluation platform with realistic constraints and automated scenario generation, the framework addresses the long-standing lack of standardized evaluation in this field. Baseline results confirm its discriminative capacity and reveal clear performance bottlenecks in large-scale dynamic scenarios. Inspired by ImageNet and GLUE, RSCMP-Bench can support systematic community evaluation and fair competition. The framework has been deployed at https://www.tianzhibei.com, and its adoption can accelerate progress in intelligent mission planning for next-generation remote sensing constellations.
Dual-MPC-Driven Modeling and Spatiotemporal Evolution of Intelligent Connected Traffic Risk Fields
JIANG Linyuan, DING Fei, FAN Xuan, YANG Xuechao, SONG Aiguo, ZHANG Dengyin
Available online  , doi: 10.11999/JEIT260194
Abstract:
  Objective  With the deployment of intelligent connected vehicle-road-cloud cooperative systems, roadside infrastructure is evolving from traffic-state sensing units into intelligent decision-support platforms for multi-vehicle interaction analysis and dynamic risk inference. In highway and urban freeway scenarios, traffic operation is affected not only by the kinematic responses of individual vehicles but also by lane-changing intentions, car-following competition, and conflict propagation under local interactions. From a roadside perspective, a unified framework is therefore needed to continuously represent traffic risk, reveal its spatiotemporal evolution, and couple risk information with behavior decision-making and trajectory planning. Existing car-following models, such as the Optimal Velocity Model (OVM), Full Velocity Difference (FVD) Model, and Intelligent Driver Model (IDM), can describe speed-spacing evolution. However, these models mainly focus on longitudinal interactions and usually embed risk implicitly in safety-distance or acceleration constraints. They cannot explicitly characterize the coupling between longitudinal following and lateral lane changing, nor can they provide a continuous risk representation suitable for regional traffic assessment. Although Artificial Potential Field (APF) methods and Model Predictive Control (MPC) methods can improve trajectory safety, existing studies still lack a unified mechanism that links risk assessment, behavior decision-making, and motion planning. In addition, discrete behavior choices and continuous control actions are difficult to process efficiently within a single optimization framework.  Methods  An intelligent connected traffic risk-field model oriented toward vehicle-road cooperation is first established. The model integrates vehicle-interaction risk, lane-marking constraint risk, and road-boundary repulsive risk (Fig. 1). In the vehicle-interaction layer, motion-state-induced risk is formulated by considering the relative speed and relative orientation between the ego vehicle and surrounding vehicles. Distance-induced risk is modeled to reflect attenuation as separation distance increases (Fig. 2(a)(b)). To represent the stronger influence of forward hazards than lateral and rear hazards, a directional non-uniformity coefficient is used. This coefficient adjusts the angular attenuation of field strength and enables anisotropic spatial risk representation around the vehicle. In the road-constraint layer, lane markings and road boundaries are modeled separately. The total driving risk field is obtained by weighting and combining the lane-marking field, road-boundary field, and multi-vehicle interaction field (Fig. 2(e)(f)). Based on this representation, a dual-MPC hierarchical decision and motion-planning architecture is designed (Fig. 3). In each control cycle, the upper layer evaluates candidate behavior modes according to the vehicle state, surrounding traffic state, and dynamic risk field, and then outputs a unique behavior mode. The lower layer activates the corresponding control branch. When lane keeping is selected, longitudinal speed-planning MPC is used. When lane changing is selected, lane-change trajectory-planning MPC is activated under road-boundary, lane-marking, and safe-gap constraints.  Results and Discussions  The proposed framework reconstructs microscopic traffic risk evolution under different datasets and time scales. In the HighD highway scenario, when the slicing interval is 0.4 s, the local evolution of following and lane-changing interactions is captured in detail. This includes the process in which the ego vehicle initially follows a preceding vehicle and then starts changing to the adjacent lane (Fig. 4(a)(d)). When the interval is increased to 1.0 s, a wider spatiotemporal interaction range becomes visible, and lane-change completion and the subsequent return maneuver are identified more clearly (Fig. 4(e)(h)). In the NGSIM scenario, a smaller interval provides a finer description of the lane-change disturbance process. By contrast, a larger interval reveals the wider reconstruction of interaction relationships among the original lane, target lane, and surrounding vehicles (Fig. 4(i)(p)). These results indicate that the proposed roadside-oriented risk field can describe both local interaction details and larger-scale evolution trends, depending on the selected reconstruction interval. Sensitivity experiments further confirm the role of the directional non-uniformity coefficient. As this coefficient increases, forward risk concentration becomes stronger, local peak risk increases, and the coverage of high-risk regions decreases (Table 2). This finding shows that the coefficient effectively regulates anisotropic field distribution. Comparative experiments with IDM, OVM, FVD, and APF show that the proposed method performs better in most representative scenarios and error metrics (Fig. 5, Table 3). In the lateral cut-in scenario, its advantage lies in the early representation of lateral intrusion risk, which enables the behavior decision layer to anticipate conflict and the motion-planning layer to generate continuous evasive actions. In congested scenarios, the superposition of forward congestion risk, lateral neighboring-vehicle influence, and road-boundary constraints allows the dual-MPC controller to evaluate safety and feasibility simultaneously in local space.  Conclusions  A unified framework for roadside-oriented traffic-risk modeling and behavior-driven trajectory planning is developed. By integrating multi-vehicle interaction risk, lane-marking constraint risk, and road-boundary repulsive risk into a continuously evolving dynamic risk field, the spatial quantification of multi-vehicle interaction risk is realized. The directional non-uniformity coefficient further enables asymmetric risk perception modeling in forward, lateral, and rear directions. On this basis, a dual-MPC hierarchical architecture is constructed to couple behavior decision-making with motion planning, so that lane-keeping and lane-changing behaviors can be adaptively selected and optimized under a unified risk-driven mechanism. Experiments based on HighD and NGSIM datasets show that the proposed method can effectively characterize the spatiotemporal evolution of traffic-risk fields and outperform representative comparison models in most typical scenarios and error metrics.
MGM-3DUNet: A Multi-scale Edge Semantic-guided GraphConvolutional Sequence Method for Brain Tumor Segmentation
ZHUANG Jianjun, LI Xiang, JING Shenghua, LÜ Zhenglong
Available online  , doi: 10.11999/JEIT260128
Abstract:
  Objective  Feature fusion in U-Net and its 3D variants mainly relies on simple single-scale concatenation, which limits the use of encoder features and weakens fine-grained segmentation of Tumor Core (TC) and Enhancing Tumor (ET) regions. Recent methods such as VM-UNet improve sequence modeling efficiency, but they mainly focus on global information modeling. Local detail preservation and edge enhancement remain insufficient. Therefore, current methods still have limitations in segmentation accuracy and clinical utility. To address these problems, this paper proposes MGM-3DUNet for brain tumor segmentation.  Methods  The Multi-Scale Edge semantic Guidance Module (MEGM) is designed to improve tumor boundary segmentation through learnable edge detection. The Graph Convolutional Sequence Module (GCSM) combines the local aggregation ability of graph convolution with efficient long-range modeling based on a Mamba-like structure. This design improves semantic consistency while preserving small tumor structures with fewer parameters. The Multi-scale Context Perception Module (MCPM) is introduced to strengthen feature complementarity across different tumor scales through dual-scale fusion.  Results and Discussions   Experiments show that the proposed method achieves better average Dice similarity coefficient (Dice) and 95th percentile Hausdorff Distance (HD95) than the comparison methods. With only 2.3M parameters, MGM-3DUNet achieves Dice values of 91.2%, 90.4%, and 89.2% for Whole Tumor (WT), TC, and ET, respectively. The visualization results (Fig. 9, Fig. 10) further show that MEGM improves boundary localization. Overall, the proposed method shows improved sensitivity to edge details and contextual correlations while maintaining a low parameter count.  Conclusions   This method improves tumor boundary prediction by introducing shallow-layer edge enhancement to emphasize tumor contours. Local and global semantic information is fused in the bottleneck layer, and multi-scale contextual features are integrated during decoding. The proposed design achieves accurate segmentation with low computational cost and is suitable for deployment on resource-constrained platforms.
A Low-latency Synchronization Header Detection Algorithm and Circuit for the JESD204C Interface
YIN Peng, ZHANG Chao, LEI Changan, HOU Weizhou, SHU Zhou, LIU Shubin, ZHU Zhangming
Available online  , doi: 10.11999/JEIT260163
Abstract:
  Objective  With rapid advances in high-speed electronics, front-end Analog-to-Digital Converters and Digital-to-Analog Converters (ADCs/DACs) continue to increase in sampling rate and resolution. Back-end Field-Programmable Gate Arrays and Application-Specific Integrated Circuits (FPGAs/ASICs) also provide stronger computing capability. These trends impose strict requirements on high-speed data interfaces, including high bandwidth, low latency, low power consumption, and reliable synchronization. As a mainstream high-speed Serializer/Deserializer (SerDes) interface, the JESD204C interface still suffers from long link initialization latency and high synchronization power consumption. These limitations restrict system real-time performance and energy efficiency. To address these issues, this study optimizes the link-layer design of the JESD204C receiver and proposes an efficient Synchronization Header (SH) detection method. The method implements exponential compression of the search set through global observation and iterative convergence. Detection efficiency is improved, fast and accurate SH positioning is achieved, link synchronization latency is reduced, and synchronization stability and energy efficiency are enhanced.  Methods  A typical JESD204C interface uses serial sliding detection for SH detection, which causes high link initialization latency and large delay jitter. To solve these problems, an Iterative Set Screening (ISS)-based SH detection algorithm is proposed. The SH detection task is modeled as the rapid localization of a deterministic pattern in a binary random sequence. A theoretical model based on information theory and stochastic processes is constructed. Expected space utilization and Bit Error Rate (BER) are introduced to support quantitative performance evaluation. In this model, SH candidate positions are defined as a dynamic set. Based on the inherent polarity inversion characteristic of the SH and global observations in each clock cycle, multilevel XOR logic is used to verify all candidate hypotheses in parallel. Non-inverting candidate positions are eliminated, and the search space is dynamically compressed. This design improves synchronization speed and position robustness, providing a low-latency and reliable initialization solution for high-speed SerDes links.  Results and Discussions  The proposed ISS-based SH detection algorithm is validated under harsh conditions, including SH crossing block boundaries and loss of lock caused by burst errors. The results demonstrate strong robustness, with rapid SH locking and link resynchronization under all test conditions (Figures 1116). To evaluate performance, four representative schemes are reproduced: a single-bit serial locking circuit, a 66-bit serial locking architecture, a register-intensive block synchronization method, and a parallel search circuit. A systematic comparison is then conducted between these schemes and the proposed design. The results show that the normalized locking time of the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method varies substantially with SH position (Figure 17(a)), especially at block boundaries (Figure 17(b)). When the SH is located at the Most Significant Bit (MSB), typical sliding detection requires about 1.8 times the time needed at the Least Significant Bit (LSB), indicating strong sensitivity to the starting position and search path. In contrast, the proposed ISS scheme maintains a stable normalized locking time within 1.0 ± 0.05 across all positions, with the standard deviation reduced by more than 70%. By evaluating all candidate positions equally through parallel filtering, the scheme eliminates position dependence. Synchronization can be completed within tens of clock cycles whether the SH is located at the LSB, the MSB, or any other position in the block. The experimental results verify that the ISS algorithm improves synchronization robustness and predictability while accelerating link initialization. Table 3 summarizes the performance metrics. The average locking time is only 24.4 clock cycles, representing an overall improvement of more than 70% compared with the single-bit serial locking circuit, 66-bit serial locking architecture, and register-intensive block synchronization method. The standard deviation of locking time is only 4.7, indicating a more stable synchronization process. In terms of resource utilization, the design consumes 509 Look-Up Tables (LUTs) and only 2.0 mW, much lower than the 3 503 LUTs and 94.1 mW required by the register-intensive scheme. Its energy efficiency reaches 0.03 mW/bit, which is better than those of the three conventional methods. Compared with the parallel search circuit, the average locking time is reduced by 6.11%, power consumption is reduced by 50.3%, and energy efficiency is improved by 53.8%. Therefore, the proposed JESD204C receiver link shows advantages in SH detection speed, stability, power consumption, and energy efficiency.  Conclusions  An ISS-based SH detection algorithm is proposed for the JESD204C receiver. By screening the data stream in parallel through multilevel XOR logic, dynamically compressing the search space, and efficiently eliminating non-inverting candidate positions, the algorithm converges to the true SH position. This approach improves the conventional serial detection mechanism. The design is verified on the Xilinx KC705 FPGA platform. A Pseudorandom Binary Sequence 31 (PRBS31) is used to emulate the random distribution of polarity transitions, and a high-speed SubMiniature version A (SMA) cable is used for data loopback transmission. The results show that the algorithm achieves an average locking time of only 24.4 clock cycles, with a standard deviation as low as 4.7. High robustness is maintained for the SH at any position within the 66-bit block, and the energy efficiency reaches 0.03 mW/bit. The algorithm is superior to existing typical schemes in locking speed, delay stability, and energy efficiency. It provides a low-latency, reliable, and energy-efficient synchronization initialization approach for high-speed SerDes links.
Survey on Intelligent Semantic Covert Communication
FENG Zhaoxin, XU Yifan, XING Chengwen, XU Yuhua, ZHAO Nan, WANG Jinlong
Available online  , doi: 10.11999/JEIT260184
Abstract:
  Significance   As the Sixth-Generation mobile communication network (6G) evolves from the Internet of Everything to the Intelligent Internet of Everything, the communication paradigm is shifting from reliable bit transmission to effective semantic transmission. Semantic communication extracts and compresses task-related semantics to reduce redundancy and resource use. However, because semantic information is highly structured and task-specific, it is vulnerable to eavesdropping, inference, and attacks. Covert communication addresses this risk by hiding transmission behavior from unauthorized monitoring. With support from Artificial Intelligence (AI), covert communication can use reinforcement learning to adjust power and resource allocation in dynamic environments. Generative models can also conceal transmitted signals by learning and reproducing environmental patterns. However, strict covertness constraints limit the achievable transmission rate and make large-scale information transmission difficult. Intelligent semantic covert communication integrates semantic extraction with covert transmission, providing a reliable approach to secure and efficient 6G communications.  Progress   With the development of AI, especially deep learning for complex feature modeling, semantic communication can support efficient semantic extraction and nonlinear compression of multimodal data. Research on semantic communication has also shifted from Separate Source-Channel Coding (SSCC) to Joint Source-Channel Coding (JSCC), which supports end-to-end training and improved transmission performance. For image transmission, Convolutional Neural Networks (CNNs) use local receptive fields to capture spatial correlations. For sequential data transmission, Long Short-Term Memory (LSTM) networks use gating mechanisms to maintain temporal coherence. In covert communication, Generative Adversarial Networks (GANs) and diffusion models can learn the statistical patterns of environmental noise in the time, frequency, and spatial domains, thereby concealing transmitted signals. These methods reduce the effectiveness of unauthorized monitoring and detection, and improve system adaptability in dynamic environments. AI also improves autonomous decision-making in dynamic covert communication. By modeling covert transmission as a Markov Decision Process (MDP), Deep Reinforcement Learning (DRL) can learn resource allocation strategies through interaction with the environment. This approach reduces computational complexity compared with traditional convex optimization methods. By integrating semantic extraction and covert transmission, intelligent semantic covert communication further supports semantic-driven covert transmission. Large Language Models (LLMs) can evaluate semantic sensitivity and contextual risks, enabling selective covert transmission of sensitive semantic information.  Conclusions  Research on intelligent semantic covert communication shows the advantages of coordinated semantic perception and physical-layer covert mechanisms. AI improves semantic extraction efficiency and strengthens adaptation to dynamic and complex environments. By integrating semantic understanding with covert transmission strategies, intelligent semantic covert communication supports both efficiency and security for ubiquitous 6G services.  Prospects   Future research on intelligent semantic covert communication should address several key challenges, including AI-enabled detection, unified semantic metrics, lightweight model design, multimodal semantic alignment, system interpretability, and semantic hallucination. Active threat detection and adaptive defense strategies are needed to counter AI-driven surveillance. Causal reasoning in Large Multimodal Models (LMMs) can help mitigate semantic hallucination and improve data transmission reliability. Advances in model compression and cloud-edge collaboration are also needed to deploy high-complexity AI models on resource-limited terminals. With the rapid development of AI, intelligent semantic covert communication is expected to provide core support for intelligent connectivity of everything and help build more secure, efficient, and reliable 6G networks.
A Point Cloud Slice-based UAV SLAM Method for 3D Reconstruction of Large Container Port Areas
HU Zhaozheng, ZUO Zhihang, XU Cong, TAO Qianwen, LIU Chao, MENG Jie
Available online  , doi: 10.11999/JEIT251112
Abstract:
  Objective  With the continuous development of port intelligence, the demand for digital management in container port areas has increased. In large container yards, Three-Dimensional (3D) reconstruction of the yard environment can be achieved using Unmanned Aerial Vehicle (UAV)-based Simultaneous Localization And Mapping (SLAM). However, container port areas contain many repetitive semantic structures. Traditional semantic matching methods therefore show low efficiency and limited accuracy. In addition, lanes between container yards form large feature-sparse regions during UAV-based 3D reconstruction, which can cause odometry degradation. Repetitive scene features also interfere with loop closure detection. To address these problems, this paper proposes a rapid feature extraction method based on point cloud slicing and further optimizes it according to the structural characteristics of container yards. A UAV point cloud slice-based SLAM method, termed Slice-SLAM, is proposed for high-precision 3D reconstruction of large container port areas.  Methods  To improve point cloud semantic extraction, a rapid point cloud slicing method is proposed. The principal direction is extracted rapidly, and the point cloud is divided into multiple layers to obtain multi-layer semantic point clouds efficiently. The slicing strategy is further optimized for container yard scenarios. Principal plane extraction is simplified using the gravity direction, and the elevation range of each container layer is obtained adaptively from point cloud density gradient changes. Multi-layer slice point clouds are then constructed. A progressive adaptive Light Detection And Ranging (LiDAR) odometry method based on slice point clouds is developed. Elevation slices are used to identify degenerate scenarios adaptively, and a layer-wise incremental slice matching and fusion strategy is used. This improves the accuracy, efficiency, and stability of LiDAR odometry. In addition, a factor graph optimization method that integrates slice point cloud information is designed. Fusion voting is performed on the matching results of multi-layer slice point clouds to remove erroneous matches and reduce the effect of repetitive structures on loop closure detection. Slice factors are then used to construct factor graph edges, which improves global optimization and supports efficient and stable 3D reconstruction.  Results and Discussions  The feasibility and effectiveness of the proposed method are verified in CARLA simulation scenarios and real-world tests at a large container port in Wuhan. First, comparisons with three semantic extraction algorithms, namely RANSAC, Region Growth, and 3DG_SEG, demonstrate the efficiency and accuracy of the proposed semantic extraction method. Second, estimated trajectories are compared with those obtained by two open-source LiDAR algorithms, FAST-LIO2 and Faster-LIO, confirming the advantages of the proposed odometry method. Finally, speed and confidence score are compared with those of six algorithms: ICP, NDT, GICP, Fast-GICP, Scan Context+ICP, and Quatro. The loop closure detection module of LIO-SAM is also integrated into FAST-LIO2, and the Scan Context module is integrated into Faster-LIO. The resulting estimated trajectories are compared with those of the proposed method, verifying the effectiveness of the proposed loop closure detection algorithm. The proposed method achieves high 3D reconstruction accuracy and is suitable for practical port operations.  Conclusions  The proposed method uses an efficient point cloud slicing technique and a multi-layer slice matching mechanism. Points within the same elevation range are defined as a slice point cloud, and the segmentation process is defined as point cloud slicing. This design enables efficient and robust 3D reconstruction in large-scale scenes with repetitive features. First, the LiDAR point cloud is aligned with the positive Z-axis using the gravity direction derived from the Inertial Measurement Unit (IMU). A sliding window records density gradient changes to determine the elevation range of each layer adaptively. This simplifies point cloud slicing and reduces the effects of non-standard containers and ground height variations on semantic extraction. Multi-layer slice information is then integrated into the odometry module to detect degenerate scenarios. Under normal conditions, progressive slice matching is used to initialize pose estimation. In degenerate scenarios, iterative Kalman filtering with increased IMU weighting is used. Finally, the fusion voting mechanism removes outliers from multi-layer slice matching results. The optimal match is used to initialize loop closure for global registration of container-region point clouds, enabling dual-stage loop closure detection and slice factor construction. By integrating slice point cloud information into factor graph optimization, the proposed method unifies point clouds in a common coordinate system and achieves efficient and robust 3D reconstruction.
An Inverse-Hybrid-Modeling Digital Twin System for Natural Gas Energy Metrology
LIU Bin, ZHONG Lu, FENG Quanyuan, CHEN Yihong
Available online  , doi: 10.11999/JEIT260289
Abstract:
  Objective  Global natural gas consumption continues to increase at an average annual rate of 3.2%. A 0.1% reduction in energy measurement error can reduce trade disputes by approximately $750 million per year. Traditional studies mainly use indirect methods for energy measurement. Among these methods, chromatographic analysis and acoustic velocity correlation are the most widely used, but both have clear application limits. Chromatographic analysis has a low interference error, but it shows delayed dynamic response at high flow rates and limited dynamic calibration capability. It also has poor adaptability to multi-gas-source switching, requires manual calibration, and has high operation and maintenance costs. The lack of interoperability standards for energy networks further increases the difficulty of system integration. Acoustic velocity correlation provides a low-latency dynamic response for flow measurement, but it has a high interference error. This error may increase when the content of a single component changes, such as when the hydrogen content increases from 5% to 10%. The method may even fail under complex operating conditions, such as multi-gas-source mixing and dynamic pressure fluctuations. To address these issues, new mechanism-modeling-oriented methods have been developed. The two most representative directions are mechanism-modeling-driven methods and hybrid-modeling methods. Both methods combine multi-source data fusion with virtual-physical interaction to establish mechanism models that link flow rate, other parameters, and energy. These methods provide a new approach for accurate energy measurement, but new challenges remain. Mechanism-modeling-driven methods are usually based on static flow modeling using Computational Fluid Dynamics (CFD). However, their dynamic parameter updates are slow, with delays of more than 30 s. They also have difficulty adapting to real-time operating-condition changes, rely on large labeled datasets, and have limited interpretability. Hybrid-modeling methods still face unresolved problems in collaborative optimization across multiple modules. In addition, existing studies lack support from industrial-grade verification platforms. These limits restrict their ability to solve the dynamic response delay, parameter identification difficulty, excessive physical simplification, and weak interference resistance of traditional natural gas energy metrology methods under complex conditions. Based on recent progress in mechanism-modeling-driven and hybrid-modeling methods, this study proposes an inverse-hybrid-modeling-driven digital twin system. The system introduces a Variational AutoEncoder (VAE)-based operating-condition feature extraction algorithm and a Dynamic Bayesian Network (DBN)-based parameter calibration mechanism. It also uses a Variational Expectation-Maximization (VEM) algorithm for offline calibration. The proposed system aims to improve the accuracy, adaptability, and interference resistance of natural gas energy metrology under complex operating conditions.  Methods   A natural gas energy metrology digital twin system based on inverse hybrid modeling is proposed. The system is built on a three-tier “algorithm-system-scenario” architecture. It integrates calorific value, flow, and energy mechanism models with multi-source real-time data streams. The VAE is used for unsupervised mining of operating-condition features. A parameter self-correction loop is then constructed by combining the DBN with VEM-based system calibration. Industrial-grade devices, including ultrasonic flowmeters and gas chromatographs, are integrated to ensure real-time data transmission and closed-loop control. The system covers key operating conditions, including dynamic pressure fluctuations, hydrogen-blended gas mixtures, and multi-gas-source switching. This design ensures strong adaptability between the model and practical applications. The system was continuously verified for 25 weeks on a full-scale industrial-grade experimental platform. The results show an operational delay of ≤3.8 s, data transmission jitter of ≤0.5 s, average daily energy consumption per device of ≤1.2 kW·h, Mean Time Between Failures (MTBF) of ≥4 100 h, energy measurement error of ≤0.25%, calorific value error of ≤0.12%, and flow indication error of ≤0.2%. The system also meets security requirements through industrial Ethernet encryption and hierarchical access control. It provides engineering support for intelligent pipeline-network optimization and standardized integration.  Results and Discussions  First, a multi-level hybrid modeling framework is established. Modular hybrid modeling is achieved through the algorithm-system-scenario three-tier architecture. Numerical methods combined with data are more flexible than purely analytical models and can represent complex multiphysics systems with fewer lumped physical parameters. These parameters may change during energy measurement under mechanical, energy, and hydrodynamic effects. The VAE and DBN are used to deeply integrate mechanism models with real-time data. This reduces the parameter synchronization delay to 3.8 s and supports fluid-acoustic co-simulation and rapid response under complex operating conditions, such as hydrogen-blended natural gas. Second, an integrated algorithm for inverse hybrid modeling and system calibration is proposed. By incorporating the VAE, DBN, and VEM algorithm, the inverse hybrid modeling algorithm forms a self-supervised, adaptive intelligent system with an internal closed-loop operation. The VAE encoder compresses high-dimensional operating-condition data into low-dimensional feature vectors. This enables unsupervised feature extraction without large labeled datasets. Based on the learned internal data distribution, the VAE can also generate perturbed data similar to the input data. These data are used to simulate abnormal operating conditions and verify interference resistance. The DBN constructs a continuous “prior-evidence-posterior” iterative cycle to support system self-correction and adaptive response to operating-condition changes. The VEM algorithm compensates for systematic errors that are difficult for the DBN to capture, thereby overcoming the limits of traditional static models.  Conclusions  This study describes and validates a hybrid digital twin system that combines experimental data-driven methods with physical models. The system successfully simulates the physical characteristics of natural gas energy metrology. A full-scale test platform was constructed, and the main system parameters were validated using experimental measurement data and compared with industry benchmarks. Each independent module in the algorithm-system-scenario three-tier hybrid modeling architecture, including calorific value measurement, flow calculation, and energy conversion, was continuously verified for 25 weeks. The results confirm strong consistency between model predictions and actual measurements. On the natural gas energy metrology digital twin experimental platform, systematic validation was performed for three core functions: flow measurement under dynamic conditions, multi-component calorific value determination, and energy accumulation. The results show that the output of the digital twin model matches the physical device measurement data with an accuracy of more than 99.5%. Under complex operating conditions, such as pressure pulsations and hydrogen-blended gas mixtures, the system maintains the measurement error within 0.5%. This performance is better than that of traditional methods and meets the Class A accuracy requirements for natural gas measurement. By introducing a multi-tier hybrid modeling framework, this study addresses the parameter identification difficulty and excessive physical simplification of traditional natural gas energy metrology methods. The integration of the VAE, DBN, and VEM algorithm enables unsupervised feature extraction under complex operating conditions and adaptive calibration of model parameters. This reduces dependence on prior physical knowledge and large labeled datasets. The experimental results show that the proposed method maintains high precision and strong stability under complex scenarios, including pressure pulsations and hydrogen-blended gas mixtures, where traditional models have difficulty providing accurate descriptions.
A Social-Aware Ant Colony Optimization Algorithm with Reproductive Division of Labor for MCS Task Allocation
SHEN Xiaoning, SHE Juan, WANG Zhilong, LI Jiayuan
Available online  , doi: 10.11999/JEIT260018
Abstract:
  Objective  With the rapid development of handheld and wearable smart devices, Mobile Crowd Sensing (MCS) has become an efficient data collection paradigm. Effective task allocation can improve system efficiency, requester and participant satisfaction, and platform sustainability. Existing models often neglect task skill requirements, do not use participants’ social networks as auxiliary execution resources in emergencies, and overlook the effect of collaboration efficiency on team-task quality. To address these issues, this paper proposes a Social-Aware MCS Task Allocation model (SAMCSTA) with two objectives: maximizing total platform revenue and total task sensing quality. Social networks are used to build a two-layer collaboration framework of platform participants and social-network friends, which expands available execution resources and improves allocation flexibility. For complex tasks, participant sensing capability is quantified, and collaboration efficiency is introduced to optimize team composition.  Methods  This paper proposes a Multi-objective Ant Colony Optimization based on Reproductive Division of Labor (MACORDL) algorithm. The main innovations are as follows. First, the ant colony is divided into four collaborative subpopulations: queen ants, male ants, scout ants, and worker ants. Local enhancement, memetic crossover, knowledge transfer, and other search strategies are designed for these subpopulations to form a hierarchical collaborative search framework. Second, a statistical-learning-based mating selection strategy is designed to support intelligent transfer of elite genes. Third, the short-term contribution of each subpopulation is predicted from historical performance, which enables dynamic and adaptive allocation of computational resources. Fourth, a cooperative update mechanism for node pheromones and participant pheromones is designed to establish a dual-layer search guidance system.  Results and Discussions  The evaluation uses 8 synthetic instances and 4 real-world instances. Performance is measured by HyperVolume Ratio (HVR) and Inverted Generational Distance (IGD). The Wilcoxon rank-sum test at a significance level of 0.05 is used for statistical comparison. The results show that MACORDL achieves the best HVR and IGD on most instances (Table 2, Table 3). On average, MACORDL improves HVR and IGD by 16.41% and 18.04%, respectively, compared with the second-best algorithm. Visual comparisons further show that the Pareto front obtained by MACORDL has better convergence, distribution uniformity, and breadth (Fig. 4). Although its fine-grained local search can still be improved for a few large-scale instances, MACORDL shows stable performance and good scalability across different problem scales. It helps the platform obtain task allocation schemes with higher revenue and better sensing quality.  Conclusions  This paper studies the task allocation problem in MCS systems by considering interactions among platform participants and between participants and their social-network friends. A social-aware MCS task allocation model is established, and MACORDL is proposed to solve it. Comparative experiments on 8 synthetic instances and 4 real-world instances with different scales show that MACORDL outperforms six representative algorithms on most instances. It obtains allocation schemes and paths that yield higher total platform revenue and better task sensing quality, indicating good scalability. MACORDL uses multiple strategies to balance local exploitation and global exploration. However, the current model assumes that all tasks are released at the initial stage and that complete information is available. Participant privacy protection is also not considered. Future work will focus on MCS task allocation models in dynamic and uncertain environments and on privacy-preserving distributed optimization.
Non-Terrestrial Network Architecture and Key Technologies for Civil Aviation
LIU Xiangnan, QIU Yu, HUANG Zhipeng, ZHANG Haijun
Available online  , doi: 10.11999/JEIT260348
Abstract:
  Significance   Civil aviation communication systems are entering a new stage of development driven by the rapid growth of global air transportation, the increasing demand for intelligent air traffic management, and the continuous expansion of in-flight connectivity services. Traditional civil aviation communication systems mainly rely on high frequency radio, high frequency radio, terrestrial air-to-ground links, and conventional satellite communication systems. These technologies have supported aircraft operation, air traffic control, airline operational communication, and low-rate data transmission for a long time. However, they still face limitations when applied to future civil aviation scenarios characterized by global coverage, high-speed mobility, low latency, high reliability, and service diversification. Particularly, terrestrial networks are difficult to deploy in transoceanic routes, polar regions, deserts, mountains, and remote airspace, while traditional geostationary satellite systems suffer from large propagation delay and limited capacity. Current systems cannot fully meet the requirements of continuous aircraft access, real-time flight monitoring, engine health data transmission, aviation safety communication, and passenger broadband services. Non-Terrestrial Networks (NTNs) provide a promising technical path for overcoming these limitations. By integrating GEOstationary satellites (GEO), Medium Earth Orbit satellites (MEO), low Earth orbit satellites (LEO), Very Low Earth Orbit satellites (VLEO), High-Altitude Platform Stations (HAPS), Unmanned Aerial vehicles (UAV), electric Vertical Take Off and Landing (eVTOL), and terrestrial infrastructures, NTN can construct a multi-layer air-space-ground integrated communication system. Such a system is able to provide continuous coverage, flexible deployment, resilient connectivity, and differentiated service support for civil aviation. NTN is becoming an important enabling technology for future civil aviation communication systems and for the digital and intelligent transformation of the aviation industry.  Progress   This paper reviews the development of NTN technologies for civil aviation and summarizes key research progress from three aspects: network architecture, access and mobility management, and resource management and scheduling. (1) We propose an aviation-oriented NTN networking framework composed of three layers: the satellite edge layer, the airborne core layer, and the terrestrial assistance layer. The satellite edge layer includes GEO, MEO, LEO, and VLEO satellites connected through inter-satellite links. GEO satellites are suitable for wide-area broadcasting and non-real-time services, MEO satellites can support navigation and intermediate-delay services, LEO satellites are suitable for low-latency and high-capacity broadband access, and VLEO satellites can further reduce propagation delay for future near-real-time aviation applications. The airborne core layer includes civil aircraft, HAPS, UAVs, and eVTOL platforms. HAPS can act as a regional relay, edge computing node, or software-defined control carrier, while UAVs and eVTOL platforms can provide flexible low-altitude coverage, emergency communication, and local access support. The terrestrial assistance layer consists of terrestrial base stations and gateway stations, which support air-to-ground communication and satellite-terrestrial interconnection. Civil aviation services can be divided into air traffic control and air traffic management services, airline operational control services, and airline passenger communication or in-flight entertainment services. Through network slicing, these heterogeneous services can be logically isolated and managed over a shared air-space-ground infrastructure. In congestion, rain attenuation, or shortened visibility-window scenarios, safety slices should be protected with the highest priority, while passenger service slices can be rate-limited, buffered, or degraded. (2) We analyze the characteristics of NR-NTN access and air-to-ground direct access in civil aviation. NR-NTN can provide continuous coverage for oceanic, polar, desert, and remote flight routes through satellites or HAPS, while air-to-ground direct access can provide low-latency and high-rate links in areas where terrestrial base stations can be deployed. However, aircraft differ significantly from ordinary terrestrial terminals because their flight trajectory, altitude, speed, and route are highly predictable. Therefore, the key issue in aviation NTN access is not only how to execute random access, but how to predict the access window, timing compensation, frequency offset, and target access node before the aircraft enters the coverage area. By using satellite ephemeris, Global Navigation Satellite System information, aircraft trajectory, and velocity parameters, civil aircraft can predict satellite visibility and pre-compute timing advance, scheduling offset, and Doppler compensation before initiating access. This transforms random access from a passive response process into a proactive and predictive access process, thereby improving access certainty and synchronization stability in highly dynamic aviation scenarios. For mobility management, a signaling interaction process for aircraft handover is designed. Based on trajectory prediction and satellite visibility prediction, the network can select a target satellite or gateway with longer residence time and better service capability. Before the aircraft reaches the handover boundary, the source and target network sides can complete context preparation, user-plane path preparation, radio resource reservation, and protocol data unit session update. When the handover condition is triggered, the aircraft performs random access to the target satellite or beam and then switches the user-plane path. This “prediction–preparation–fast handover” mechanism can reduce service interruption and maintain session continuity. For safety-critical traffic, priority and isolation policies should remain consistent and auditable throughout session preparation, handover execution, and path switching. (3)We discuss computing and caching resource management in civil aviation NTN. As onboard computing capability is limited and aviation applications generate increasing computing demands, NTN can provide mobile edge computing and caching services through LEO satellites, HAPS, UAVs, and inter-satellite cooperation. The paper introduces several computing offloading modes, including on-orbit satellite collaborative offloading, network-level integrated offloading, and cloud-edge-terminal hybrid offloading. These mechanisms can support tasks such as aviation monitoring, trajectory analysis, intelligent inference, and in-flight service optimization. In addition, caching mechanisms such as onboard satellite caching, inter-satellite cooperative caching, and named-data-networking-based content caching can improve content delivery efficiency and service continuity. Cache placement should consider content popularity, regional demand prediction, visibility windows, cache prefetching, and cooperative cache sharing among different satellite layers.  Conclusions   NTN can effectively complement traditional civil aviation communication systems by filling coverage gaps in remote and oceanic airspace, enhancing service continuity, and supporting differentiated aviation services. The proposed aviation-oriented NTN architecture integrates multi-orbit satellites, HAPS, UAVs, civil aircraft, and terrestrial infrastructures into a unified framework. The on-demand isolated slicing mechanism can provide differentiated protection for ATC/ATM, AOC, and APC/IFE services. Ephemeris-map-assisted access and predictive mobility management can improve access reliability and reduce handover interruption in high-speed aviation scenarios. Computing offloading and cooperative caching further enhance the ability of NTN to support intelligent and data-intensive aviation applications.  Prospects   Future civil aviation NTN should evolve toward deeper integration of low-altitude networks, space networks, and terrestrial networks. Cross-domain topology visualization, link-state sharing, policy distribution, and programmable logical networks are essential for improving controllability and scalability. In mobility management, integrated cross-domain handover mechanisms should be developed to cope with satellite beam switching, terrestrial cell handover, and air-to-air relay reconstruction. In resource management, communication, navigation, computing, and caching resources should be jointly scheduled and transformed according to aviation service requirements. With continuous advances in NTN architecture, network slicing, predictive access, mobility management, computing offloading, and caching, NTN is expected to provide more efficient, stable, and intelligent communication support for civil aviation and to promote the digital transformation of future air transportation systems.
Data-driven Sliding-mode Disturbance-rejection Formation Control for Quadrotor UAV Swarms Under Uncertain Disturbances
LI Qianxiong, LU Xiaoqing
Available online  , doi: 10.11999/JEIT260050
Abstract:
  Objective  Quadrotor Unmanned Aerial Vehicle (UAV) cooperative formation can increase payload capacity and extend the operational range. However, quadrotor UAVs are highly nonlinear and underactuated systems. Differences in size and actuator hardware further weaken the effectiveness of model-based formation-control methods. Therefore, disturbance-rejection formation control is needed for quadrotor UAV swarms with unknown internal models and uncertain external disturbances.  Methods  To address the difficulty of precise modeling for quadrotor UAV swarm formation under uncertain disturbances, this paper proposes a data-driven sliding-mode disturbance-rejection formation control method. First, a data-driven formation-control model is established using the input and output states of each UAV and its neighboring UAVs. Then, an extended state observer and an integral sliding-mode formation controller are designed to estimate uncertain disturbances online and achieve robust formation control. Finally, stability analysis is conducted to derive sufficient conditions under which all UAVs achieve sliding-mode disturbance-rejection formation. The proposed method is verified through simulations and experiments under an unknown system model and uncertain disturbances.  Results and Discussions  The simulation results show that multiple quadrotor UAVs can maintain the desired formation geometry in a wind-disturbed environment (Fig. 4). The formation position error converges to within 0.1 m in 15 s and reconverges rapidly after a 7 m/s gust is applied (Fig. 6). The velocity curves also show rapid convergence among the UAVs (Fig. 5). The experimental results indicate that three UAVs can follow the trajectory of the virtual leader while maintaining the desired triangular formation (Fig. 17). The formation error is mostly kept within 0.1 m (Fig. 18). When the observation matrix fluctuates strongly between 10 s and 20 s, the corresponding formation error is relatively large. When the observation matrix curve becomes smoother between 20 s and 30 s, the formation error also decreases (Fig. 20). Compared with traditional model-based formation-control methods and existing data-driven methods, the proposed method reduces the formation error by 41% and shortens the formation response time by 40%.  Conclusions  This paper proposes a data-driven sliding-mode disturbance-rejection formation control method for quadrotor UAV swarms with unknown internal models and uncertain external disturbances. Under an unknown quadrotor UAV model and a 7 m/s wind disturbance, the proposed method keeps the formation error below 0.1 m. It also reduces the formation error by 41% and shortens the formation response time by 40% compared with traditional model-based formation-control methods and existing data-driven methods. Future work will study multilayer data-driven formation control for heterogeneous UAV-UGV swarm systems. It will also optimize computational cost and scalability in large-scale and complex application scenarios.
Research on Secure and Covert Transmission for UAV-assisted Visible Light Communication Systems
WU Mengru, LIN Jiale, LU Weidang, LI Bo, GUO Lei
Available online  , doi: 10.11999/JEIT260239
Abstract:
  Objective  Unmanned Aerial Vehicles (UAVs) can serve as aerial base stations for Visible Light Communication (VLC) because of their mobility and on-demand coverage capabilities. However, air-ground communication links are exposed to open environments, which makes VLC vulnerable to data eavesdropping and malicious detection. To address this issue, this paper proposes a secure and covert transmission strategy for a UAV-assisted VLC system from the perspectives of Physical Layer Security (PLS) and Covert Communication. The proposed strategy jointly optimizes UAV transmit power and hovering altitude to maximize the system secrecy capacity. The optimization is subject to covert communication requirements, illumination requirements, and operational constraints on UAV transmit power and hovering altitude.  Methods  This paper investigates secure and covert communication in a UAV-assisted VLC system. A UAV-assisted VLC system model is first established. In this model, a mobile UAV equipped with a Light-Emitting Diode (LED) is used to establish a VLC link with a legitimate ground user in the presence of an eavesdropper (Eve) and a warden (Willie). An optimization problem is then formulated to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude. To solve this problem, a Two-Layer OPtimization (TLOP) algorithm is proposed. The transformed problem is decomposed into two subproblems: an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude design problem. A closed-form expression for the optimal transmit power is derived for the inner-layer problem. A Particle Swarm Optimization (PSO) algorithm is then developed to solve the outer-layer problem.  Results and Discussions  In the simulations, the proposed optimization scheme is compared with two baseline schemes. First, the convergence of the proposed TLOP algorithm is verified (Fig. 3). The results show that the algorithm converges rapidly within a limited number of iterations. Second, the optimal UAV hovering altitude with respect to the UAV horizontal coordinates is illustrated under the spatial distribution (Fig. 4). The results indicate that the optimal hovering altitude decreases as the UAV approaches the legitimate ground user. The secrecy capacity with respect to the UAV horizontal coordinates is then presented (Fig. 5). The secrecy capacity increases as the UAV approaches the legitimate ground user. This is because the legitimate VLC channel gain increases when the UAV is closer to the user. In contrast, when the UAV approaches Eve and Willie, the security and covertness constraints become stricter. The UAV is then forced to reduce its transmit power or increase its hovering altitude, which decreases the system secrecy capacity. Furthermore, the secrecy capacity of all schemes increases as ϵ increases (Fig. 6). This is because a larger ϵ relaxes the covertness requirement. The UAV can therefore adjust its hovering altitude and transmit power more flexibly to increase the system secrecy capacity. In addition, the secrecy capacity decreases as the number of symbols increases (Fig. 7). This occurs because more symbols provide Willie with more signal samples for detection, thereby improving Willie’s detection capability. Finally, the secrecy capacity of all schemes decreases as the uncertainty-region radius of illegal nodes increases (Fig. 8). This trend occurs because greater location uncertainty forces the UAV to address potential threats over a wider area. The UAV must therefore adopt a more conservative strategy under worst-case eavesdropping and detection conditions. Overall, the simulation results confirm that the proposed scheme improves the secrecy capacity of the UAV-assisted VLC system.  Conclusions  This paper investigates secure and covert communication in a UAV-assisted VLC system. The objective is to maximize the system secrecy capacity by jointly optimizing UAV transmit power and hovering altitude under covert communication, illumination, transmit power, and hovering altitude constraints. Because the formulated problem is highly non-convex, a PSO-based TLOP algorithm is designed to solve it. The proposed algorithm decomposes the problem into an inner-layer transmit power optimization problem and an outer-layer UAV hovering altitude optimization problem. Simulation results show that the proposed algorithm converges rapidly and improves the system secrecy capacity compared with the baseline schemes.
KE-HNS: Knowledge-Enhanced Personalized Recommendation Model with Hierarchical Noise Suppression
XIE Jun, WANG Dantong, ZHANG Bo, CHEN Guijun, LV Jiaqi, LUO Xiongyan
Available online  , doi: 10.11999/JEIT260051
Abstract:
  Objective  In the era of Big Data and Artificial Intelligence (AI), rapid information growth has increased the difficulty of filtering valuable content from redundant data. Personalized recommender systems are key tools for accurate information matching and resource allocation. Knowledge Graphs (KGs) can enrich user-item representations. However, current KG-based recommendation models still face weak noise suppression, coarse-grained user-interest modeling, and imbalanced use of heterogeneous information, which reduce recommendation accuracy. This paper proposes Knowledge-Enhancedpersonalized recommender model with HierarchicalNoise Suppression (KE-HNS), which integrates knowledge enhancement with hierarchical noise suppression. By combining graph representation learning and contrastive learning, KE-HNS addresses noise interference, fine-grained preference modeling, and multi-source information balance, thereby improving recommendation performance.  Methods  KE-HNS adopts a hierarchical noise-suppression paradigm. At the input stage, Input Noise Reduction (INR) is used to reduce noise from two sources. For user-item interactions, a learnable binary mask matrix is used to remove noisy edges. For KG denoising enhancement, triples are scored by importance, low-score triples are identified with a Bottom-K strategy, and noisy triples are masked. At the feature-fusion stage, Isolated Noise Suppression (INS) is used to preserve spatial independence by partitioning entity-attribute spaces according to relation type. This design limits high-order noise propagation and semantic contamination. At the representation-optimization stage, Comparative Noise Suppression (CNS) is implemented through contrastive learning to suppress irrelevant entity noise and strengthen robust semantic signals. To capture fine-grained user interests, Graph Convolutional Networks (GCNs) are used to enhance user representations from historical interactions and related entities. Adaptive weight layers further refine item representations by using entity attributes and relations. To balance heterogeneous information, a dual-view contrastive learning mechanism is constructed between the user-item view and the item-entity view. Positive and negative sample pairs are used to adaptively adjust the weights of different information sources. Finally, user and item representations are matched by inner product to generate the Top-K recommendation list.  Results and Discussions  KE-HNS is evaluated on three public datasets, Book-Crossing, MovieLens-1M, and Last.FM, through performance comparison, ablation experiments, denoising evaluation, case analysis, and complexity assessment. For Click-Through Rate (CTR) prediction, KE-HNS outperforms the best baseline models by 0.94%~1.01% in Area Under the Curve (AUC) and 0.43%~0.90% in F1-score (Table 3). For Top-K recommendation, its Recall@K is higher than those of most advanced methods across nearly all K values, with only a slight gap behind CG-KGR on Last.FM (Fig. 7). The ablation results show that all three denoising components contribute to the performance gains (Table 4). The denoising evaluation shows that KE-HNS effectively suppresses noise and maintains high prediction accuracy under noisy conditions (Fig. 8). The complexity analysis further indicates that the model remains feasible for practical deployment (Table 5).  Conclusions  This paper presents KE-HNS, a personalized recommendation model that combines knowledge enhancement with hierarchical noise suppression. By reducing noise interference and balancing collaborative filtering signals with knowledge-aware semantics, KE-HNS improves recommendation accuracy across multiple benchmark datasets. The model still has limitations in computational efficiency and depends on the coverage and completeness of the KG. Future work may focus on computational optimization and dynamic knowledge integration.
Energy-Efficient Trajectory Planning and Resource Optimization for UAV Relay Communications over Hybrid RF/FSO Links
LI Baolong, PAN Wenwei, JIANG Hao, FENG Simeng, WU Qihui
Available online  , doi: 10.11999/JEIT260139
Abstract:
  Objective  In low-altitude communication networks, hybrid Radio Frequency/Free-Space Optical (RF/FSO) Unmanned Aerial Vehicle (UAV) relaying can ease RF spectrum congestion and improve uplink data aggregation. However, in obstacle-rich urban environments, FSO backhaul links are vulnerable to blockage and intermittent outages. This creates a severe mismatch between the RF access-link rate and the FSO backhaul-link rate. UAV trajectory planning is also constrained by obstacle avoidance and flight dynamics. To address these coupled issues, this paper investigates an energy-efficiency maximization problem. Multiuser Non-Orthogonal Multiple Access (NOMA)-based RF access and the Three-Dimensional (3D) obstacle-avoiding UAV trajectory are jointly optimized, and buffer-assisted RF/FSO rate decoupling is incorporated.  Methods  A time-slotted UAV relaying model is considered, in which multiple ground users upload data to the UAV through an RF link using NOMA. The UAV decodes superposed signals by Successive Interference Cancellation (SIC), and the decoding order in each slot is determined according to the received-power ranking. The successfully received data are then forwarded to a Base Station (BS) through an FSO backhaul link. Urban blockage is modeled using 3D geometric obstacles. A visibility test is used to determine whether each relevant link is in Line-Of-Sight (LOS) or Non-Line-Of-Sight (NLOS), which captures the spatially correlated and time-varying RF access-link rate and intermittent FSO backhaul capacity. To suppress blockage-induced rate mismatch between the RF access link and the FSO backhaul link, an onboard finite-capacity buffer is deployed at the UAV. In each slot, the forwardable data amount is jointly limited by the instantaneous FSO backhaul capacity and the data available in the buffer, and buffer-capacity constraints are imposed to prevent overflow. System energy efficiency is defined as the ratio of cumulative data successfully delivered to the BS over the mission horizon to UAV propulsion energy consumption. Propulsion power is modeled as a function of UAV velocity and acceleration to reflect the effect of flight dynamics. Under 3D flight-region boundaries, prescribed start and end locations, discrete-time kinematic equations, maximum velocity and acceleration limits, and obstacle collision-avoidance constraints, a non-convex optimization problem is formulated. The decision variables are cross-slot multiuser transmit powers and the 3D UAV trajectory. An alternating optimization framework is then developed. For a fixed trajectory, propulsion energy is fixed, so maximizing energy efficiency is equivalent to increasing end-to-end successfully forwarded data. This yields a power-optimization subproblem. Because of NOMA coupling and logarithmic rate expressions, this subproblem remains non-convex and is solved by Successive Convex Approximation (SCA). For fixed transmit powers, Particle Swarm Optimization (PSO) is used to search candidate 3D trajectories in continuous space. To ensure feasibility under strict dynamics and safety constraints, Quadratic Programming (QP) projection is used to enforce velocity and acceleration constraints. Collision checks are performed for trajectory waypoints and inter-slot line segments to ensure obstacle-free flight. These two optimization procedures are performed alternately. The resulting joint design satisfies flight-dynamics feasibility and collision-avoidance requirements and improves energy efficiency.  Results and Discussion   Simulations are conducted in an urban airspace with multiple users, a BS, and dense 3D obstacles. Blockage causes frequent LOS/NLOS switching as the UAV moves. Fig. 2 and 3 compare the 3D trajectory and its planar projection, respectively. Compared with the initial trajectory, the optimized trajectory shows clear detours and necessary altitude adjustments. It achieves collision-free flight while satisfying velocity and acceleration constraints, thereby verifying the feasibility and safety of the proposed trajectory planning method. Fig. 4 shows the convergence of energy efficiency under different user transmit-power budgets. The proposed alternating optimization generally stabilizes within a small number of outer iterations. The converged energy efficiency increases with the power budget, indicating synergy between power control and trajectory adaptation. Fig. 5 shows buffer evolution over time. The buffer gradually accumulates data when the backhaul is blocked or experiences strong fading. It is quickly drained when the UAV enters regions with LOS backhaul and improved FSO capacity. To quantify buffering gain, Fig. 6 compares system energy efficiency between the proposed buffering mechanism and the no-buffer scheme. The proposed mechanism enables store-and-forward temporal smoothing during backhaul interruptions and improves system energy efficiency. Fig. 7 shows energy-efficiency convergence under different buffer capacities. As buffer capacity increases, the converged energy-efficiency level improves. A larger buffer enhances the UAV’s ability to temporarily store incoming data and reduces data accumulation and transmission blockage when RF access-link and FSO backhaul-link rates are mismatched or the backhaul link is constrained. Figure 8 compares four benchmark schemes, namely a non-optimized baseline, a power-optimization scheme, a trajectory-optimization scheme, and the proposed joint power-and-trajectory optimization scheme. The coordinated design of power allocation and obstacle-avoiding trajectory improves end-to-end energy efficiency. Trajectory optimization also plays a more dominant role under blockage-limited conditions.  Conclusion  This paper investigates a hybrid RF/FSO UAV relaying scheme with NOMA and an onboard buffering mechanism for low-altitude urban communication. Given dense obstacles, frequent blockage, FSO-link susceptibility, and strict flight-dynamics constraints, an energy-efficiency maximization problem is formulated for the joint optimization of multiuser NOMA power allocation and UAV trajectory. An SCA-based power-allocation method and an obstacle-avoiding trajectory design that combines PSO with QP projection are developed. The obtained trajectory satisfies flight-dynamics feasibility and collision-avoidance requirements and improves throughput per unit propulsion energy. Simulation results show that the planned trajectory can avoid obstacles, and that the onboard buffer provides an effective cushion between RF access and FSO backhaul to mitigate rate mismatch. The proposed method consistently outperforms benchmark schemes in energy efficiency. Trajectory optimization is also shown to be generally more effective than power allocation in improving overall system performance.
Secure and Covert MIMO Short packet Communication with Location-Uncertain Malicious Nodes
TIAN Bo, YANG Weiwei, YANG Xiaoqin, BAI Mengmeng
Available online  , doi: 10.11999/JEIT260059
Abstract:
  Objective  This paper investigates secure and covert short-packet communication in Multiple-Input Multiple-Output (MIMO) wireless systems with location-uncertain malicious nodes over quasi-static Rician fading channels. In the considered scenario, a legitimate transmitter sends confidential short packets to a legitimate receiver. Meanwhile, multiple monitoring nodes (Willie nodes) attempt to detect whether transmission occurs, and multiple eavesdropping nodes (Eve nodes) attempt to intercept the confidential information. Because malicious nodes may remain silent and their exact locations are unavailable to the legitimate system, their spatial uncertainty poses major challenges to joint covertness and secrecy analysis. To address this problem, a unified analytical and optimization framework is established for secure and covert short-packet transmission. The framework is used to characterize the coupling among covertness, secrecy, and reliability and to improve the Average Effective Secrecy and Covert Rate (AESCR).  Methods  The transmitter adopts Singular Value Decomposition (SVD)-based precoding, and the legitimate receiver applies Maximum Ratio Combining (MRC) to enhance the legitimate link. Monitoring nodes and eavesdropping nodes are modeled as two independent Poisson Point Processes (PPPs) outside a circular protection zone centered at the transmitter. This model captures the spatial randomness of malicious nodes. For covertness analysis, each monitoring node is assumed to perform optimal Likelihood Ratio Test (LRT)-based detection with full knowledge of the system model, noise power, channel state, and codebook information. Using the Chernoff bound and the Bhattacharyya coefficient, a theoretical lower bound on the minimum detection error probability of a single monitoring node is first derived. Stochastic geometry is then combined with the distribution of the strongest monitoring node to obtain a tractable lower bound on the average minimum detection error probability. For secrecy analysis, the finite blocklength normal approximation is used to account for decoding error and information leakage penalties. The legitimate channel is statistically characterized under Rician fading conditions, and the strongest eavesdropping node is analyzed through stochastic geometry. Based on these results, an approximate analytical expression for the average secrecy rate is derived. AESCR is proposed as a comprehensive performance metric that jointly reflects reliability, secrecy, and covertness. Under the average covertness constraint and the short-packet length constraint, a joint optimization problem for transmit power and packet length is formulated. By using the monotonic properties of the objective function and the covertness constraint, the original coupled optimization problem is transformed into a one-dimensional search problem.  Results and Discussions  Simulation results verify the accuracy of the theoretical derivations and reveal the effects of key system parameters. Both the simulated average minimum detection error probability and its theoretical lower bound decrease as the packet length increases. Higher transmit power further reduces the detection error probability, indicating that excessive power makes transmission more exposed to monitoring nodes (Fig. 2). Increasing the number of monitoring-node antennas strengthens spatial reception capability and further degrades covertness (Fig. 2). Enlarging the protection zone improves covertness because malicious nodes are forced to remain farther away from the transmitter. However, increasing the monitoring-node density weakens this benefit by raising the probability that a strong monitoring node appears near the protection-zone boundary (Fig. 3). The average secrecy rate increases with packet length and gradually approaches the asymptotic secrecy-capacity upper bound because the finite blocklength rate penalty decreases as the packet length grows (Fig. 4). AESCR first increases and then decreases with packet length, confirming the existence of an optimal packet length. This behavior results from the tradeoff between the reduced finite blocklength penalty and increased detection exposure (Fig. 5). Higher malicious-node density and more malicious-node antennas degrade system performance because they enhance both monitoring and eavesdropping capabilities (Fig. 5). Relaxing the covertness constraint improves the achievable AESCR because the system can select a higher transmit power or a more favorable packet length (Fig. 6). Results under different Rician factors show that the proposed analytical framework is applicable to both Rician and Rayleigh fading conditions (Fig. 6). Increasing the number of legitimate receive antennas improves AESCR, and a larger transmit antenna array provides additional SVD precoding gain (Fig. 7). Compared with benchmark schemes, the proposed joint optimization of transmit power and packet length consistently outperforms the scheme with fixed packet length and power-only optimization. This result demonstrates the need to jointly balance reliability, secrecy, and covertness in MIMO short-packet transmission (Fig. 8).  Conclusions  This paper develops a stochastic-geometry-based analytical framework for secure and covert MIMO short-packet communication with location-uncertain multi-antenna malicious nodes. By deriving a lower bound on the average minimum detection error probability, obtaining an approximate analytical expression for the average secrecy rate, and proposing AESCR, the framework reveals the fundamental tradeoff among covertness, secrecy, and reliability under finite blocklength transmission. The results show that increasing the number of legitimate transmit and receive antennas improves secure and covert performance, whereas higher malicious-node density and more malicious-node antennas degrade system performance. The existence of an optimal packet length further shows that packet length and transmit power should be jointly designed. The proposed joint optimization method therefore provides an effective solution for secure and covert short-packet transmission in mission-critical and low-latency wireless systems.
Joint Optimization Method for Pairwise Constrained Projection Clustering Integrating a Two-row Simultaneous Update Strategy
ZHU Jianyong, CHEN Kun, YANG Hui, NIE Feiping
Available online  , doi: 10.11999/JEIT260111
Abstract:
  Objective  As data structures become increasingly complex, conventional unsupervised clustering methods often fail to achieve satisfactory performance. Semi-supervised clustering has therefore attracted growing attention because it uses limited prior information to improve clustering quality. However, existing methods have two major limitations. First, traditional constrained projection clustering algorithms usually use a two-step independent strategy, in which the projection matrix is learned before k-means clustering is performed. This separation allows projection errors to be propagated directly to the clustering stage, causing accumulated learning errors. In addition, applying pairwise constraints only during projection deviates from the goal of using prior information to guide clustering. Second, many existing methods, including spectral clustering-based approaches, handle pairwise constraints implicitly, for example through eigen-decomposition of a modified similarity matrix. Such implicit processing may not strictly satisfy the constraints, especially Cannot-Link (CL) constraints, which are non-transitive, resulting in high constraint violation rates. To address these issues, this paper proposes a joint optimization method for pairwise constrained Projection Clustering Integrating a Two-row simultaneous Update Strategy (PCITUS). The objective is to unify dimensionality reduction and clustering within a single framework to reduce information loss, while designing an explicit optimization strategy that lowers constraint violations and improves computational efficiency.  Methods  The proposed PCITUS model integrates constrained projection and clustering into a unified objective function for collaborative optimization, with pairwise constraints optimized directly. First, the algorithm uses the transitive property of Must-Link (ML) constraints. Samples belonging to the same ML connected component are merged into a single hyper-point in the feature space. This preprocessing step ensures that all ML constraints are naturally satisfied. A trade-off parameter is then introduced to incorporate projection learning into the clustering framework as a regularization term, allowing both components to be jointly optimized under one objective. Prior information is further embedded into the clustering process by transforming pairwise constraints into row-wise constraints on the indicator matrix. An improved coordinate descent method is then used to optimize the discrete indicator matrix directly, which improves computational efficiency and produces better clustering results. A key feature of PCITUS is the two-row simultaneous update strategy for CL constraints. PCITUS explicitly checks CL conflicts by simultaneously evaluating objective function values obtained by moving conflicting rows to suboptimal classes and then selects the case with the higher value.  Results and Discussions  Extensive experiments are conducted on eight benchmark datasets and compared with nine state-of-the-art semi-supervised clustering algorithms. Quantitative results based on ACCuracy (ACC) and Normalized Mutual Information (NMI) demonstrate the superiority of PCITUS (Table 4 and Table 5). PCITUS achieves the best performance on most datasets. In particular, on the Mushroom dataset, NMI is improved by 7.29% compared with the second-best algorithm. The comparison with CNP, a two-step projection method, confirms that the unified framework effectively reduces error propagation and information loss. This effect is also supported by the mutual reinforcement between projection and clustering: a better projection space produces a clearer clustering structure, while a more reasonable clustering structure guides the formation of a more discriminative projection space. The effectiveness of explicit constraint handling is further illustrated (Fig. 1). PCITUS produces no ML constraint violations because of the hyper-point merging strategy. For CL constraints, the two-row simultaneous update strategy enables PCITUS to maintain an extremely low violation rate, such as 0.57% on Mushroom and 0.41% on Satimage, greatly outperforming methods that handle constraints implicitly. Additionally, the parameter sensitivity analysis (Fig. 2) shows that PCITUS remains stable across a wide range of trade-off parameter values. The noise sensitivity experiments (Fig. 3a and Fig. 3b) confirm its robustness. The convergence curves (Fig. 3c and Fig. 3d) and runtime comparisons (Table 7) further verify its computational efficiency, showing rapid convergence and a stable objective function value within approximately 10 iterations in most cases.  Conclusions  This paper presents PCITUS, a semi-supervised clustering framework that jointly optimizes pairwise constrained projection and clustering structures. The method addresses the difficulty of optimizing CL constraints and overcomes the limitations of traditional constrained projection clustering frameworks based on a two-step separation scheme. By integrating the projection objective into the clustering framework as a regularizer, the proposed method enables subspace learning and data partitioning to reinforce each other and jointly approach the global optimum. Pairwise constraints are used throughout the learning process, allowing prior knowledge to guide optimization more fully. The coordinate descent method with the two-row simultaneous update strategy directly and accurately allocates samples under CL constraints, significantly reducing constraint violations. Experimental results show that PCITUS outperforms existing algorithms in clustering performance.
MG-MoE: Routed Multi-Granularity Expert Ensemble
XIAN Fengyu, JIAN Haifang, XIE Zihui, DU Jun, ZHANG Yuanyuan, NING Xin, DONG Miaomiao, WANG Hongchang
Available online  , doi: 10.11999/JEIT260219
Abstract:
  Objective  Fine-Grained Image Recognition (FGIR) aims to distinguish visually similar subcategories that differ only in subtle local patterns. It must also remain robust to large intra-class variations caused by pose changes, occlusion, illumination shifts, and complex backgrounds. In real-world scenarios, these challenges are further intensified by long-tailed category distributions. Rare or difficult classes are more likely to overfit spurious contextual cues and suffer from unstable decision boundaries. Therefore, a conditional computation paradigm is needed, in which complementary inductive biases are separated into specialized expert branches and adaptively combined for each sample. This work aims to develop a routed multi-granularity mixture-of-experts framework that improves discriminative performance under controllable inference cost. It also enhances robustness for difficult samples and long-tailed categories through adaptive sparse expert activation.  Methods  A Multi-Granularity Mixture-of-Experts (MG-MoE) model is proposed. It is a routed ensemble architecture composed of a shared backbone, four heterogeneous experts, and a learnable router that predicts input-conditioned expert weights (Fig. 2). The experts are designed with complementary inductive biases to address key factors in FGIR. MPSA emphasizes global structure and contour-level semantics. PMG captures fine local details through multi-granularity part modeling. TransFG focuses on pose and deformation modeling. PIM improves robustness in cluttered backgrounds through background suppression. To limit interference and reduce unnecessary computation, MG-MoE adopts sparse fusion. Only the Top-K experts, with K=2 by default, contribute to the final prediction during inference. To improve routing stability and generalization, a two-stage optimization strategy is designed. In the first stage, dynamic cluster-level training is performed. A cluster-level soft teacher distribution is constructed from validation-set statistics and imposed through Kullback-Leibler (KL) divergence regularization. This process stabilizes routing behavior and promotes effective expert specialization. In the second stage, residual fine-tuning is conducted. The feature-driven routing mechanism is kept unchanged, while the classification heads of the Top-2 experts associated with each cluster are selectively unfrozen. The router and expert heads are then jointly optimized with grouped learning rates. This design reduces fusion bias and strengthens discrimination for difficult samples and long-tailed categories.  Results and Discussions  MG-MoE achieves strong performance on standard FGIR benchmarks. On CUB-200-2011, it obtains 92.89% Top-1 accuracy. This result is higher than those of representative expert backbones used individually, including MPSA (91.23%), PIM (91.17%), and TransFG (90.49%). It also outperforms the multi-granularity baseline PMG (88.32%) (Table 1). On the Bird-1445 sampled set, MG-MoE achieves 96.80% Top-1 accuracy and consistently improves over strong baselines (Table 2). These results indicate that routed multi-expert specialization remains effective in data-limited and highly similar fine-grained scenarios. The efficiency-accuracy trade-off is summarized in Table 3. With Top-2 sparse routing, MG-MoE reaches 92.89% accuracy with a compute budget of 143.9 GFLOPs. It avoids dense expert activation during inference by selecting only the Top-2 experts for each sample, thereby achieving a favorable balance between accuracy and efficiency. Ablation experiments show that increasing K beyond 2 does not yield consistent gains, which suggests that indiscriminate fusion can dilute discriminative evidence. Top-2 fusion produces the best performance, whereas Top-1 fusion is more sensitive to routing errors and larger K values may introduce noise and reduce accuracy (Table 4). The role of expert diversity and composition is also analyzed. Two- and three-expert variants generally underperform the full four-expert configuration, indicating that each inductive bias contributes to different fine-grained difficulty factors. In contrast, adding homogeneous experts without new functional diversity brings diminishing or negative gains, which is consistent with increased routing ambiguity and limited expert complementarity (Table 5). These results support the use of a compact set of heterogeneous experts combined with sparse routing. To interpret the learned specialization, category-wise routing statistics are visualized. The expert-category heatmap shows that MPSA receives dominant routing weights across many categories, reflecting the central role of global structure in fine-grained discrimination. PIM and TransFG show higher activation for specific difficult categories, which is consistent with their roles in background suppression and pose and deformation modeling (Fig. 3). Finally, t-SNE visualizations illustrate the qualitative effect of expert fusion on class separability. Shared backbone features show stronger inter-class entanglement among visually similar subcategories. In contrast, fused outputs form clearer clusters with better between-class separation and within-class compactness, indicating a more reliable decision space shaped by routed expert aggregation (Fig. 4).  Conclusions  MG-MoE is a multi-granularity routed mixture-of-experts framework for fine-grained recognition. By combining four complementary experts, Top-2 sparse fusion, and a two-stage optimization strategy for stable routing and calibrated fusion, MG-MoE improves recognition accuracy on CUB-200-2011 and the Bird-1445 sampled set. It also provides interpretable evidence of expert specialization (Table 1, Table 2, Fig. 3, Fig. 4). Ablation results confirm that controlled Top-2 fusion and heterogeneous expert design are key to the observed performance gains. Overly dense fusion or homogeneous expert expansion provides limited benefit (Table 4, Table 5).
Rotatable-Antenna-Aided Near-Field Wideband Integrated Sensing and Communication System: Hybrid Beamforming Design
XU Hongbo, MO Minghui, XIN Wei, WANG Shuli, WANG Ji, LI Xingwang, ZHENG Le
Available online  , doi: 10.11999/JEIT260023
Abstract:
  Objective  Near-field wideband Integrated Sensing and Communication (ISAC) systems face two main challenges: pronounced near-field effects and wideband beam splitting. These effects reduce communication throughput and sensing reliability, particularly when fixed-orientation antenna arrays and phase-shifter-based beamforming architectures are used. Because such architectures provide limited spatial adaptability and frequency-independent phase control, the spatial-frequency degrees of freedom available in near-field wideband channels cannot be fully used. To address this issue, a Rotatable-Antenna-assisted near-field wideband ISAC architecture is investigated to improve the system sum rate under sensing constraints.  Methods  A near-field wideband ISAC architecture assisted by Rotatable Antennas (RAs) is proposed. By allowing the antenna boresight direction to be adjusted mechanically or electronically, additional angular degrees of freedom are provided at the element level, which enables more flexible spatial coverage and more accurate energy focusing. A True Time Delay (TTD)-based hybrid beamforming architecture is further adopted to provide frequency-dependent phase shifts and compensate for the frequency-independent property of conventional phase shifters. Consistent beam focusing across subcarriers is thus maintained, and wideband beam splitting is effectively suppressed. Based on a spherical-wave near-field channel model that incorporates propagation distance, angular information, and the orientation gain of RAs, a joint optimization problem is formulated to maximize the system sum rate under transmit power constraints, sensing power thresholds, and antenna rotation constraints. Because the resulting problem is highly non-convex, a Penalty-Based Fully Digital Approximation (PBFDA) algorithm is developed. In each iteration, the RA orientations are first optimized by Particle Swarm Optimization (PSO) to improve the weighted channel gain. Then, with the antenna orientations fixed, a reduced-dimensional formulation with Successive Convex Approximation (SCA) is used to solve the fully digital beamforming problem. Finally, a manifold-based Block Coordinate Descent (BCD) algorithm is used to jointly optimize the analog beamformer, digital beamformer, and TTD units, so that the hybrid beamforming solution gradually approaches the fully digital solution (Algorithm 1–Algorithm 4).  Results and Discussions  Simulation results verify the effectiveness of the proposed RA-assisted near-field wideband ISAC framework. The proposed PBFDA algorithm converges monotonically within a limited number of iterations, which confirms its numerical stability and efficiency (Fig. 2). Compared with fixed-antenna architectures, the proposed RA-assisted scheme achieves a clear improvement in system sum rate under the same transmit power constraint (Fig. 3). When the system bandwidth increases, the spectral efficiency of TTD-based hybrid beamforming decreases because the limited number of TTD units and the restricted maximum delay weaken frequency-dependent compensation and aggravate beam splitting. By contrast, the optimal fully digital beamforming scheme maintains nearly unchanged spectral efficiency because each subcarrier can be controlled accurately (Fig. 4). When the sensing power threshold increases, the achievable sum rate decreases for all schemes, which reflects the trade-off between communication and sensing. The proposed method, however, consistently outperforms the benchmark schemes (Fig. 5). The effects of antenna number, antenna directivity factor, and maximum rotation angle are also evaluated. Spectral efficiency increases with the number of antennas because of the higher array gain (Fig. 6). As the antenna directivity factor increases, the RA-assisted system attains further gains through adaptive orientation, whereas fixed-orientation and isotropic schemes degrade (Fig. 7). A larger allowable rotation range also provides greater spatial alignment flexibility and further improves system performance (Fig. 8). Overall, the proposed architecture improves near-field energy focusing and achieves performance close to that of fully digital beamforming with lower hardware complexity.  Conclusions  A Rotatable-Antenna-assisted near-field wideband ISAC system with a TTD-based fully connected hybrid beamforming architecture is investigated. By jointly using antenna rotation and true time delay, the proposed framework effectively mitigates near-field effects and wideband beam splitting. The developed PBFDA algorithm solves the resulting highly non-convex optimization problem efficiently. Numerical results show that the proposed scheme significantly improves the system sum rate under sensing constraints and approaches the performance of fully digital beamforming, which supports its use in near-field wideband ISAC systems.
A Multimodal Sentiment Analysis Model with Multi-source Knowledge guided Visual Confidence Perception
PENG Juhong, ZHANG Zhi, LIU Peng, GE Wenhui, LIU Chen, LIAO Lingxin, ZHANG Kai
Available online  , doi: 10.11999/JEIT260063
Abstract:
  Objective  Multimodal sentiment analysis is often affected by visual noise from complex environments, image-text sentiment inconsistency, and imbalanced modality contributions. When all modalities are treated without distinction, visual noise can degrade model performance. A robust mechanism is therefore needed to evaluate visual confidence and filter redundant visual information.  Methods  A Multimodal Sentiment Analysis Model with Multi-source Knowledge-guided Visual confidence Perception (MKVP) is proposed (Fig. 1). A multi-source knowledge guidance matrix is constructed using syntactic-dependency, sentiment-intensity, and aspect-focused operators (Fig. 2). Guided by this matrix, the Visual Confidence Perception (VCP) module measures semantic affinity and dynamically suppresses irrelevant visual noise (Fig. 3). A dual-stream parallel interaction module is then used to support deep cross-modal alignment, and a global gated fusion mechanism further adjusts the fusion weights of different modalities.  Results and Discussions  Extensive experiments are conducted on the MVSA-Single, MVSA-Multiple, and HFM datasets. The proposed MKVP model achieves accuracy and F1 scores of 77.56% and 76.70%, 72.72% and 70.66%, and 87.26% and 86.78%, respectively. Compared with the baseline models, the accuracy and F1 score are improved by 2.45% and 3.68%, 2.19% and 2.21%, and 1.83% and 1.91%, respectively (Table 3). Ablation studies show that each component contributes to performance, especially the VCP module, which filters visual noise and improves feature quality (Table 5). Feature-space visualization further confirms that the VCP module refines semantic representations by promoting clearer clustering of samples with the same sentiment polarity (Fig. 4). Case studies on mismatched image-text samples also verify the ability of the model to resolve cross-modal semantic conflicts (Table 6). Model-complexity analysis shows that MKVP maintains high computational efficiency and low inference latency (Table 8).  Conclusions  The proposed MKVP framework reduces the effects of visual noise and image-text sentiment inconsistency in multimodal sentiment analysis. By using multi-source knowledge to guide visual confidence perception and combining dual-stream interaction with dynamic gated fusion, the model learns robust sentiment representations from noisy multimodal data. This method provides an efficient and reliable solution for complex social media scenarios.
Robust Optimization of Low-altitude Communication and Computation Resources in Uncertain Environments
GONG Yucheng, LI Bin, WANG Xinyi, FEI Zesong
Available online  , doi: 10.11999/JEIT260090
Abstract:
  Objective  Low-altitude edge computing networks provide flexible computing services and extended coverage for user equipment. However, quality of service is often degraded by uncertainty in task data size and by Unmanned Aerial Vehicle (UAV) position jitter caused by environmental disturbances. Existing robust methods commonly rely on deterministic uncertainty sets, which tend to be conservative and cannot accurately describe the stochastic distribution of task demands. To address these challenges, a robust energy minimization framework is proposed for multi-UAV-assisted Mobile Edge Computing (MEC) networks. The objective is to minimize the weighted sum of system energy consumption. This is achieved by developing a joint optimization model that coordinates UAV flight trajectories, task splitting decisions, and computation and communication resource allocation. The model explicitly accounts for the dual uncertainties of task data size and UAV trajectory.  Methods  To handle the nonconvexity and strong coupling among optimization variables, the problem is first modeled as a Markov Decision Process (MDP). A comprehensive state space is defined to characterize real-time system dynamics, and a continuous action space is designed for trajectory control and resource management. A Distributionally Robust Optimization Soft Actor-Critic (DRO-SAC) algorithm is then developed to solve the MDP. In this framework, an ambiguity set based on the L1-norm distance is constructed to characterize the distributional uncertainty of the task demand distribution. A maximum-entropy reinforcement learning mechanism is used to learn an optimal policy under the worst-case distribution within the ambiguity set. In this way, UAV trajectories, task splitting, and computation and communication resource allocation are jointly optimized to improve system robustness under dynamic environmental fluctuations.  Results and Discussions  The performance of the proposed DRO-SAC algorithm is evaluated through simulations. DRO-SAC achieves faster convergence and higher rewards than Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO) algorithms (Fig. 3). For energy consumption, the proposed method consistently achieves higher efficiency under different user densities (Fig. 4). The robustness of the system against position errors is also verified, with energy fluctuations kept at a low level (Fig. 5). Dynamic trajectory adjustment further confirms that the proposed method can provide effective user coverage while reducing system energy consumption (Fig. 6).  Conclusions  A DRO-SAC-based joint optimization framework is proposed to address uncertainty in task data size and UAV position jitter in multi-UAV-assisted MEC networks. By constructing an ambiguity set for the task demand distribution and optimizing the worst-case expected objective, the proposed method mitigates the limitations of traditional deterministic models in dynamic environments. Weighted system energy consumption is minimized while latency and safety constraints are satisfied. Simulation results demonstrate that the proposed scheme achieves stable convergence and high energy efficiency, even when communication and computation resources are limited and environmental parameters fluctuate strongly.
A Joint Source-Channel Coding Modulation Scheme for the Transmission of Gaussian Sources
LV Yaping, MA Xiao
Available online  , doi: 10.11999/JEIT251224
Abstract:
  Objective  The Separated Source-Channel Coding (SSCC) scheme has been proven, which will not incur performance loss as long as the source block length goes to infinity. However, the SSCC scheme usually leads to a large buffer and long delay, and may cause error propagation in case of a single symbol error incurred in the communication channel. In order to alleviate these issues, Joint Source Channel Coding (JSCC) schemes have been investigated to transmit Gaussian sources. In this paper, a JSCC modulation scheme for the transmission of Gaussian sources is proposed, and a Gaussian source reconstruction scheme and its reconstruction expression are provided.  Methods  In this paper, the Gaussian source sequence is quantified as a sequence of M-ary symbols by a Lloyd-Max quantizer. For the M-ary quantization symbol sequence, the matching M-ary Fourier Transform Pair (FTP) code is constructed, and the modulation mode adopts the corresponding M-ary Pulse Amplitude Modulation (M-PAM). In particular, the modulated M-ary symbol sequences are transmitted in a block Markov superposition way, which constructs the Block Markov Superposition Transmission FTP (BMST-FTP) code. In addition, in order to obtain the shaping gain, the constellation Geometry Shaping (GS) scheme is also proposed. For the proposed source reconstruction scheme, the system output is the weighted average of the representative elements of the Lloyd-Max quantizer, which replaces the representative elements.  Results and Discussions  The simulations are conducted over M-PAM modulated AWGN channels using GF(3) and GF(5) BMST-FTP codes. For FTP codes employing random mapping, the WER approaches the Union Bound (UB) at high SNR. Similarly, the FTP codes with m repeated transmissions exhibit WER performances that approach the corresponding UBs. Furthermore, the WER performance of BMST-FTP codes with memory m matches UBs in the high SNR region (Fig. 6). For Symbol Error Rate (SER), the GF(3) BMST-FTP code outperforms the GF(5) BMST-FTP code (Fig. 7(a)). For the GF(5) BMST-FTP code, the GS can achieve an SER performance gain of approximately 0.3 dB (Fig. 8(a)). In terms of distortion performance, the GF(3) BMST-FTP code outperforms the BMST-FTP code in the low SNR region, whereas in the high SNR region, the GF(5) BMST-FTP code performs better (Fig. 7(b)). Furthermore, compared with other work, the GF(3) BMST-FTP code with m=1 has a similar performance, and the GF(5) BMST-FTP code with m=1 performs better (Fig. 7(b)).  Conclusions  This work has proposed a joint source-channel coding modulation scheme for the transmission of Gaussian sources. In the proposed scheme, two types of BMST-FTP codes were constructed, each matched with a corresponding Lloyd-Max quantizer and M-PAM modulator. Additionally, a Gaussian source reconstruction scheme and its reconstruction expression were provided. Simulation results demonstrate that the appropriate transmission scheme can be selected according to the aim performance. The proposed GS scheme can obtain a gain of SER of about 0.3dB, which can improve the distortion performance of the waterfall area.
A Lightweight True Random Number Generator Based on Chain-Coupled Oscillation Rings
ZHANG Yuan, YING Haixuan, GAO Kai, YE Jin, WANG Shuang, ZHANG Jiliang
Available online  , doi: 10.11999/JEIT260377
Abstract:
  Objective  With the rapid growth of the Internet of Things, 5G/6G, and satellite Internet, resource-constrained devices increasingly require high-quality random numbers for key generation, authentication, masking, and other security functions. Although pseudo-random number generators are efficient, their outputs may be predictable once the seed or internal state is compromised. True random number generators (TRNGs) offer a hardware root of trust by extracting entropy from physical randomness, but many existing designs rely on multiple entropy sources or complex post-processing, leading to increased area and power consumption. To address this issue, this paper proposes a lightweight TRNG based on chain-coupled oscillation rings for high-quality randomness with very low FPGA overhead.  Methods  Starting from the state evolution of a Galois oscillation ring (GARO), this work demonstrates that ideal matched-delay conditions can result in periodic and predictable oscillation. However, in practical circuits, delay mismatch, jitter, and process variation disturb the ideal evolution and can be exploited as entropy sources. On this basis, a compact delay-feedback XOR ring is proposed to enhance state uncertainty, introduce feedback competition, and improve randomness through inter-stage delay differences. In addition, a second-order oscillation ring is incorporated to eliminate the all-zero stop state and provide continuous excitation. Multiple rings are then chain-coupled, enabling adjacent rings to mutually interfere with one another and thereby generate stronger irregular oscillations. The proposed design is modeled in MATLAB and implemented on a Xilinx Artix-7 FPGA. Finally, we evaluate its performance by NIST SP 800-22, NIST SP 800-90B, bias, autocorrelation, and voltage-temperature robustness tests.  Results and Discussions  Simulation confirms that the proposed structure avoids stable periodic locking and produces sustained irregular oscillation. Experimental results show that the TRNG passes all NIST SP 800-22 tests and achieves an average minimum entropy of 0.9936 in NIST SP 800-90B test, outperforming conventional RO and GARO-based TRNGs under similar conditions. The measured bias is only 0.0228%, and the autocorrelation remains well below the threshold, indicating excellent statistical independence. The design also maintains high entropy over temperatures from 0 °C to 80 °C and supply voltages from 0.9 V to 1.1 V. Implemented on Artix-7, our proposed TRNG achieves 200 Mbps throughput using only 11 LUTs and 4 DFFs, with 0.108 W power consumption.  Conclusions  This paper presents a lightweight chain-coupled oscillation-ring TRNG that exploits delay mismatch, phase disturbance, and feedback competition to generate high-quality physical randomness. The theoretical analysis clarifies how practical nonidealities transform ideal periodic oscillation into irregular oscillation, providing a design basis for compact oscillator-based entropy sources. By combining delay-feedback XOR rings with chain-coupled mutual disturbance and continuous excitation, the proposed design enhances entropy while avoiding excessive hardware overhead and complex post-processing. FPGA implementation and statistical evaluations verify high entropy, low bias, and high randomness under voltage and temperature variations. Therefore, the proposed TRNG achieves high randomness quality and high throughput while effectively reducing hardware overhead, making it suitable for resource-constrained security applications such as IoT terminals, lightweight cryptographic modules, and embedded authentication systems.
From Touch to Semantics: A Cross-Modal Framework for Zero-Shot Spiking Tactile Object Recognition
CHI Wei, XU Jin
Available online  , doi: 10.11999/JEIT260158
Abstract:
  Objective  Tactile perception enables robots to understand object properties and perform dexterous interactions. However, tactile data are costly to collect and difficult to scale, which limits conventional supervised learning in open-world scenarios. Zero-Shot Learning (ZSL) provides a promising solution by transferring knowledge from seen to unseen categories through semantic representations. Existing tactile ZSL methods either rely on auxiliary visual information or use manually designed attributes, which are often subjective and limited in generalization. Event-based spiking tactile signals are sparse and asynchronous, with rich spatiotemporal dynamics. These properties make semantic modeling more challenging. Systematic studies on zero-shot recognition for such data remain limited. To address these issues, this paper proposes a zero-shot object recognition framework for spiking tactile perception. The framework aims to bridge low-level tactile dynamics and high-level semantics in a scalable manner.  Methods  The proposed framework consists of three components (Fig. 1): spiking tactile feature extraction, semantic prototype construction, and cross-modal tactile-semantic alignment. First, a biomimetic Spiking Graph Neural Network (SGNN) is used to model raw event-based spiking tactile signals. By integrating Leaky Integrate-and-Fire (LIF) neurons with graph-based message passing, the SGNN captures temporal firing dynamics and spatial relationships among tactile sensing units. It then generates discriminative and biologically interpretable high-level tactile embeddings. Second, instead of using manually annotated attributes, a Large Language Model (LLM) is used to generate structured, fine-grained, and extensible tactile attribute descriptions for each object category. These textual descriptions are encoded as continuous semantic vectors to form class-level semantic prototypes with consistent dimensionality across categories. This strategy supports flexible semantic expansion and avoids labor-intensive attribute engineering. Third, a bidirectional tactile-semantic alignment mechanism is designed to improve generalization to unseen categories. A forward mapping projects tactile embeddings into the semantic space for classification, whereas a reverse mapping reconstructs tactile features from semantic representations. A cycle-consistency constraint is imposed between the two mappings to preserve structural coherence and semantic stability across modalities. The overall framework is trained only on seen categories. During zero-shot inference, tactile embeddings of unseen samples are matched with their corresponding semantic prototypes in the shared embedding space.  Results and Discussions  The proposed method is evaluated on the Ev-Object event-based tactile dataset under a strict zero-shot setting, with disjoint seen and unseen category sets. Performance is assessed using Mean Class Accuracy (MCA), Top-k accuracy, and the Semantic Alignment Score (SAS). The proposed framework consistently outperforms representative tactile ZSL baselines across all metrics. It achieves an MCA of 73.48%, a Top-1 accuracy of 62.68%, and a Top-2 accuracy of 88.75%. Ablation studies show that removing the LLM semantic module, bidirectional mapping, or cycle-consistency constraint reduces recognition performance and semantic alignment quality. Removing the LLM semantic module causes a substantial decrease in MCA, which confirms the role of structured LLM-generated tactile semantics in knowledge transfer. Removing the bidirectional mapping or the cycle-consistency constraint also reduces performance, indicating that both components help maintain stable cross-modal alignment. The t-SNE visualization further shows that cycle-consistent alignment yields more compact intra-class clusters and clearer inter-class separation for unseen categories. Semantic prototypes are also better located near the centers of tactile feature clusters. These results indicate that combining biologically inspired spiking models with LLM-generated tactile semantics provides an effective solution for open-world tactile perception.  Conclusions  This paper presents a zero-shot object recognition framework for spiking tactile perception by integrating SGNN-based tactile representation with semantic prototypes. The proposed method addresses key limitations of existing tactile ZSL approaches by avoiding visual data and manual attribute design while effectively modeling the spatiotemporal dynamics of event-based spiking tactile signals. Experimental results under strict zero-shot settings confirm the effectiveness and robustness of the proposed framework. This work provides a strong baseline for zero-shot spiking tactile recognition and offers a principled path toward open-world tactile cognition in robotic systems. Future work will explore generalized zero-shot tactile perception, multimodal extensions, and real-world robotic deployment under noisy and dynamic sensing conditions.
Design of a Timing-Controlled Non-Volatile Flip-Flop with Low-Switching-Ratio FeFET
DU Shimin, YANG Chang, WANG Lunyao, ZHANG Zhe
Available online  , doi: 10.11999/JEIT251059
Abstract:
  Objective  Nonvolatile processors(NVPs) have become a key technology for Internet-of-Things (IoT) and energy-harvesting systems, where maintaining computational states during unexpected power interruptions is essential. Conventional volatile processors rely on external nonvolatile memory(NVM) for state retention; however, this approach incurs significant latency and energy overheads. Integrated nonvolatile flip-flops using ferroelectric field-effect transistors(FeFETs) offer a promising alternative by enabling on-chip state backup and recovery. Nevertheless, existing single-ended FeFET-based flip-flops are prone to contention-induced failures during power-up recovery, especially when the FeFET on/off ratio degrades. This issue originates from competing discharge paths that lead to uncertainty in internal node voltage settling, thereby resulting in unreliable state restoration. To address this challenge, this work proposes a novel flip-flop architecture that replaces contention-based recovery with a timing-controlled two-phase mechanism. The primary objectives of this design are to achieve high-reliability recovery even under degraded FeFET on/off ratios as low as 102, optimize timing parameters such as Hold-Time and Clock-to-Q delay, and maintain low energy consumption suitable for IoT applications.  Methods  The proposed design is an extension of the Static Contention-Free Single-Phase-Clocked Flip-Flop(SSCFF), which inherently eliminates internal node contention through its fully static structure. Based on this foundation, one FeFET device and five additional MOSFETs are integrated to construct a single-ended nonvolatile flip-flop(NVFF). Two control signals, RES and MOD, are introduced to manage the recovery process.In the normal operation mode where MOD=0, the circuit functions as a conventional SSCFF and supports state backup during runtime. In the recovery mode where MOD=1, the recovery operation is divided into two distinct phases.In the pre-charge phase, when RES=0, the internal nodes are pre-charged to VDD. In the selective discharge phase, as RES transitions from low to high, the resistance state of the FeFET determines whether discharge occurs. If the FeFET is in the low-resistance state(LRS), a discharge path is formed, pulling the node voltage down to ground. If the FeFET remains in the high-resistance state(HRS), the node retains its charge until the next clock edge.This sequence of pre-charging followed by selective discharge eliminates contention during recovery and ensures that the internal node voltages settle deterministically and reliably.The design is implemented in a 130nm CMOS process with integrated FeFET models. Simulations, including Monte Carlo analysis, were performed in Cadence Virtuoso across a supply voltage range of 0.6–0.9 V and FeFET on/off ratios ranging from 102 to 104. Key performance metrics, such as Setup-Time, Hold-Time, Clock-to-Q delay, restore energy, and recovery success rate, were evaluated and compared against traditional Transmission Gate Flip-Flop.  Results and Discussions  Simulation results show that the timing-controlled recovery improves reliability even under severe FeFET degradation. The proposed flip-flop achieves 100 % restore yield when the FeFET on/off ratio drops to 102. This is because the proposed structure eliminates the competing discharge paths. Timing metrics are also improved: the 3σ worst-case Hold-Time is reduced by 64.6 % , and the Clock-to-Q delay is shortened by 33.9 %. Although Setup-Time increases slightly, it can be compensated by device sizing. Restore energy remains in the low-fJ (10–15 J) range across all supply voltages, rising only modestly compared with the TGFF because of the added pre-charge phase.  Conclusions  A Ferroelectric FET Nonvolatile Flip-Flop with timing-controlled two-phase recovery has been presented, addressing the contention-induced failure modes that limit low-voltage NVFF reliability. By integrating a single FeFET with an enhanced SSCFF structure and using RES signal to manage the pre-charge and discharge steps, high restore yield is maintained even under severely degraded FeFET on/off ratios, while Hold-Time and Clock-to-Q delay are significantly improved relative to traditional transmission-gate NVFFs. The proposed architecture offers a compelling solution for energy-constrained IoT processors requiring fast, reliable state preservation under unpredictable power conditions.
A Noise Reduction Strategy via Coprime-Spacing Subarrays for Biodiversity Acoustic Indices
CHEN Lei, XU Zhiyong, ZHAO Zhao
Available online  , doi: 10.11999/JEIT260237
Abstract:
  Objective  As a popular tool for rapid biodiversity assessment, acoustic indices have attracted increasing attention in the field of soundscape ecology in recent years. Nevertheless, most commonly used acoustic indices are susceptible to background noise. Traditional single-channel noise reduction strategies, including spectral subtraction, high-pass filtering, and threshold detection, have been widely adopted as preprocessing approaches to optimize the calculation of acoustic indices. However, when dealing with anthropogenic interference that overlaps with biotic signals in both time and frequency domains, the denoising capability of single-channel methods degrades severely. Although spatio-temporal adaptive whitening filtering based on microphone arrays provides a feasible approach for suppressing directional interference, it suffers from a non-uniform two-dimensional spatio-temporal amplitude response and the self-cancellation of target signal in the unconstrained interference cancellation. These disadvantages lead to distortion in the time-frequency distribution of target signals, causing acoustic index calculations to deviate from the ground truth. Therefore, this study aims to propose a noise reduction strategy via coprime-spacing subarrays for biodiversity acoustic indices. This method effectively suppresses directional interference while maximally preserving the time-frequency distribution structure of biotic signals.  Methods  The noise reduction strategy based on microphone array spatio-temporal adaptive whitening filtering is proposed, incorporating the Frequency-dependent Acoustic Diversity Index (FADI), which is insensitive to fluctuations in the array's two-dimensional spatio-temporal amplitude response. A noise-robust acoustic index method, termed Adaptive Interference Cancellation–Frequency-dependent Acoustic Diversity Index (AIC-FADI), is subsequently developed. Specifically, a non-uniform linear array is first constructed using three microphones to form two dual-element subarrays with coprime spacing. This design fully exploits the high spatial resolution of wide-spacing arrays to narrow the null width in the direction of interference. Meanwhile, it avoids the physical implementation difficulties and mutual coupling effects associated with small-spacing array designs caused by the ultra-wideband characteristics of target signals. The spatio-temporal adaptive whitening filtering is then performed on each coprime-spacing subarray separately, adaptively forming two-dimensional nulls within the interference support region, thereby suppressing directional anthropogenic interference in analytical data before index calculation. Next, a frequency-dependent threshold scheme is utilized to obtain the binary spectrogram for each coprime-spacing subarray output, abating the influence from gain differences along the frequency axis for a certain direction. Afterwards, by leveraging the high spatial resolution of wide-spacing arrays and the interleaved characteristics of spatial aliasing null positions between the spatio-temporal frequency responses of the two subarrays with coprime spacing, a pointwise maximum fusion is applied to the above two binary spectrograms. This process reconstructs the binary time-frequency distribution structure of target signals outside the interference support region, leading to a single binary spectrogram where biological sound components are preserved to a great extent and anthropogenic interference is considerably suppressed. Ultimately, from this single binary spectrogram, the proportions of non-zero time-frequency bins within each frequency band are calculated and forwarded to the entropy function, resulting in the final AIC-FADI result.  Results and Discussions  The simulation result indicates that the proposed AIC-FADI maintains numerical robustness across an SINR range down to –15 dB (the yellow line in Fig. 5), substantially outperforming the classical ADI version based on single-channel noise reduction algorithm (FADI) and other ADI versions based on single-array interference suppression processing mentioned in this paper (AIC-FADI-s, AIC-FADI1, and AIC-FADI2). The real-world experiment confirms that the proposed spatio-temporal adaptive whitening filtering effectively suppresses wideband interference signals in complex scenarios, thereby improving the SINR of the analyzed recording. This enables some weaker biotic signals to exceed their corresponding frequency-dependent adaptive thresholds, greatly reducing missed detection of the target signal. In addition, by performing pointwise maximum fusion of the binary spectrograms from the two coprime-spacing subarray outputs, AIC-FADI further alleviates the extent of target signal missed detection (Fig. 8). Nevertheless, the real-world experiments also reveal that the interference suppression performance of AIC-FADI degrades for highly time-varying interference components.  Conclusions  This paper addresses the challenge of calculating acoustic indices reliably in complex soundscapes where directional anthropogenic interference overlaps with biotic signals in both time and frequency domains. A noise reduction strategy using coprime-spacing subarrays is proposed, and a new noise-robust acoustic index (AIC-FADI) is then developed. The method is evaluated through simulations and real-world recordings, and the results show that: (1) By applying spatio-temporal adaptive whitening filtering on each coprime-spacing subarray followed by pointwise maximum fusion, the proposed method achieves both wideband interference suppression capability and target information fidelity in complex soundscapes containing strong interference. (2) As a result, the proposed AIC-FADI maintains numerical robustness down to –15 dB SINR, substantially outperforming the classical FADI algorithm and other ADI versions based on single-array interference suppression methods. (3) The proposed method provides a feasible technical solution for extending the practical application scenarios and spatio-temporal coverage of biodiversity acoustic indices in human-dominated areas. However, this study only considers directional interference that is relatively stable or slowly time-varying. Hence, the interference suppression performance degrades for highly time-varying or uncorrelated noise components. These challenges should be addressed in future work through more advanced signal processing techniques to further improve the robustness of acoustic indices in highly complex acoustic environments.
A Survey of Quantum Covert Communication Integration Schemes and Application Scenarios
SUN Yiheng, XU Yongjun, ZHANG Haibo, HUANG Zishan
Available online  , doi: 10.11999/JEIT260282
Abstract:
  Significance   With the growing demand for network communication security, research and development in covert communication and quantum communication have continued to evolve. However, current covert communication suffers from inherent security vulnerabilities; the transmission reliability of quantum communication has been limited by information eavesdropping and harmful interference. Therefore, quantum covert communication has become a research hotspot, integrating the advantages of both covert and quantum communication while addressing their respective security limitations. To this end, this paper provides a comprehensive survey of quantum covert communication integration schemes and application scenarios, including the principles of covert communication and typical enabling techniques; protocols for quantum communication and important quantum techniques; and three types of quantum covert communication integration schemes summarized by different application scenarios. This paper contributes to the design of advanced secure communication networks while offering guidance for the development of future quantum covert communication systems.  Progress   This paper presents a comprehensive survey of recent advances in quantum covert communication integration schemes and application scenarios, with an in-depth discussion of the principles of covert communication and key enabling techniques, such as Fluid Antenna (FA), Reconfigurable Intelligent Surface (RIS), and Unmanned Aerial Vehicle (UAV). FA actively reshapes wireless channel characteristics, particularly the spatial correlation of multipath components, by dynamically adjusting the transmitter physical configuration, thereby reducing information leakage. In Non-Line-of-Sight (NLoS) scenarios, RIS can dynamically alter the direction of reflected transmission of the incident signal, not only enhancing the Channel State Information (CSI) quality of the covert signal but also reducing signal leakage. In flexible or temporary communication networks, UAVs can increase CSI uncertainty, preventing unauthorized users from establishing a stable monitoring model and thereby complicating eavesdropping. Then, key protocols and significant techniques of quantum communication are introduced, including BB84, B92, and E91 for Quantum Key Distribution (QKD), and BF02, Two-Step for Quantum Secure Direct Communication (QSDC). Additionally, the quantum repeaters and Quantum Random Number Generator (QRNG) are reviewed. Based on different application scenarios, quantum covert communication integration schemes can be categorized into enabling, covert, and symbiotic integration schemes, depending on the integration mechanisms. To be specific, the enabling integration scheme leverages the unconditional security of quantum communication to address the security vulnerabilities in covert communication, the covert integration scheme utilizes enabling techniques in covert communication to reduce the detection probability of quantum communication, and the symbiotic integration scheme combines both advantages of covert communication and quantum communication to achieve mutual empowerment and deep symbiosis. Finally, critical challenges are highlighted, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles in large-scale applications. Promising directions for future research are also identified, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development.  Prospects   Despite remarkable progress in preliminary applications and specific scenarios, research on quantum covert communication remains in its infancy. As quantum covert communication scenarios become increasingly diverse and complex, future studies should prioritize challenges that restrict further development and large-scale application of quantum covert communication. The stringent hardware precision requirements are the primary challenge, limiting reliable transmission distance and stability. Low resource allocation efficiency is another challenge, as the quantum covert communication system that generates quantum entanglement over lossy channels remains subject to the Square Root Law (SRL) constraints, while signal transmission exhibits burstiness and dynamics. Additionally, high deployment costs and the lack of standardization present significant hurdles. To address the challenges mentioned, future directions should include R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development to facilitate the development of high-performance, large-scale, and multi-scenario quantum covert communication.  Conclusions  This paper provides a comprehensive survey of quantum covert communication with particular emphasis on integration schemes and application scenarios. The fundamentals and typical enabling techniques of covert communication are first reviewed, highlighting its Low Probability of Detection (LPD) secure paradigm and unique channel characteristics. The typical protocols and important techniques of quantum communication are then examined, including QKD, QSDC, quantum repeaters, and QRNG. Three types of quantum covert communication integration schemes have been further classified by different integration mechanisms and corresponding application scenarios. Finally, several existing challenges are identified, including stringent hardware precision requirements, low resource allocation efficiency, and obstacles to large-scale applications. Relevant research directions are also outlined, including R&D on precision communication equipment, dynamic resource management, cost control during deployment, and the promotion of standardized development. These directions are expected to serve as a valuable reference for advancing and standardizing quantum covert communication in future secure networks.
Full-Space Covert Integrated Sensing and Communications Assisted by Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface
XIE Wenwu, ZHANG Qinke, YANG Liang, WANG Ji, YU Chao, LIU Xinzhong, CUI Yaru
Available online  , doi: 10.11999/JEIT260145
Abstract:
  Objective  The evolution of Sixth Generation (6G) mobile communications toward higher frequencies and larger antenna arrays has made Integrated Sensing And Communication (ISAC) a key enabling technology. However, ISAC systems still face limited communication covertness and resource competition between sensing and communication. Covert communication and Reconfigurable Intelligent Surface (RIS) techniques provide promising solutions. However, most existing studies use reflective RISs with half-space coverage and assume far-field propagation. These assumptions limit deployment flexibility and fail to capture near-field spherical-wave characteristics. To address these issues, this paper proposes a near-field full-space ISAC framework assisted by an Extremely Large-Scale Simultaneously Transmitting And Reflecting Reconfigurable Intelligent Surface (XL-STAR-RIS). The objective is to jointly optimize active transmit beamforming and passive XL-STAR-RIS coefficient design to improve the covert communication rate while satisfying sensing performance and covertness requirements.  Methods  The detection capability of warden Willie is first analyzed, and a closed-form lower-bound expression for the minimum Detection Error Probability (DEP) is derived. A non-convex optimization problem is then formulated to maximize the covert communication rate under sensing Signal-to-Noise Ratio (SNR), covertness, and total transmit power constraints. Direct solution is difficult because the active transmit beamforming vectors and passive XL-STAR-RIS coefficients are strongly coupled. An Alternating Optimization (AO) framework is therefore adopted to decompose the original problem into two tractable subproblems. The active transmit beamforming subproblem is solved using SemiDefinite Relaxation (SDR) combined with a penalty-based successive convex approximation method. The passive XL-STAR-RIS coefficient design subproblem is solved using the Dinkelbach algorithm and a rank-one penalty method. The two subproblems are solved alternately until convergence.  Results and Discussions  Simulation results verify the effectiveness of the proposed framework. The algorithm converges within approximately 10 iterations and achieves a covert communication rate of about 11.5 bit/(s·Hz). This rate is higher than those of the passive-RIS scheme (9.8 bit/(s·Hz)) and the non-RIS scheme (8.0 bit/(s·Hz)). The performance gain becomes more evident as the transmit power increases, which indicates strong power adaptability. The proposed framework also maintains robust performance under strict operational constraints. When the sensing SNR threshold increases, it achieves a higher covert communication rate than the benchmark schemes. Under a stricter covertness requirement, it also preserves a higher communication rate. These results show that joint active transmit beamforming and passive XL-STAR-RIS coefficient design can effectively balance communication, sensing, and covertness in near-field ISAC systems.  Conclusions  This paper presents an XL-STAR-RIS-assisted covert communication framework for near-field ISAC systems. By jointly designing active transmit beamforming and passive XL-STAR-RIS coefficients through an efficient AO algorithm, the proposed framework balances communication rate, sensing performance, and communication covertness. Simulation results confirm its advantages over conventional passive-RIS and non-RIS schemes, especially under strict sensing and covertness constraints. The results also indicate the potential of XL-STAR-RIS for secure full-space 6G applications. Future work will consider imperfect Channel State Information (CSI), dynamic propagation environments, and multi-RIS collaboration to improve practical robustness.
Millimeter-Wave Air-to-Ground Channel Prediction Assisted by Visual Information of the Propagation Environment
CHENG Yuanxun, HU Qingsong, ZHANG Xiaomin, WANG Xuesong
Available online  , doi: 10.11999/JEIT260274
Abstract:
  Objective  Accurate prediction of air-to-ground (A2G) channel states is essential for adaptive transmission and resource optimization in unmanned aerial vehicle (UAV) communications. In urban millimeter-wave scenarios, however, A2G links are highly sensitive to blockage, reflection, scattering, and the rapidly changing geometric relationship among the transmitter, the receiver, and surrounding buildings. As a result, the channel exhibits strong spatial and temporal nonstationarity, and conventional pilot- or feedback-based acquisition methods may become ineffective because the obtained channel state information is easily outdated. Recent data-driven approaches have shown potential, but many of them rely heavily on historical channel observations or directly use raw images as network inputs, which may introduce redundant visual information and weaken physical interpretability. To address these limitations, this paper proposes a vision-assisted millimeter-wave A2G channel prediction method that extracts low-dimensional geometric features from the propagation environment instead of using raw visual data directly. The objective is to preserve the key structural information governing channel evolution while reducing irrelevant redundancy, thereby improving the prediction of channel.  Methods  A communication-and-sensing integrated dataset with strict spatial and temporal alignment is established for millimeter-wave UAV A2G channel prediction. On the sensing side, a high-fidelity three-dimensional urban scenario containing 23 buildings, roads, and intersections is constructed in Unreal Engine 4.27, where synchronized RGB and depth images are collected through AirSim using a multirotor UAV equipped with RGB and depth cameras. The UAV flies along 10 preset trajectories at a height of 55 m with a spatial sampling interval of 1 m, yielding 2160 valid visual samples (Fig. 1, Fig. 2). On the communication side, the same scene is reconstructed in Wireless InSite, and the transmitter-receiver positions are synchronously updated along the same trajectories to ensure frame-level alignment between visual and channel data (Fig. 3). To obtain compact and physically meaningful environmental representations, a cross-modal spatial feature extraction scheme is developed. Buildings are first detected from RGB images using YOLO-V8 (Fig. 4), and the detected regions are then registered with depth images to reconstruct three-dimensional point clouds. After Euclidean clustering and axis-aligned bounding-box fitting, key geometric attributes, including planar position, height, and volume, are extracted. These features are combined with the transmitter-receiver distance to form the spatial feature vector of each frame, and their relevance to path loss, received power, and RMS delay spread is evaluated through cosine-similarity-based correlation analysis (Fig. 6). Based on the extracted features, a hybrid Transformer-MLP network is designed for channel prediction (Fig. 5). Building features are first projected into a latent space, and a stacked Transformer encoder is employed to capture global interactions among buildings through masked multi-head self-attention. Masked average pooling is then used to aggregate building-level representations into a scene-level environmental descriptor, which is concatenated with the link distance feature and fed into a multilayer perceptron regressor to predict the three target channel parameters.  Results and Discussions  The results confirm the effectiveness of the proposed spatial feature representation. Correlation analysis shows that the extracted geometric features are consistently related to path loss, received power, and RMS delay spread under different aggregation strategies (Fig. 6), indicating that compact building descriptors can effectively characterize the propagation environment. Among them, building height exhibits the strongest correlation with all three channel parameters, highlighting its important role in blockage, attenuation, and multipath propagation in urban millimeter-wave A2G channels. In prediction experiments, the proposed method accurately tracks the variation trends of all three targets. It remains effective in deep-fading and sharp-fluctuation regions for path loss prediction (Fig. 7), achieves high consistency with the ground truth for RMS delay spread (Fig. 8), and follows rapid local fluctuations of received power with good fidelity (Fig. 9). In contrast, the benchmark model only captures the general trend and shows larger deviations in peaks, valleys, and abrupt-changing intervals. Residual analysis further demonstrates the superiority of the proposed method. Its errors are more concentrated around zero and fluctuate within narrower ranges than those of the benchmark model across all three tasks (Fig. 10). Quantitatively, both the mean absolute error and the root mean squared error are reduced (Fig. 11). In addition, the model maintains acceptable complexity, with about 5.5 M parameters and a single-frame inference delay of about 3.4 ms, indicating good potential for real-time deployment.  Conclusions  A vision-assisted millimeter-wave A2G channel prediction method for UAV communications is proposed. By constructing a strictly aligned communication-and-sensing dataset and extracting low-dimensional spatial features with clear physical meaning, the method establishes an effective mapping from environmental geometry to channel parameters. The proposed Transformer-MLP framework achieves accurate prediction of path loss, received power, and RMS delay spread, while offering better interpretability, robustness, and efficiency than the benchmark model.
Transfer Learning Aided CNN for Efficient Data Detection in ReRAM with Sneak-Path Interference
DAI Bin, WU Anni
Available online  , doi: 10.11999/JEIT260354
Abstract:
  Objective  Sneak path interference (SPI) in resistive random-access memory (ReRAM) introduces unpredictable inter-cell correlations, significantly increasing the complexity of signal detection. Traditional detection methods typically rely on assumptions about known channel noise states, resulting in limited generalization capability in practical applications. To address this issue, three data detection methods based on convolutional neural networks (CNNs) are proposed, which can effectively model and mitigate interference without relying on prior channel information: first, a method combining constrained coding with a multi-layer CNN, which uses constrained coding to determine the sneak path interference state and recover data; second, a dual-CNN framework that first employs a lightweight CNN for sneak path interference identification, followed by a multi-layer CNN for refined detection; third, an approach incorporating transfer learning, which maintains detection accuracy while reducing the required training sample size to one-thousandth of that of traditional methods. Simulation results demonstrate that the proposed method achieves superior bit error rate (BER) performance under unknown channel conditions, with a BER reduction of at least half relative to existing algorithms, approaching the theoretical performance limit. Moreover, the integration of transfer learning reduces the required training samples from \begin{document}$ {10}^{6} $\end{document} to \begin{document}$ 1000 $\end{document}, corresponding to a reduction of three orders of magnitude.  Methods  To address distinct challenges in sneak path interference detection, this paper proposes three methods sequentially:1. The integrated constrained coding aided convolutional neural network (CC-CNN) detection framework effectively addresses the complex inter-cell correlations introduced by sneak path interference. This approach first employs constrained coding to detect the presence of interference and subsequently utilizes a CNN to learn and capture the random correlations under the influence of interference, thereby achieving accurate signal recovery.2. The dual-CNN-based detection method resolves the code rate loss associated with traditional constrained coding. By directly leveraging a CNN to learn and identify sneak path interference patterns from raw data, this method eliminates the need for redundant coding or additional overhead. It ensures high-precision interference detection while preserving the overall code rate performance of the system.3. The transfer learning-based CNN (TL-CNN) detection method overcomes the dependence of high-performance CNNs on large-scale training datasets. By reusing knowledge from pre-trained models, this method enables rapid adaptation to ReRAM signal detection tasks. It significantly reduces the required number of training samples while maintaining high detection accuracy and resource efficiency, thereby enhancing the feasibility of the solution in practical scenarios.  Results and Discussions  Simulation results demonstrate that the performance of the three proposed methods consistently approaches the theoretical lower bound (Fig.6), outperforming baseline methods such as the Belief Propagation (BP) detector, Deep Neural Network (DNN) detector, and Elementary Signal Estimator (ESE) detector. The two-step network achieves performance comparable to that of the single-step network while successfully avoiding code rate loss. Notably, the transfer learning-aided CNN attains near-optimal BER with only 1000 target domain samples, and its performance stabilizes when the sample size exceeds 1000 (Fig.7), fully validating its data efficiency. The integration of SK modules enables the models to effectively capture SPI-induced spatial correlations, while the transfer learning strategy ensures the models’ robust performance under different noise conditions.  Conclusions  The crossbar array architecture of ReRAM is susceptible to sneak-path interference during storage operations, leading to reduced data reliability. To address this issue, this paper proposes three deep learning-based detection methods. Type-I integrates constrained coding with a CNN to achieve efficient and fast interference detection. Type-II adopts a two-stage processing approach: it first classifies interference patterns in the memory array and then performs detection specifically on affected units, thereby ensuring high detection accuracy while minimizing coding rate loss. Type-III introduces a transfer learning framework that leverages a pre-trained model from the source domain, significantly reducing the number of training samples required in the target domain and effectively lowering training overhead. Experimental results show that under different noise conditions, all three proposed methods achieve performance close to the theoretical lower bound, providing an effective solution for enhancing the reliability of ReRAM storage systems.
Physical-layer Security in Visible Light Communications: Fundamental Theories, Key Techniques, and Future Challenges
WANG Jinyuan, YAN Xinrun, LIN Zihan, LI Yuanyuan, LI Zheng, ZHANG Xin
Available online  , doi: 10.11999/JEIT260338
Abstract:
  Significance   Due to the broadcast nature of optical signals, information security represents a critical research direction in visible light communication (VLC). Conventional encryption techniques address network security issues at the upper layers of the protocol stack through access control, cryptographic protection, and end-to-end encryption. However, their security relies on the assumption that eavesdroppers possess limited computational capabilities, an assumption that currently faces significant challenges. In recent years, physical layer security (PLS) has emerged as a novel information security paradigm and has attracted considerable attention from researchers worldwide. PLS exploits the randomness, heterogeneity, and distinctiveness between the main channel and the eavesdropping channel to achieve secure information transmission at the physical layer. To date, extensive research achievements have been made regarding PLS techniques in conventional radio frequency wireless communications (RFWC). Nevertheless, due to substantial differences in frequency bands, transmitted signals, power representations, and channel characteristics, PLS research results from RFWC systems cannot be directly applied to VLC. Although scholars worldwide have conducted research on VLC PLS technology, the foundational theories, key techniques, and future challenges involved in VLC PLS still lack a systematic review. To bridge this gap, this paper presents a comprehensive survey of VLC PLS technology.  Progress   To evaluate and enhance system performance, a classic VLC PLS system model—comprising the received signal model, the input constraint model, and the channel gain model—is initially established. A comprehensive theoretical framework for performance evaluation is then developed, encompassing instantaneous performance metrics, statistical performance metrics, and asymptotic performance metrics. Specifically, to characterize instantaneous performance, existing works on instantaneous secrecy capacity and instantaneous secrecy rate across different scenarios are summarized. As statistical performance metrics, average secrecy capacity, average secrecy rate, secrecy outage probability, probability of strictly positive secrecy capacity, and interception probability are analyzed. To demonstrate asymptotic performance, secrecy diversity order and secrecy degrees of freedom are derived. Furthermore, to enhance the PLS performance, advanced technologies, including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity, are summarized.  Prospects   Despite existing research achievements, numerous challenges remain in VLC PLS. This paper identifies four critical challenges: (i) Accurate PLS performance limit: Deriving exact expression of secrecy capacity under VLC's unique physical constraints remains challenging. (ii) Incomplete evaluation framework: Some key metrics widely used in RFWC have not been investigated in VLC, and the construction of a comprehensive VLC PLS performance evaluation framework remains unresolved. (iii) Limitations of existing methods: Conventional PLS performance enhancement methods typically adopt a “modeling-optimization-verification” separated research paradigm, often falling into a vicious cycle of “inaccurate modeling-suboptimal solutions-limited performance gains”. Therefore, it is imperative to integrate novel technologies (such as deep learning, reinforcement learning, and digital twins) to construct a data-model dual-driven framework for VLC PLS performance enhancement. (iv) Hardware platform gap: The absence of dedicated hardware platforms featuring adversarial topologies and real-time processing capabilities significantly impedes the practical deployment of VLC PLS technologies. Therefore, addressing these challenges is essential for transitioning VLC PLS from theoretical advances to commercial applications.  Conclusions  The broadcast nature of optical signals renders VLC systems vulnerable to eavesdropping attacks. This paper presents a comprehensive survey of PLS in VLC, covering system models, performance metrics (instantaneous, statistical, and asymptotic), and key performance enhancement technologies including secure beamforming, artificial noise, physical region protection, secure coding, and secure diversity. Despite significant progress, challenges remain in establishing accurate performance bounds, complete evaluation frameworks, novel enhancement techniques, and practical hardware implementations. By exploiting channel disparities at the physical layer without relying on complex encryption, PLS represents a paradigm shift in security assurance, paving the way for next-generation secure and reliable VLC networks.
Aerial Spatio-Temporal Image Generation via Latent Diffusion Models
SHANG Yuying, HOU Yingyan, LIU Zinan, LU Wanxuan, HUANG Yuhong, WANG Yixiao, YU Hongfeng, FU Kun
Available online  , doi: 10.11999/JEIT260165
Abstract:
  Objective  Aerial Earth observation plays a pivotal role in environmental monitoring, disaster warning, and urban planning. However, constraints such as flight-platform endurance and mission-window timeliness often prevent acquired aerial imagery from fully characterizing the long-term evolution of the Earth's surface. Although pre-trained latent diffusion models have shown strong potential for image generation, their application in aerial scenarios remains challenging because of the scarcity of high-quality temporal annotation data and semantic-visual misalignment caused by variable observation scales. To address these challenges, this paper proposes ASTIG, a training-free framework for Aerial Spatio-Temporal Image Generation. By leveraging the generative priors of pre-trained latent diffusion models and Large Language Models (LLMs), ASTIG provides a new paradigm for semantically controllable aerial spatio-temporal image generation.  Methods  ASTIG consists of three coordinated components. First, a dynamic semantic decomposition process is proposed to parse complex descriptions of aerial scene evolution into frame-level visual prompts, thereby compensating for the lack of temporal semantic annotations in existing aerial image-text datasets. Second, a Linguistic Binding (LB) strategy is proposed to establish explicit associations between key ground objects and their corresponding visual attributes within the cross-attention mechanism of the diffusion model, thereby improving the semantic response precision of the generated images. Third, a Temporal Anchor Attention (TAA) mechanism is incorporated. It uses dual reference frames to maintain subject stability and background consistency across the generated spatio-temporal image sequence, thus suppressing inter-frame temporal drift under training-free conditions.  Results and Discussions  ASTIG and the baseline methods are evaluated on 7,236 high-quality aerial spatio-temporal descriptions using six automated metrics, including subject consistency, background consistency, temporal flickering, motion smoothness, aesthetic quality, and imaging quality. Quantitative results (Tables 1 and 2) show that ASTIG outperforms the baseline methods in spatio-temporal image generation, with improvements of 3.91% in subject consistency and 4.57% in motion smoothness over the frame-prompt baseline. Qualitative comparisons (Fig. 4) further show its strong ability to model long-term surface evolution in aerial imagery. Ablation studies validate the individual effectiveness of the LB strategy and the TAA mechanism (Table 3 and Fig. 5). Sensitivity analyses of the intervention steps (Table 4 and Fig. 6) and binding strength (Table 5 and Fig. 7) further identify suitable parameter settings. Extension experiments from satellite perspectives (Figs. 8 and 9) also show that ASTIG has the potential to generalize beyond aerial platforms to broader Earth observation scenarios.  Conclusions  This paper proposes ASTIG, a training-free framework for aerial spatio-temporal image generation that addresses the scarcity of high-quality long-term temporal data and semantic-visual misalignment. By leveraging the generative priors of pre-trained latent diffusion models and LLMs, ASTIG integrates a dynamic semantic decomposition process, an LB strategy, and a TAA mechanism to improve temporal semantic construction, semantic response precision, and inter-frame consistency. Experimental results show that ASTIG outperforms existing baseline methods across multiple automated evaluation metrics, providing a new paradigm for aerial spatio-temporal image generation. As a training-free method, ASTIG is still limited by the prior knowledge of the backbone model. Future work will examine geometric correction and nadir-view prior constraints to better align the generated results with the physical properties of satellite imagery.
Joint Channel Estimation and Diagnosis for Blocked RIS-Assisted Multi-User Multipath Millimeter-Wave Systems
LI Shuangzhi, LIU Cong, WANG Ning, HAN Gangtao, GUO Xin
Available online  , doi: 10.11999/JEIT260093
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS) can effectively modulate Millimeter-Wave (mmWave) signals and reshape the wireless propagation environment. In practical deployments, however, RIS elements are vulnerable to adverse weather and physical obstructions, which cause unpredictable distortion and motivate joint channel estimation and blockage diagnosis. Most existing studies focus on single-user systems, whereas multi-user scenarios remain insufficiently studied. This gap creates an opportunity to exploit the common RIS blockage vector and the shared RIS-Base Station (BS) channel across users. This paper therefore proposes a low-complexity framework for joint channel estimation and blockage diagnosis by exploiting the sparsity and correlation of multi-user cascaded channels.  Methods  Under the assumption that all User Equipment (UE) shares the same RIS-BS channel and is affected by a common RIS blockage vector, the problem is divided into two stages. First, a target UE is selected. The sparsity of the mmWave channel and blockage vector, together with the linear dependence among RIS-BS paths, is used to formulate a sparse recovery problem. A hierarchical Bayesian model is then adopted, and an efficient Sparse Bayesian Learning (SBL) algorithm is used for joint recovery. Second, partial Channel State Information (CSI) obtained from the target UE is used to construct a common channel matrix that combines the RIS-BS channel and blockage information. Channel estimation for the remaining UEs is then reformulated as another sparse recovery problem.  Results and Discussions  A low-complexity strategy for cascaded channel estimation and blockage diagnosis is developed by exploiting the sparsity and correlation of multi-user cascaded channels and the commonality of the RIS blockage vector. Ideal estimation results are used as a theoretical lower bound, and the proposed algorithm is compared with two benchmark schemes. Simulation results show that the proposed algorithm consistently outperforms the benchmark schemes (Fig. 1). Specifically, a higher target-user Signal-to-Noise Ratio (SNR) improves the Normalized Mean Square Error (NMSE), which confirms the importance of target-user selection (Fig. 2). The algorithm also shows good convergence as the number of iterations increases (Fig. 3), and its performance approaches the ideal case more closely as the number of time frames increases (Fig. 4). In addition, the method remains robust as the number of blocked elements increases (Fig. 5). More BS antennas further improve performance by enhancing array orthogonality (Fig. 6). By exploiting path correlation, the proposed method achieves better estimation accuracy with slightly lower runtime (Table 1). However, estimation accuracy decreases as the number of paths increases because the model becomes more complex (Figs. 7 and 8).  Conclusions  This paper proposes a joint channel estimation and blockage diagnosis framework for blocked RIS-assisted multi-user multipath mmWave systems. Simulation results show that the method approaches the theoretical performance bound in complex multipath environments. It also maintains clear performance advantages under high blockage rates while reducing computational complexity through the use of common channel structures. This study provides a practical solution to performance degradation in RIS deployment, clarifies the effects of key parameters, and offers guidance for system design. Because practical blockages often exhibit block-sparse or structured-sparse characteristics, future work may incorporate structured priors, such as group sparsity and Markov random fields, into the SBL framework to capture spatial correlation and improve diagnostic accuracy and robustness.
Optimal Weighted Subspace Fitting-based Direct Position Determination with HF/VHF Collaboration
YANG Gao-yuan, YIN Jie-xin, WANG Ding, YANG Bin
Available online  , doi: 10.11999/JEIT260001
Abstract:
  Objective   Passive localization is essential for target detection, navigation, and track tracking, particularly in military applications involving maritime and aerial targets. These targets often transmit across multiple frequency bands, including shortwave High Frequency(HF) and Very High Frequency (VHF). Existing localization methods largely rely on single-band approaches or two-step positioning techniques. Single-band methods underutilize the positional information available across different bands, while two-step methods lose information during intermediate parameter estimation (e.g., Direction-Of-Arrival (DOA); Time-Difference-Of-Arrival (TDOA)), reducing localization accuracy. Collaborative fusion of HF signals (via ionospheric reflection) and VHF signals (via Doppler effects from moving arrays) has been rarely addressed. To overcome low positioning accuracy and limited spatial resolution in over-the-horizon multi-target scenarios, this study proposes a novel collaborative Direct Position Determination (DPD) method designed to integrate the complementary strengths of HF and VHF signals, enhancing localization precision and robustness in complex electromagnetic environments.  Methods  An Optimal Weighted Subspace Fitting (OWSF) DPD algorithm is proposed. Comprehensive signal propagation models are established for heterogeneous observation platforms (Fig. 1). HF signal propagation is modeled using a two-dimensional DOA framework based on ionospheric reflection, incorporating azimuth and elevation angles to handle nonlinear over-the-horizon propagation. VHF signals are modeled using a space-time extended signal framework for a moving Unmanned Aerial Vehicle (UAV), exploiting Doppler effects to create a virtual large-aperture array that captures both one-dimensional angle and Frequency-Of-Arrival (FOA) information. Unlike traditional methods that process each band separately, the OWSF algorithm constructs a unified cost function that fuses the signal and noise subspaces of both HF and VHF data using optimal weighting matrices, balancing the contributions of different signal qualities. Target positions are then estimated by minimizing this cost function via grid search or Newton iteration. The Cramér-Rao Bound (CRB) under Earth-ellipsoid constraints is derived to provide the theoretical performance limit.  Results and Discussions   Simulations are conducted in a centralized processing scenario, where HF stations and UAV VHF signals are transmitted to a central station for joint processing (Fig. 2). The simulation involves three stationary targets and a collaborative system comprising HF stations and a UAV (Fig. 3, Table 2, Table 3). Performance comparisons demonstrate that the OWSF method consistently outperforms traditional two-step positioning methods and single-system DPD methods (DOA-only or FOA-only) in Root Mean Square Error (RMSE) (Fig. 4). When HF SNR is 5 dB lower than VHF SNR, OWSF exhibits superior robustness compared to Subspace Data Fusion (SDF) and Minimum Variance Distortionless Response (MVDR) methods, approaching the CRB at high SNR (Fig. 5). The impact of system parameters is further analyzed, showing that increasing the number of sampling points (Fig. 6) and array elements (Fig. 7) improves accuracy, particularly in low SNR regimes. Regarding spatial resolution, the OWSF algorithm generates sharper spectral peaks for distant targets and successfully resolves closely spaced targets that the SDF-DPD algorithm fails to distinguish (Fig. 8, Fig. 9).  Conclusions   The HF/VHF collaborative DPD method effectively integrates multidimensional observational information from ionospheric reflection and Doppler-based propagation. Simulation results demonstrate substantial improvements in localization accuracy, spatial resolution, and robustness, especially under low-SNR conditions or heterogeneous signal quality between bands. The derived CRB provides a solid theoretical benchmark, confirming that the method overcomes the limitations of single-band and two-step approaches. This approach offers a highly effective solution for over-the-horizon passive localization of multiple stationary targets.
Household Appliance Plastics Identification by Fusing Multi-Level Feature Enhancement and Hierarchical Classification
CHONG Penghao, ZHENG Yunlong, YANG Aosong, GUO Mengci, LI Shifeng
Available online  , doi: 10.11999/JEIT260084
Abstract:
  Objective  Accurate plastic identification remains challenging in waste household appliance recycling under low-resolution spectral conditions. In practical recycling environments, plastics often have complex compositions, surface contamination, and aging effects, which increase classification difficulty. Black plastics are especially difficult to identify because their strong light absorption and spectral overlap in the Visible-Near Infrared (Vis-NIR) range reduce feature separability and degrade classification performance. Under these conditions, conventional single-stage classification models often fail to maintain stable accuracy. To address this problem, an automated identification method is proposed for low-dimensional multispectral feature spaces. The method aims to improve the discriminative capability of limited spectral information and enhance classification accuracy for complex plastic categories.  Methods  A compact Vis-NIR multispectral acquisition system based on the AS7265x sensor is used to collect 18-channel reflectance data in the 410~940 nm range. A handheld acquisition device with a controlled optical structure is designed to reduce environmental interference and ensure measurement consistency (Fig. 3). A total of 576 samples are collected from five typical household appliance plastics, including Acrylonitrile Butadiene Styrene (ABS), High-Impact PolyStyrene (HIPS), PolyPropylene (PP), Acrylonitrile Styrene copolymer (AS), and Polycarbonate/Acrylonitrile Butadiene Styrene (PC+ABS) blends. These samples are obtained from waste household appliances and are subjected to preliminary surface cleaning before spectral acquisition. To improve feature representation, a multi-level feature engineering strategy is adopted. This strategy integrates original spectral intensity features, nonlinear polynomial expansion features, and adjacent-channel ratio features to characterize both global and local spectral information. The nonlinear expansion enhances the representation of reflectance variations, whereas the ratio features capture local spectral-shape changes and reduce external disturbances. These features are combined into a 53-dimensional feature vector. Linear Discriminant Analysis (LDA) is then applied to enhance interclass separability. To address spectral overlap and class imbalance, a Hierarchical Joint Classifier (HJC) is constructed. HJC uses a two-stage classification framework. In the first stage, an XGBoost-based primary classifier performs coarse classification to separate easily distinguishable samples and group spectrally similar black plastics. In the second stage, a TabTransformer-based secondary classifier performs fine-grained classification of difficult samples (Fig. 6). This hierarchical design reduces classification complexity and improves discrimination for challenging categories. Model performance is evaluated using five-fold cross-validation and an independent test set. Accuracy, precision, recall, and F1-score are calculated from confusion matrices (Fig. 7). Comparative experiments are conducted with traditional machine learning methods, ensemble learning models, and deep learning approaches under different feature-processing strategies (Fig. 8, Fig. 9).  Results and Discussions  The proposed HJC achieves a classification accuracy of 97.4% in five-fold cross-validation and 93.1% on the independent test set (Table 4). Compared with single-stage classifiers and methods without feature enhancement, the proposed method provides higher performance and greater stability under low-resolution spectral conditions. Comparative results show that the proposed method outperforms baseline approaches, such as PCA combined with CNN, which achieves an accuracy of approximately 71.3% on the same dataset (Fig. 8). This improvement indicates that the proposed feature engineering strategy effectively strengthens the discriminative capability of low-dimensional spectral data. Combining LDA with feature engineering further improves class separability compared with conventional PCA-based methods. Confusion matrix analysis shows that misclassifications mainly occur between spectrally similar black ABS and black HIPS samples, whereas most other categories achieve high classification accuracy (Fig. 9). These results indicate that spectral overlap remains the main challenge under low-resolution conditions. The hierarchical classification strategy reduces this problem by focusing classification resources on difficult samples, thereby improving the overall generalization ability of the model. Overall, the proposed method shows robustness under practical conditions, including spectral noise, limited channel resolution, and material heterogeneity. These results indicate its suitability for real-world recycling applications.  Conclusions  A hierarchical classification method with multi-level spectral feature engineering is developed for plastic identification under low-resolution Vis-NIR conditions. Nonlinear and spectral-shape features are incorporated into a two-stage framework to improve the identification of spectrally similar materials. The results show stable accuracy across different plastic types. The method is suitable for automated sorting in waste household appliance recycling and can be extended to other material identification tasks with limited spectral information.
SG-DDPG-based Low-intercept Point Beam Design for FDA-MIMO Short-range Detectors
JIA Jinwei, GAO Min, HAN Zhuangzhi, LIU Limin, YIN Yuanwei
Available online  , doi: 10.11999/JEIT260010
Abstract:
  Objective  Radio short-range detectors are widely used in many detection systems. However, in modern battlefields, the electromagnetic environment is increasingly complex, and radio short-range detectors must withstand various forms of electromagnetic interference. In particular, fourth-generation jammers based on Digital Radio Frequency Memory (DRFM) can implement repeater deception jamming. Such jamming may cause failures such as premature detonation in radio short-range detectors and reduce their damage effectiveness. Anti-repeater deception jamming has therefore become a key issue for short-range detectors. Improving the Low Probability of Intercept (LPI) performance of radio short-range detectors is an effective means of resisting repeater deception jamming. According to the Chinese manuscript, this study focuses on the effect of FDA-MIMO array-element frequency-offset settings on beam synthesis and proposes an SG-DDPG-based method for LPI point beam design.  Methods  Frequency Diverse Array-Multiple-Input Multiple-Output (FDA-MIMO) technology is used in this study, and the key factors affecting beam convergence are analyzed. For the spatial LPI beam design of radio short-range detectors, a performance evaluation model for spatial LPI beams is constructed. An FDA-MIMO LPI point beam design method based on the Stage Guidance-Deep Deterministic Policy Gradient (SG-DDPG) algorithm is then proposed. In the SG-DDPG algorithm, a multidimensional staged guidance reward function is designed. An Actor-Critic model is used to maximize the reward value through gradient ascent. The array-element frequency offsets that provide better beam convergence in the current environment are then obtained. The SG-DDPG algorithm is suitable for LPI point beam design under different fall angles of radio short-range detectors. It overcomes the technical limitation of formula-based frequency-offset calculation, which is only applicable when the detector fall angle is close to vertical.  Results and Discussions  The simulations show that, after the array-element frequency offsets are optimized by the SG-DDPG algorithm, the FDA-MIMO beam achieves a half-power beam width of 1 m in the range dimension and 9.9° in the angular dimension. The proposed method provides better beam convergence and LPI performance than classical frequency-offset design methods. These results indicate that the proposed algorithm offers an effective approach for array-element frequency-offset optimization and LPI point beam design, thereby improving the LPI performance of radio short-range detectors.  Conclusions  This paper presents an FDA-MIMO LPI point beam design method based on the SG-DDPG algorithm, with the array-element frequency offset used as the optimization objective. The simulation results support two main conclusions. First, the proposed method removes the restriction that the fall angle of the radio short-range detector must be close to vertical when the array-element frequency offset is calculated by a formula-based method. The algorithm can be applied to LPI beam design under different fall angles and improves the LPI performance of radio short-range detectors. Second, the proposed method achieves a half-power beam width of only 1 m in the range dimension and 9.9° in the angular dimension, which is better than that of traditional methods. Under different fall angles, the beam formed by the proposed method has the smallest intercept area, indicating the best LPI performance.
Performance Optimization and Gate Oxide Electric Field Analysis of 1200V Trench SiC MOSFET Based on PCL-CSL Collaborative Design
FANG Shaoming, LI Hongda, GAO Yuan
Available online  , doi: 10.11999/JEIT260164
Abstract:
  Objective  1 200 V Silicon Carbide (SiC) trench Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) are key devices in medium- and high-voltage power conversion systems. They feature high switching performance, low conduction loss, and high-temperature stability. However, conventional trench structures suffer from electric-field concentration at the trench corner and bottom gate oxide. This effect can cause the peak gate oxide electric field to exceed the industrial reliability criterion of 3 MV/cm, reducing long-term reliability. In addition, strong trade-offs exist among breakdown voltage, specific on-resistance, threshold voltage, and peak gate oxide electric field. These trade-offs make it difficult to achieve high efficiency and high reliability at the same time. To address these issues, this work studies a synergistic structure that combines deep P-type Column (PCL), Carrier Storage Layer (CSL), and locally thickened gate oxide. The aim is to regulate the electric-field distribution, suppress electric-field concentration, improve carrier transport, and achieve balanced device performance. This study provides a systematic design method for high-reliability and high-performance 1 200 V Trench SiC MOSFETs for industrial applications.  Methods  Numerical device simulations were performed using a Technology Computer-Aided Design (TCAD) platform to analyze and optimize the electrical performance of 1 200 V Trench SiC MOSFETs. To ensure reliable simulations, physical models were used for bandgap narrowing, Shockley-Read-Hall (SRH) recombination, Auger recombination, avalanche breakdown, incomplete dopant ionization, doping- and temperature-dependent mobility, and high-field mobility saturation. A device structure with deep PCL, CSL, and locally thickened bottom gate oxide is constructed to reduce the peak gate oxide electric field and improve device reliability. Key structural and process parameters were swept and quantitatively analyzed. These parameters included epitaxial layer thickness (TEpi), epitaxial layer doping concentration (NEpi), trench width, trench depth, P-Well (PW) implantation dose, PCL spacing, and CSL implantation dose. Static electrical characteristics, including threshold voltage (Vth), specific on-resistance (Ron,sp), Breakdown Voltage (BV), and peak gate oxide electric field (Eox,max) are extracted and evaluated. The final parameter combination is finally determined through a trade-off analysis between conduction performance and long-term device reliability.  Results and Discussions  The simulation results show that the deep PCL structure redirects electric-field lines away from the trench bottom gate oxide and reduces electric-field concentration. When this structure is combined with the locally thickened bottom gate oxide, Eox-max is reduced below 3 MV/cm, meeting the industrial reliability criterion. The CSL broadens the vertical conduction path, reduces current crowding, and decreases Ron,sp. Parameter optimization shows that TEpi, NEpi, trench dimensions, PW implantation dose, and CSL implantation dose determine the trade-off between BV and conduction performance (Fig. 5, Fig. 6, Fig. 9, Fig. 10, and Fig. 19). PCL spacing has a strong effect on electric-field shielding and gate oxide protection (Fig. 16 and Fig. 17). After multi-parameter optimization, the device achieves VTH=4.7 V, BV=1 708 V, Ron,sp=1.57 mΩ·cm2, and Eox-max=2.5 MV/cm (Table 2). These results indicate balanced performance for high-voltage power applications.  Conclusions  A synergistic PCL-CSL structural design for 1 200 V Trench SiC MOSFETs is studied and validated through TCAD simulation. The design addresses key limitations of conventional Trench SiC MOSFETs, including high peak gate oxide electric field, limited breakdown capability, and the trade-off between conduction performance and reliability. The effects of TEpi, NEpi, trench dimensions, PW implantation dose, PCL spacing, and CSL implantation dose on device performance and gate oxide reliability are clarified through parameter sweeping and comparative analysis. With coordinated structural optimization, the optimized device achieves low Ron,sp, high BV, suitable VTH, and suppressed electric-field concentration near the trench bottom oxide. Eox-max is controlled below the 3 MV/cm industrial reliability criterion, which reduces the risk of oxide degradation under high-bias operation. The proposed structural strategy and optimization method provide guidance for the design, simulation, and process development of high-voltage, high-reliability SiC power devices.
Multipath Scheduling Algorithm for UAV Video Streaming
CAO Changlong, LI Lingzhi, SHI Lianmin, ZHAO Qingyue
Available online  , doi: 10.11999/JEIT260002
Abstract:
  Objective   With the rapid growth of the low-altitude economy, Unmanned Aerial Vehicle (UAV) technology has been widely used in emergency rescue, disaster monitoring, urban security, and other applications. In these scenarios, stable, low-latency, and high-fidelity video backhaul is critical for task execution. Multipath transport protocols can improve Quality of Experience (QoE) through bandwidth aggregation, providing an effective basis for UAV video streaming. However, under dynamic and heterogeneous network conditions, the performance of multipath transport protocols depends strongly on the design of multipath scheduling algorithms. Existing heuristic schedulers use predefined rules to reduce head-of-line blocking and inter-path load imbalance, but their adaptability remains limited in highly dynamic environments. Learning-based schedulers can learn the mapping between network states and scheduling rewards from real-time feedback, enabling adaptive performance optimization. However, most existing learning-based schedulers are designed for general network scenarios. They are not optimized for UAV networks, and their ability to guarantee QoE has not been fully validated. A multipath scheduling algorithm tailored to UAV video streaming is therefore needed to better exploit the performance potential of multipath transport protocols.  Methods   To address the dynamic and heterogeneous challenges of UAV video streaming, this paper proposes NeuroFly, a multipath scheduling framework based on the NeuralUCB algorithm. In NeuroFly, multipath traffic scheduling is formulated as a Contextual Multi-Armed Bandit (CMAB) problem. The context space is constructed by integrating path state information, video encoding features, and UAV mobility parameters, which jointly characterize the current transmission environment. In the action space, a frame-priority-driven redundant transmission mechanism is proposed. Video frames are assigned different frame priorities according to decoding dependencies, and differentiated redundancy strategies are used to improve the probability of successful video-frame delivery. A multi-objective reward function is further designed to guide policy learning and support adaptive optimization under dynamic and heterogeneous network conditions. In addition, a context monitoring mechanism is integrated into NeuroFly to handle abrupt environmental changes caused by high UAV mobility. This mechanism detects context distribution shifts and triggers a two-stage restart strategy. A soft restart is activated when gradual context drift is detected, removing outdated historical experience. A hard restart is performed under abrupt context changes by clearing the experience replay buffer and reinitializing model parameters, allowing learning to restart under a new distribution.  Results and Discussions   The proposed NeuroFly framework is evaluated in both simulation and field environments. First, Mininet-WiFi is used to simulate realistic UAV network environments and evaluate overall QoE performance. The results (Fig. 4) show that, compared with state-of-the-art heuristic and learning-based schedulers, NeuroFly achieves broad performance gains by fully using aggregated multipath bandwidth. Specifically, the 99th-percentile latency is reduced by 19.9%~51.0%, the average video frame rate is increased by up to 24.6%, image structural similarity is improved by up to 49.2%, and the buffering time ratio is reduced by 13.4%~77.6%. These results demonstrate the strong ability of NeuroFly to guarantee QoE. Field experiments (Fig. 6) further confirm that NeuroFly provides favorable optimization in real UAV operation scenarios. Compared with mainstream transport solutions widely deployed in production environments, NeuroFly achieves better real-time transmission performance and shows strong practical applicability for future large-scale UAV deployment.  Conclusions   This paper addresses network dynamics, path heterogeneity, and time-varying transmission conditions in UAV video streaming over multipath transport protocols. An intelligent multipath scheduling framework, NeuroFly, is proposed based on the NeuralUCB algorithm. In this framework, multipath traffic scheduling is modeled as a CMAB problem. Through the design of the context space, action space, and multi-objective reward function, online learning and adaptive optimization of traffic allocation policies are achieved. To further improve robustness under severe environmental changes, a lightweight context monitoring mechanism is introduced to detect context distribution drift and restart the learning process when needed. Systematic evaluations are conducted on both simulation platforms and real UAV operation environments. The simulation results show that NeuroFly achieves consistent improvements across QoE metrics compared with state-of-the-art heuristic and learning-based schedulers. The field results further indicate that NeuroFly provides reliable guarantees in actual UAV operation scenarios when compared with mature solutions that have been widely deployed in production environments. These results validate the practicality, robustness, and engineering feasibility of NeuroFly, and suggest its potential for large-scale deployment in UAV applications that are sensitive to real-time video quality, including emergency response, power inspection, agricultural monitoring, and logistics delivery.
Research on Energy Efficiency Optimization of Rotatable Hybrid Intelligent Reflecting Surface Communication
ZHANG Guangchi, GUO Xuan, WANG Luyao, CUI Miao, FU Hao
Available online  , doi: 10.11999/JEIT260119
Abstract:
  Objective  With the evolution of 6G communication networks, reconfigurable intelligent surfaces (RIS) have emerged as a pivotal technology for reshaping wireless environments and enhancing spectral efficiency. However, conventional fixed RIS architectures face two critical challenges in practical deployment: the “angle mismatch” loss, where the effective aperture significantly diminishes when users are located at large angles from the RIS normal, and the “energy consumption bottleneck,” caused by the high cumulative power consumption of radio frequency (RF) circuits and static control elements in large-scale arrays. Existing research often treats mechanical rotation and element switching in isolation, lacking a unified framework to balance the trade-off between mechanical/circuit energy consumption and communication gain. To address these limitations, this paper investigates a rotatable and switchable hybrid RIS (H-RIS) assisted downlink communication system. The primary objective is to maximize the system’s energy efficiency (EE) by jointly optimizing the base station transmit power, subarray activation states, physical rotation angles, and electronic phase shifts. This approach aims to introduce mechanical rotation degrees of freedom to compensate for path loss and employ dynamic switching mechanisms to reduce redundant power consumption, thereby achieving sustainable green communication.  Methods  A joint optimization framework is established for the H-RIS aided single-user multiple-input single-output (MISO) system. The system model explicitly accounts for the dynamic power consumption induced by mechanical rotation and the static power consumption of active subarrays. The resulting optimization problem is formulated as a non-convex mixed-Integer non-linear programming (MINLP) problem, involving coupled binary variables (activation status) and continuous variables (power, angles, phases). To solve this challenging problem, a block coordinate descent (BCD)-based alternating optimization (AO) algorithm is proposed to decouple the variables into three sub-problems.Firstly, to tackle the exponential complexity caused by binary switching variables, a channel contribution-based ranking strategy is developed. By performing eigenvalue decomposition on the cascaded channel correlation matrix, the priority of each subarray is quantified, reducing the search space from exponential to linear.Secondly, for the power allocation sub-problem, the non-convex fractional objective function is transformed into a parametric subtractive form using the Dinkelbach algorithm, which is then solved via the interior-point method.Thirdly, for the physical rotation and electronic phase optimization, the problem is decomposed into single-variable sub-problems. A Golden Section Search algorithm is employed to iteratively find the optimal rotation angle and phase shift for each subarray within bounded constraints, ensuring the monotonic convergence of the objective function.  Results and Discussions  Extensive simulations are conducted to evaluate the performance of the proposed H-RIS scheme compared with benchmark schemes, including “Only-Rotation” (always on), “Only-Switching” (fixed angle), and “Conventional” (fixed and always on).The simulation results regarding the maximum transmit power Pmax(Fig. 2 and Fig. 3) demonstrate that the proposed method achieves the highest energy efficiency across the entire power range. Specifically, in the low power regime, the proposed algorithm intelligently turns off redundant subarrays where the rate gain cannot offset the circuit power cost, thereby significantly outperforming the “Only-Rotation” scheme which suffers from high static power consumption.The impact of user distance is also analyzed (Fig. 4 and Fig. 5). Results indicate that the proposed scheme maintains high spectral efficiency comparable to the “Only-Rotation” scheme by dynamically adjusting the rotation angles to align with the Line-of-Sight (LoS) path, effectively compensating for the angle mismatch loss observed in the “Only-Switching” and “Conventional” schemes.Furthermore, the activation pattern of the subarray varies in a “U” shape with distance (Table 1), which allows for flexible adjustment of array size and orientation according to user-RIS geometry.  Conclusions  This paper proposes an energy-efficient transmission scheme for H-RIS aided communication systems by integrating mechanical rotation and dynamic switching capabilities. A low-complexity BCD-based algorithm is developed to jointly optimize the transceiver design. The results confirm that introducing mechanical rotation significantly mitigates the angle mismatch loss, while the proposed channel contribution-based switching strategy effectively eliminates redundant energy consumption. The proposed H-RIS architecture offers a superior trade-off between spectral efficiency and energy efficiency compared to traditional fixed RIS architectures, providing a viable solution for future green 6G networks.
CRLB Optimization for O-RIS-Assisted VLP Systems
ZHANG Zengjie, WU Qi, ZHANG Jian, DUAN Ruijie, FENG Yunhan
Available online  , doi: 10.11999/JEIT260120
Abstract:
  Objective  With the rapid development of indoor location-based services, Visible Light Positioning (VLP) has emerged as a promising high-accuracy positioning technology. The integration of Optical Reconfigurable Intelligent Surfaces (O-RIS) into VLP systems can effectively enhance signal coverage and improve positioning performance. However, optimizing the positioning accuracy and fairness across different user areas in RIS-assisted VLP systems remains a challenging issue. This study focuses on optimizing the Cramer-Rao Lower Bound (CRLB) of the system under both near-field and far-field channel models, aiming to enhance overall positioning precision and fairness through RIS configuration.  Methods  Under the far-field channel model assumption, the RIS orientation optimization problem is formulated as a received power maximization problem. A positioning algorithm combining Particle Swarm Optimization (PSO) and N-step iteration is proposed to dynamically adjust the RIS orientation optimally without prior knowledge of the receiver’s position. Under the near-field channel model assumption, the allocation problem between RIS elements and LEDs is constructed as a Markov Decision Process (MDP). A reinforcement learning method based on experience replay and knowledge utilization is designed to solve this problem, aiming to minimize the CRLB while ensuring positioning fairness for users in different regions.  Results and Discussions  Simulation results demonstrate that the proposed algorithms effectively enhance system positioning performance under both models. In the far-field model, the PSO-based iterative algorithm achieves dynamic optimization of RIS orientation, significantly improving positioning accuracy (Fig. 3). Under the near-field model, the reinforcement learning approach not only minimizes the CRLB but also considerably improves positioning fairness across the entire area, with a noticeable reduction in performance disparity among users in different zones (Fig. 5, Fig. 6). Comparative experiments show that the proposed methods outperform conventional RIS configuration strategies in terms of both average positioning error and fairness index (Table 1).  Conclusions  This paper investigates CRLB optimization methods for O-RIS-assisted VLP systems under near-field and far-field channel models. In the far-field scenario, a PSO-based iterative algorithm is proposed to optimize RIS orientation, enhancing positioning accuracy without requiring prior receiver location information. In the near-field scenario, a reinforcement learning-based approach is designed to optimize RIS element–LED allocation, which effectively minimizes the CRLB and improves positioning fairness across the whole area. Simulation results validate the effectiveness of the proposed algorithms in both models. Future work may consider more practical channel impairments and multi-user scenarios to further improve the robustness and scalability of the system.
Research on UAV-assisted Dynamic-weight Edge Computing Offloading Strategy
WANG Yijun, WANG Yachu, SHAHD Batool, MIAO Ruixin
Available online  , doi: 10.11999/JEIT260054
Abstract:
  Objective  The increasing demands of the Internet of Things (IoT) for computational resources and real-time processing have highlighted the significance of Mobile Edge Computing (MEC). Traditional MEC relies on terrestrial base stations, resulting in coverage blind spots in remote or specialized environments. Unmanned Aerial Vehicle (UAV)-assisted MEC architectures exploit UAVs’ flexible deployment to expand service coverage. However, existing approaches for multi-terminal, multi-UAV scenarios often fail to optimize task offloading latency, system energy consumption, and adaptability to dynamic environments simultaneously. They also overlook optimal UAV selection when terminal devices are covered by multiple UAVs and lack adaptive mechanisms to adjust optimization objectives during task execution. This study addresses these challenges by integrating cooperative caching, offloading decision-making, and resource allocation strategies.  Methods  A three-tier microcloud-edge-terminal architecture is constructed, comprising a central cloud, multiple UAV edge servers with caching capabilities, and numerous mobile terminal devices. A cooperative caching mechanism reduces transmission delay during task execution. Task offloading adopts a fine-grained partial offloading mode, dividing complex tasks into dependent subtasks modeled through a Directed Acyclic Graph (DAG). The Cooperative Caching-Adaptive Hierarchical MultiVerse Optimizer (CCAH-MVO) algorithm is proposed. A hybrid coding scheme encodes offloading decisions, caching decisions, and resource allocation uniformly. A dynamic weight mechanism adaptively balances delay and energy consumption according to the system’s real-time energy state. Additionally, a UAV selection strategy is implemented for scenarios where terminals are covered by multiple UAVs. By simulating inter-universe material exchange and local refined search, the algorithm efficiently determines the optimal offloading strategy. MATLAB simulations validate the method under various experimental settings.  Results and Discussions  The simulation scenario involves 50 randomly distributed terminal devices and 5 UAVs in a 400 m × 400 m area. UAVs are deployed above terminal cluster centers, while terminals at cluster edges are simultaneously within the coverage of multiple UAVs (Fig. 5). The optimal UAV for each terminal is selected using the UAV selection function (Fig. 6), preventing resource bottlenecks and achieving balanced load distribution. In terms of delay performance, the CCAH-MVO algorithm maintains the lowest task delay across all task volumes, with a gradual increase as the number of tasks grows (Fig. 7). Delay under CCAH-MVO is consistently lower than that under fixed-weight strategies across the full task range, demonstrating the effectiveness of the dynamic adaptive mechanism in preserving low latency (Fig. 10). For energy consumption, differences among the algorithms are minor when task quantities are low. Under high task loads, the activation of the dynamic weight mechanism flattens the energy consumption curve (Fig. 8). When the number of tasks reaches 100, total energy consumption under CCAH-MVO is the lowest among all strategies and remains lower than the fixed-weight approach, reflecting effective control under critical energy conditions (Fig. 9). Regarding total system overhead, the CCAH-MVO algorithm consistently achieves the best performance. The gap with fixed-weight strategies widens when task numbers exceed 80, illustrating the dynamic weight mechanism’s collaborative optimization of delay and energy consumption (Fig. 11). Overall, by integrating the dynamic weight mechanism and balancing load through UAV selection, the CCAH-MVO algorithm effectively mitigates resource constraints and high task processing overhead in complex, dynamic UAV-assisted MEC environments. It ensures precise coordination between task delay and energy consumption across different load stages.  Conclusions  The proposed CCAH-MVO framework, incorporating a microcloud-edge-terminal architecture, cooperative caching mechanism, fine-grained partial offloading, dynamic weight adjustment, and UAV selection strategy, effectively addresses resource scheduling in complex multi-UAV MEC environments. Simulations show adaptive optimization of objectives, intelligent energy management, low latency, and reduced total system overhead, improving service stability and user experience. This research provides a practical solution for efficient UAV edge computing in dynamic environments. Future work will explore dynamic energy efficiency optimization and multi-node collaboration while maintaining low-latency performance.
A Radio Frequency Fingerprint Open-set Identification MethodCombining Multi-scale Wavelet Front-end and Hyperspherical Metric Learning
TIAN Xinyu, LI Zirui, ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260214
Abstract:
  Objective  Open-set Radio Frequency Fingerprint (RFF) identification under low Signal-to-Noise Ratio (SNR) conditions is challenging because fingerprint features are easily masked by noise, multipath effects induce nonlinear distortions, and existing methods struggle with feature extraction and unknown device detection. This study proposes a deep learning framework that integrates a multi-scale wavelet front-end with hyperspherical metric learning to achieve robust open-set RFF identification.  Methods  The proposed method, MS-RANet, comprises three key components. First, a multi-scale wavelet front-end based on one-dimensional stationary wavelet transform performs full-resolution, multi-scale decomposition of I/Q signals, preserving discriminative fingerprint information while suppressing noise. Second, a multi-scale residual attention network incorporates deep residual learning, global self-attention, and Bidirectional LSTM (BiLSTM) to enhance sensitivity to subtle fingerprint features and capture long-range temporal dependencies. Third, hyperspherical metric learning constrains the feature space onto a unit hypersphere, optimizing angular margins to produce compact intra-class and separable inter-class feature distributions. Unknown devices are subsequently detected using cosine similarity.  Results and Discussions  Experiments on a high-fidelity IEEE 802.11 simulation dataset demonstrate the effectiveness of MS-RANet. The method achieves an average classification accuracy of 65.34% across SNR levels from –5 dB to 20 dB, and an Area Under the Curve (AUC) of 0.81 at –5 dB SNR, outperforming DNN, GRU, CNN-LSTM, ResNet50, and DRSN-CA. Confusion matrices and Receiver Operating Characteristic (ROC) curves confirm robustness under extreme channel conditions. t-SNE visualization shows well-separated, compact clusters for known devices, while unknown samples are effectively isolated from known class regions. Ablation studies verify the contributions of the multi-scale wavelet front-end, global attention, BiLSTM, and hyperspherical metric learning modules.  Conclusions  This study presents a robust open-set RFF identification method combining a multi-scale wavelet front-end with hyperspherical metric learning. The framework exhibits strong noise resilience, enhanced feature discrimination, and reliable detection of unknown devices under low-SNR and multipath fading conditions. Future work will focus on reducing computational complexity, improving inference speed, evaluating generalization across diverse scenarios and protocols, and integrating the method with complementary physical-layer security mechanisms for collaborative authentication.
A Survey of Processor Security
CHEN Congcong, GU Zhiyang, ZHANG Jiliang
Available online  , doi: 10.11999/JEIT260026
Abstract:
  Significance   Processor security is a cornerstone of modern information security. Cryptographic algorithms, operating systems, and applications have long relied on processors as trusted computing bases. However, as Moore’s Law slows, modern processors increasingly adopt aggressive microarchitectural optimization techniques to improve performance and energy efficiency, often without sufficient security consideration. This trend has led to frequent security vulnerabilities in recent years. In particular, microarchitectural timing channels, exemplified by Meltdown and Spectre, exploit timing differences caused by microarchitectural state changes to break fundamental hardware and software isolation, affecting billions of devices worldwide. At the same time, the boundary between architectural and microarchitectural behavior has become less clear, giving rise to new attack paradigms and turning timing channels from isolated hardware flaws into cross-layer system security problems.  Progress   Although substantial progress has been made in the study of timing channels, existing surveys still have several limitations. First, the mechanisms of timing channels are highly diverse, and the set of exploitable components continues to grow. Hardware-centric classification schemes are therefore insufficient to capture emerging and previously unknown attacks, and they often obscure the common features shared across different techniques. Second, as traditional microarchitectural channels become better understood and partially mitigated, leakage increasingly shifts to higher-level shared resources, including operating system policies and software-managed shared resources. However, previous studies have often treated software mainly as an execution context rather than a direct source of timing leakage. In addition, current discussions of defenses tend to emphasize individual techniques, with limited analysis of their scope and failure modes.  Contributions   This survey systematically reviews timing channels from a cross-layer perspective and unifies hardware- and software-based timing channels under a common abstraction. Four necessary conditions for timing channel exploitation are identified, and a unified classification framework is established based on the nature of shared mutable state and the mechanisms that make timing differences observable. Within this framework, representative attacks from the past decade are comprehensively reviewed, their attack procedures are systematically analyzed, and their common features are clarified. In addition, existing defense mechanisms are classified according to the leakage conditions they are intended to disrupt, and their scope and possible failure modes are examined. This survey also reviews current automated vulnerability detection methods.  Prospects   Future research on timing channels faces several emerging challenges. New microarchitectural optimization techniques continue to create new attack surfaces, while resource sharing at the software level may produce additional forms of timing leakage. Moreover, emerging platforms, including chiplet-based architectures, cloud computing environments, hardware accelerators, and heterogeneous systems, are likely to expose new types of timing channels that require systematic study.
A Cross-Precision Motion Compensation Technique for Security Surveillance Video Coding
JIANG Wei, MA Wei, LU Jinghui, ZHANG Yue, ZHANG Yundong
Available online  , doi: 10.11999/JEIT251301
Abstract:
  Objective  In the field of modern security surveillance, high-altitude dome cameras are often deployed at critical locations such as bridges and tower tops that are susceptible to external interference, resulting in problems such as jitter and blurring in captured videos, which pose great challenges to video coding. In video compression coding, high-precision motion compensation is the key to improving coding efficiency. The existing Ultimate Motion Vector Expression (UMVE) technique suffers from insufficient precision and lack of flexibility in adaptive adjustment. Although high-precision coding tools such as Registration-Based Coding Mode (RCM) and Affine Motion Compensation Prediction (AFFINE) can improve compensation accuracy, they have disadvantages of high computational complexity and hardware cost, making it difficult to meet the multiple requirements of coding efficiency, power consumption and real-time performance in high-altitude surveillance scenarios. Therefore, aiming at the core pain points of video coding for high-altitude dome cameras, it is of important academic value and practical application significance to design an optimized UMVE scheme that combines high-precision motion compensation, low computational complexity and scene adaptability, so as to improve coding efficiency and balance resource consumption.  Methods  This study proposes an Ultimate Motion Vector Expression technique supporting Cross-Precision Motion Compensation (UMVE_CPMC). Its core is to improve motion compensation accuracy by constructing an extended Up-Precision Motion Vector (UPMV), whose mathematical expression is UPMV = BaseMV + MMV(p, angle), where BaseMV is the basic motion vector obtained by the existing UMVE method, and MMV is the refined fine-tuning motion vector based on specific precision p and angle, with incremental candidates only provided at the 1/8 precision level to balance computational complexity and compression efficiency. For step-size adaptive adjustment, an improved scheme with six modes is proposed, covering enhanced UMVE, conventional UMVE and four precision-improved modes, allowing the encoder to switch flexibly according to scene characteristics. The average image gradient is adopted as an objective evaluation index; test scenes are divided into Class A (high-definition motion scenes) and Class B (low-definition scenes), and different coding configurations, sequences and parameters are set to compare coding gains and computational efficiency under different modes.  Results and Discussions  Experiments show that UMVE_CPMC achieves effective performance improvement in various scenes and modes. In Class A high-definition motion scenes, with the adaptive strategy disabled and RCM disabled, the average gains of Y, U and V components in Fusion Mode 1 reach -2.912%, -1.656% and -1.654% respectively, and the average coding time is reduced to 94.55% of the baseline; the average gain of the Y component in Independent Mode 1 reaches -2.925%, with coding time reduced to 91.91% of the baseline. Compared with traditional UMVE, when CPMC Independent Mode 1 is enabled under the scenario where RCM is enabled and other tools work collaboratively, the gain is improved from -0.276% to -1.310%, showing significantly higher cost performance. In Class B low-definition scenes, after enabling adaptive adjustment, the gain losses of Fusion Mode 1 and Mode 0 are significantly reduced, with average gain losses controlled at 0.071% and 0.108% respectively, successfully maintaining the original coding gain. In multi-scene comprehensive tests, when RCM and AFFINE are disabled, 9 out of 10 test sequences in adaptive Fusion Mode 1 show positive gains, including a Y-component gain of -10.691% for the yuxuedaolu sequence and -11.400% for the BQTerrace sequence. When all existing coding tools are enabled, the Y-component gains of dianjing, yuxuedaolu and BQTerrace sequences reach -1.29%, -2.05% and -1.21% respectively, with coding time reduced to 94%–96% of the baseline. In addition, correlation analysis between average image gradient and gain reveals a significant positive correlation: images with high average gradient (high definition) achieve greater gains from UMVE_CPMC, while those with low average gradient (low definition) hardly benefit. Principle analysis indicates that pixel changes in low-definition images are gentle, and high-precision interpolation fails to generate effective pixel values, resulting in insignificant compensation effects. Performance differences among modes match computational complexity: the fusion mode balances gain and stability, while the independent mode further reduces computation. The six step-size adaptive modes can meet real-time and precision requirements of different scenes.  Conclusions  The proposed UMVE_CPMC technique, by integrating cross-precision motion compensation with the UMVE algorithm, effectively solves the core problems of insufficient precision in traditional UMVE and high computational complexity of high-precision coding tools, achieving a favorable balance among coding efficiency, computational complexity and scene adaptability. This technique delivers remarkable coding gains in Class A high-definition motion scenes, with gains exceeding 10% for some sequences without other high-precision compensation tools and 1%–2% when cooperating with other tools. In Class B low-definition scenes, the original coding gain can be maintained through frame-level adaptive adjustment interfaces. Meanwhile, the fusion mode does not increase hardware complexity, and the independent mode significantly reduces coding time, suitable for encoder designs with limited resources or simplified requirements. UMVE_CPMC provides a new effective approach to solving the low coding efficiency caused by jitter and blurring in high-altitude dome camera video coding, enriches the video coding toolset, and offers important practical guidance for the optimization of video coding technologies in the security surveillance field. Future work can further optimize the adaptive strategy, explore integration with other advanced coding technologies, develop personalized coding schemes, and improve performance in complex scenarios.
An Ultra-Wideband Low-Profile Dipole Patch Antenna for VHF-Band Probing Radars
TIAN Yuxiao, ZHANG Feng, MA Zhangjun, WANG Jiacheng, JI Yicai
Available online  , doi: 10.11999/JEIT260105
Abstract:
  Objective  In radar systems, the limitations of traditional narrowband antennas in data transmission rate and resolution have become increasingly evident. Ultra-WideBand (UWB) antennas therefore receive broad attention because they provide high range resolution and strong interference suppression capability. However, at low frequencies, existing UWB antennas usually suffer from excessively large physical size, which makes installation on airborne or vehicle-mounted platforms difficult. By contrast, compact antennas that are easier to deploy often exhibit insufficient gain and cannot satisfy the penetration-depth requirement of deep subsurface detection. Thus, achieving a proper balance among antenna size, bandwidth, and gain over an ultra-wideband range remains a major challenge for VHF-band probing radars. To address this issue, a planar dipole antenna loaded with an Artificial Magnetic Conductor (AMC) structure and metallic shorting walls is proposed. The antenna maintains stable radiation performance over a wide frequency range while preserving a low-profile and structurally simple configuration.  Methods  The reflection-phase characteristics of AMC unit cells with different geometries are compared, and square unit cells are selected to construct a 9 × 7 AMC reflective layer. Owing to its in-phase reflection property, the AMC structure removes the conventional requirement for a quarter-wavelength spacing between the antenna and a metallic ground plane, thereby reducing the profile height. The dipole patch adopts an optimized meandered current-bending structure to reduce the lateral size. Metallic shorting walls are further loaded at both ends of the antenna. According to image theory, equivalent currents are generated on the outer surfaces of these metal walls during operation, which effectively extends the electrical length and improves low-frequency performance without increasing the physical size. In addition, two vertical metallic walls are connected to the ground plane on both sides of the antenna to form a reflective back cavity, which strengthens unidirectional radiation and improves antenna gain. As part of the overall co-design, four 125 Ω resistors are inserted between the feed region and the metallic sidewalls. This resistive loading suppresses strong low-frequency resonances and broadens the impedance bandwidth at the cost of acceptable Ohmic loss.  Results and Discussions  A prototype with favorable simulated performance is fabricated and measured in a microwave anechoic chamber. The measured impedance bandwidth for VSWR<2 is 50~400 MHz, which agrees well with the simulated range of 84~366 MHz. The measured impedance matching is slightly better than the simulated result, mainly because cable loss and power-divider loss in the feeding network reduce the reflected power. The measured gain follows the same trend as the simulated gain, with deviations within 1 dBi. Radiation-pattern measurements show that at 100, 200, and 300 MHz, the measured copolarization patterns agree well with the simulated results, and the maximum radiation direction remains normal to the antenna plane, which confirms the effectiveness of the proposed design. As shown in Fig. 5, the current on the radiating patch layer mainly flows along the +x direction and generates a radiated electric field along the +z direction. The current on the AMC unit can be represented by an equivalent current loop oriented along the +z direction. At this frequency, the x-direction current and the parasitic current loop on the AMC jointly enhance the antenna gain. This result explains the gain-improvement mechanism of the AMC structure. When the operating frequency increases to 400 MHz, the electrical size of the antenna reaches approximately \begin{document}$ 1.6\lambda $\end{document}, which causes main-lobe splitting and shifts the maximum radiation direction toward 90°. Although this high-frequency beam splitting introduces spatial clutter, it is an acceptable physical trade-off for achieving the ultra-low profile of 0.07 λL, while the overall UWB characteristic still supports high time-domain resolution in probing radar systems. At 400 MHz, the measured H-plane co-polarization level is slightly higher than the simulated value, possibly because of coupling between the feeding cable and the vertically mounted antenna.  Conclusions  A low-profile UWB planar dipole antenna is proposed for VHF-band probing radar applications. By combining the AMC layer, metallic shorting walls, and resistive loading, the proposed design improves impedance matching while preserving a compact size. The reflective back cavity further improves the realized gain. The fabricated prototype shows good agreement between measurement and simulation. The antenna operates over 100–366 MHz and exhibits a measured VSWR<2 bandwidth of 50~400 MHz. It maintains a compact electrical size of 0.38λL × 0.18λL × 0.07λL, and the maximum measured gain within the operating band reaches 6 dBi. The proposed co-design provides a practical solution for low-frequency probing radar antennas that require wide bandwidth, low profile, and relatively high gain.
Semantic Relation-enhanced Adaptive Graph Representation Learning for Next POI Recommendation
WANG Zhuolu, XU Shenghua, WANG Yong, JIANG Shunshun
Available online  , doi: 10.11999/JEIT251357
Abstract:
  Objective  In recent years, next Point Of Interest (POI) recommendation has played an increasingly important role in Location-Based Social Networks (LBSNs). However, existing Graph Representation Learning (GRL)-based recommendation methods have struggled to balance node distributions across different domains (i.e., node types) effectively and have often overlooked feature differences among heterogeneous relations. Thus, complex semantic dependencies in contextual information cannot be fully captured when users’ temporal preference patterns are modeled.  Methods  To address these issues, a next POI recommendation method based on Semantic Relation-enhanced adaptive Graph Representation Learning (SR-GRL) is proposed. A heterogeneous transition graph is constructed to integrate three entity types, namely POIs, POI categories, and regions, and their complex interrelationships. An adaptive balanced random walk sampling strategy is designed to balance node distributions across different domains dynamically and to reduce information redundancy. A type-aware attention mechanism is then used to learn semantic associations among nodes through relation-specific transformation matrices, so that feature differences across node types can be identified effectively. The obtained disentangled POI representations are then used for spatiotemporal encoding of user check-in sequences, and a self-attention mechanism is applied to aggregate users, temporal preference features. Finally, next POI recommendation is generated through a Softmax function.  Results and Discussions  Experiments on the Foursquare datasets from Tokyo and New York and the Sina Weibo dataset from Shanghai show that, compared with state-of-the-art baselines, the SR-GRL method achieves Recall@10 improvements of 2.22%\begin{document}$ \sim $\end{document}24.16%, F1@10 improvements of 1.16%\begin{document}$ \sim $\end{document}10.48%, and NDCG@10 improvements of 3.01%\begin{document}$ \sim $\end{document}17.37%, indicating better recommendation performance.  Conclusions  Overall, the SR-GRL approach can balance the distributions of different node types dynamically and strengthen the modeling of complex semantic dependencies in heterogeneous contextual information.
Multi-Agent Deep Reinforcement Learning Strategy for Multi-Spacecraft Long-Distance Orbital Game
DI Peng, YIN Zengshan, LIN Zheng, YAO Ye
Available online  , doi: 10.11999/JEIT251384
Abstract:
This paper introduces a novel research scenario for multi-spacecraft Orbital Pursuit-Evasion Game (OPEG), which has not yet been systematically studied. To enhance the decision-making capabilities of spacecraft and enable them to formulate more robust policies in complex multi-agent games, this paper proposes a multi-agent deep reinforcement learning algorithm based on a progressive adversarial training framework to solve the game policies of each spacecraft. Two sets of examples with different orbital characteristics and various simulation conditions were set up for simulation verification, and behavioral deviation analysis is conducted to verify the robustness of the policy. The impact of different orbital characteristics, simulation conditions, and behavioral deviations on the game policy was analyzed. Simulation results show that the proposed method enables each spacecraft to formulate an effective game policy that satisfies all set constraints and has good robustness.  Objective  As the space environment becomes increasingly complex, space security has become a hot research area. The existence of a large amount of space debris and failed spacecraft poses a serious threat to high-value spacecraft in orbit. Therefore, the study of Orbital Pursuit-Evasion Game (OPEG) for non-cooperative target spacecraft has attracted widespread attention. Existing research focuses on OPEG for two spacecraft, but less on OPEG for multiple spacecraft. When there are more than two players in the game, zero-sum game design is not feasible, and it is difficult to solve using traditional methods. Furthermore, existing research ignores engineering dynamic constraints and simplifies or defines the dynamics as a two-dimensional scene when modeling the problem, which can cause considerable errors. To overcome the limitations of existing spacecraft game scenarios, this paper proposes a novel multi-spacecraft OPEG research scenario. The aim is to investigate the application of the MADRL algorithm in solving the approximate steady-state policies of each spacecraft in long-distance multi-spacecraft OPEG, highlighting the significant advantages of the MADRL algorithm in solving multi-spacecraft OPEG, and providing a feasible solution for truly realizing autonomous multi-spacecraft game play in the future.  Methods  The Multi-Agent Proximal Policy Optimization (MAPPO) algorithm based on the Progressive Adversarial Training Framework (PATF) is used to solve the optimal game policy for each spacecraft in the Multi-Spacecraft OPEG. First, a multi-constrained multi-spacecraft OPEG model is established based on actual engineering constraints, and the problem is transformed into a Decentralized Partially Observable Markov Decision Process (Dec-POMDPs). Secondly, in order to improve the decision-making ability of agents in complex multi-agent game environments and formulate more robust game policies, a novel PATF is introduced, with different reward functions designed for the specific missions of each spacecraft. Finally, two sets of simulation examples with different orbital characteristics were set up, and four different simulation conditions were set up for simulation and behavioral deviation analysis was performed.  Results and Discussions  The MAPPO algorithm based on the PATF proposed in this paper is compared with the original MAPPO (Fig. 3). The results show that the proposed method can learn effective policies more quickly, reduce ineffective exploration, and achieve a higher final convergence reward value with less fluctuation in the reward curve. This also demonstrates that the PATF can significantly enhance the decision-making ability of agents, enabling them to formulate robust policies more effectively. Simulation verification was performed using two sets of examples in four different settings (Figs. 4, 5, 6, and 7). Simulation results (Tables 3 and 4) show that the proposed method performs well in both sets of examples. Furthermore, it was verified that when the pursuer and the interceptor are on the same orbital plane, the pursuer is more likely to be intercepted. When the interceptor and the target are not on the same orbital plane, the interceptor has a relatively easier time carrying out the interception mission. This paper also analyzes the situation where both sides of the game have behavioral biases, and models this by adding control noise. Simulation results (Tables 5 and 6) show that both sides adopt relatively conservative policies to counter the control noise. The game policy formulated by the method in this paper is an approximate steady-state policy. Behavioral deviations will lead to a decrease in one’s own payoff and an increase in the opponent's payoff, and the game policy has good robustness.  Conclusions  The method proposed in this paper can be well applied to solving the long-distance OPEG problem involving multiple spacecraft in non-coplanar elliptical orbits, enabling each spacecraft to formulate excellent game policies. The PATF facilitates better decision-making by the spacecraft in complex multi-spacecraft dynamic systems, with robust control policies developed by the pursuer and interceptors. The results also demonstrate the accuracy and effectiveness of the reward function design. Through two sets of examples and simulation results with different settings, the impact on the policies of both parties when the pursuer and interceptor have different orbital characteristics is analyzed. When interceptors have different maximum thrusts, the decision-making of each spacecraft changes accordingly. The behavior deviation analysis proves that the game policies of each spacecraft have good robustness. When one party’s behavior deviates, the approximate steady-state policy balance will change, resulting in a decrease in its own benefits and an increase in the other party’s benefits. The research scenario formulated in this paper expands the scope of existing research on multi-spacecraft game problems.
Joint Power Allocation and AP On-Off Control for Long-Term Energy Efficient Cell-Free Massive MIMO Systems
WEI Siqi, GUO Fengqian, CHONG Baolin, CHENG Guo, LU Hancheng
Available online  , doi: 10.11999/JEIT260014
Abstract:
  Objective   With the rapid development of wireless communication technologies, Cell-Free Massive Multiple-Input Multiple-Output (CF-mMIMO) has emerged as an effective paradigm to overcome the limitations of traditional cell-centric networks, such as limited performance for edge users. By deploying a large number of distributed Access Points (APs) connected to a Central Processing Unit (CPU) to cooperatively serve users, CF-mMIMO improves spectral efficiency and macro-diversity gain. However, dense AP deployment also introduces a critical challenge: high energy consumption. In practical systems, if all APs remain continuously active, especially during periods of low traffic load, substantial and unnecessary energy consumption occurs. This behavior reduces network sustainability and conflicts with global “dual-carbon” goals. Existing studies on energy efficiency in CF-mMIMO systems mainly focus on short-term performance optimization. These short-term approaches often ignore long-term traffic dynamics and the requirement of queue stability. Therefore, they lack robustness under time-varying traffic conditions and may cause queue congestion and significant performance fluctuations, which are unacceptable for next-generation wireless networks with strict reliability requirements. Although several recent studies examine long-term energy efficiency optimization, most assume that all APs remain active at all times. Therefore, the energy-saving potential of adaptive AP on-off control is not fully utilized.  Methods   To address these issues, a joint power allocation and AP on-off control strategy is proposed for downlink CF-mMIMO systems. The optimization problem aims to maximize long-term energy efficiency subject to user queue stability and AP power constraints. Because the problem has stochastic and long-term characteristics, the Lyapunov optimization framework is applied to transform the original long-term fractional programming problem into a sequence of deterministic drift-plus-penalty minimization problems solved in each time slot. The resulting per-slot problems remain nonconvex. Therefore, each problem is decomposed into two subproblems: power allocation and AP on-off control. The Successive Convex Approximation (SCA) method is used to convert the nonconvex formulations into solvable convex problems. An alternating optimization algorithm is then developed to jointly solve the two subproblems, which enables adaptive resource configuration under dynamic network conditions and stochastic traffic arrivals.  Results and Discussions   The proposed algorithm is evaluated through extensive simulations. First, the convergence behavior is examined. Numerical results (Fig. 2) show that per-slot energy efficiency increases rapidly and stabilizes after several iterations, which verifies the convergence of the alternating optimization procedure. Second, the effect of the control parameter is analyzed. As the parameter increases, the algorithm places greater emphasis on energy efficiency. Average power consumption decreases and then stabilizes (Fig. 3), whereas long-term energy efficiency increases and eventually stabilizes (Fig. 4). These results confirm the trade-off between energy efficiency and queue stability. Third, the proposed scheme is compared with three baseline methods. The results (Fig. 5) show that the proposed joint optimization approach consistently achieves higher long-term energy efficiency than the baseline methods. Fourth, the necessity of long-term optimization is demonstrated by comparing queue lengths with a short-term baseline (Fig. 6). Under the same traffic arrival rate, the short-term method shows cumulative queue growth, whereas the Lyapunov-based approach maintains queue lengths within a stable range and ensures network stability. Finally, robustness under imperfect Channel State Information (CSI) is evaluated (Fig. 7). Although energy efficiency decreases as channel uncertainty increases, the proposed method consistently outperforms the baseline approaches, which demonstrates strong robustness to channel estimation errors.  Conclusions   A long-term energy efficiency optimization framework is proposed for CF-mMIMO systems with stochastic traffic arrivals. By applying Lyapunov optimization theory, the stochastic long-term problem is transformed into slot-level drift-plus-penalty problems based on queue states. This transformation enables per-slot resource scheduling decisions while maintaining queue stability. On this basis, an efficient joint resource scheduling algorithm that integrates power allocation and AP on-off control is developed. The original problem is decomposed into power allocation and AP on-off control subproblems and solved through alternating optimization. Simulation results show that the proposed method adapts to dynamic traffic conditions. By placing underutilized APs into sleep mode, the algorithm improves long-term system energy efficiency and maintains queue stability. These results provide guidance for the design of green and sustainable wireless networks.
Communication, Computation, and Caching Resource Collaboration for Heterogeneous Artificial Intelligence Generated Content Service Provisioning
WU Mengru, GAO Yu, ZHAO Bo, XU Bo, SUN Hao, GUO Lei
Available online  , doi: 10.11999/JEIT251300
Abstract:
  Objective  In the Artificial Intelligence of Things (AIoT), Edge Servers (ESs) provide intelligent content generation services to AIoT devices by utilizing cached Artificial Intelligence Generated Content (AIGC) models. However, the limited computing resources and caching capacity of ESs make it difficult to support the large-scale caching demands of heterogeneous AIGC services. To address this issue, a communication, computation, and caching resource collaboration scheme is proposed based on a combined cloud-edge and edge-edge collaborative framework. The scheme considers three representative AIGC services: lightweight AIGC services, computation-intensive AIGC services, and preprocessing-based AIGC services. The objective is to minimize the total AIGC service latency through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions.  Methods  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. First, an AIGC service-oriented AIoT system model is established to incorporate both cloud-edge and edge-edge collaboration. An optimization problem is then formulated to minimize the total latency of AIGC services through joint optimization of transmit power, computing resource allocation, model caching strategies, and offloading decisions. Because the formulated problem is non-convex, an Alternating Optimization (AO) algorithm is proposed. The original problem is decomposed into three subproblems. These subproblems are solved using the Successive Convex Approximation (SCA) method, Karush-Kuhn-Tucker (KKT) conditions, and an improved Harris Hawks Optimization (HHO) algorithm.  Results and Discussions  Simulation experiments compare the proposed joint optimization scheme with three baseline methods: Particle Swarm Optimization (PSO), fixed resource allocation, and random offloading and caching. First, the convergence of the proposed AO algorithm is verified (Fig. 2). The results show that the algorithm converges rapidly within a limited number of iterations across different subproblems. Second, increasing transmission bandwidth significantly reduces the total AIGC service latency (Fig. 3). This occurs because each device obtains more bandwidth resources for task transmission, and the ES can allocate more bandwidth to deliver generated content in the downlink. Furthermore, the total AIGC service latency decreases as the ES storage capacity increases for all schemes (Fig. 4). Greater storage capacity enables the ES to store more AIGC models, which reduces the transmission delay between the ES and the cloud server. Moreover, when the required floating-point operations per bit increase, the total AIGC service latency rises significantly across all schemes (Fig. 5). Finally, the total AIGC service latency decreases as the maximum transmit power of the Base Station (BS) increases (Fig. 6). This occurs because higher BS transmit power improves the downlink signal-to-noise ratio, which increases the downlink transmission rate and reduces overall service latency. The proposed scheme demonstrates better performance than the baseline schemes, particularly under high computational demand.  Conclusions  Communication, computation, and caching resource collaboration for heterogeneous AIGC services is investigated. The objective is to minimize total AIGC service latency through joint optimization of the transmit power of AIoT devices and BSs, computing resource allocation, AIGC model deployment, and service offloading decisions under computation and caching resource constraints. Because the formulated problem is a mixed-integer nonlinear programming problem, an efficient AO algorithm is developed. The original optimization problem is decomposed into three subproblems, which are solved using the SCA algorithm, KKT conditions, and the HHO algorithm, respectively. Simulation results show that the proposed algorithm reduces the total AIGC service latency compared with the baseline schemes.
Adversarial Attacks on 3D Target Recognition Driven by Gradient Adaptive Adjustment
LIU Weiquan, SHEN Xiaoying, LIU Dunqiang, SUN Yanwen, CAI Guorong, ZANG Yu, SHEN Siqi, WANG Cheng
Available online  , doi: 10.11999/JEIT251264
Abstract:
  Objective   Robust environmental perception is essential for intelligent driving systems. Light Detection And Ranging (LiDAR) provides high-resolution 3D point cloud data and serves as a core information source for object detection and recognition. However, deep learning models for 3D point cloud recognition show notable vulnerability to adversarial attacks. Small, imperceptible perturbations can cause severe classification errors and threaten system safety. Existing attack methods have improved the Attack Success Rate (ASR), but the perturbations they generate often lack concealment, create outliers, and show poor imperceptibility because they do not adequately preserve the geometric structure of point clouds. This reduces their suitability for realistic security evaluation of optoelectronic perception systems. Developing an attack method that maintains a high success rate while preserving geometric consistency and imperceptibility is therefore critical. This study addresses this need by proposing a framework that incorporates point cloud geometry into perturbation generation.  Methods   A Gradient Adaptive Adjustment (GAA) adversarial attack method for 3D point cloud recognition is proposed. The framework (Fig. 2) includes three coordinated modules. The 3D Point Cloud Salient Region Extraction module evaluates decision-level vulnerability using Shapley value analysis to identify and rank point subsets with the strongest influence on classifier output. Perturbations are then concentrated in these sensitive regions. A curvature-weighted gradient mechanism integrates local geometric priors. For each point in the salient region, a local covariance matrix is computed from its k-nearest neighbors. Principal component analysis generates eigenvalues and eigenvectors, which are used to compute a curvature measure. A Gaussian kernel function produces curvature-dependent weights that are applied to backpropagated gradients. This suppresses perturbations in high-curvature areas and encourages them in low-curvature regions to preserve local shape morphology. A principal curvature direction constrained 0ptimization module further refines the perturbation direction. The weighted gradient is projected onto the principal curvature directions, and the projection components are fused using coefficients derived from the corresponding eigenvalues. This aligns the perturbation with natural geometric trends and avoids unnatural deformation. An adaptive optimization algorithm then minimizes a multi-objective loss balancing attack success, geometric similarity (via chamfer distance and hausdorff distance), and perturbation sparsity. The adversarial point cloud is iteratively updated based on the saliency map, curvature-weighted gradients, and principal direction constraints.  Results and Discussions   Experiments on ModelNet40, ShapeNetPart, and KITTI were conducted using PointNet, DGCNN, and PointConv. The GAA method showed strong performance. On ModelNet40 with PointNet, it achieved a 97.69% ASR with an average of 28 perturbed points, outperforming ten baselines such as AL-Adv (92.92% ASR, 40 points) and Kim et al. (89.38% ASR, 36 points) (Table 1). It also produced lower geometric distortion, as indicated by smaller Chamfer Distance and Hausdorff Distance values. Visual results (Fig. 4) show that GAA produces fewer outliers and more natural adversarial point clouds compared with methods such as AL-Adv. The method generalized well across architectures, reaching 99.78% ASR on DGCNN and 96.91% on PointConv (Table 2), with similar performance on ShapeNetPart (Table 3). Ablation experiments on the number of salient regions (K) showed consistent improvements in ASR and reduced geometric distortion as K increased from 1 to 6 (Table 4, Fig. 5), confirming the advantage of targeting multiple critical regions. Tests on the KITTI dataset demonstrated strong performance in real-world, noisy environments. The method maintained high ASRs, such as 99.33% on PointNet, with limited perturbations (Table 5). An ablation study on K indicated that K=4 offers an effective balance between success rate and perturbation cost for PointNet (Table 6).  Conclusions   This study presents a GAA method for adversarial attacks on 3D point cloud recognition. By combining a Shapley value-based saliency analyzer, a curvature-weighted gradient mechanism, and a principal curvature direction constraint, the method generates adversarial examples that achieve high attack success while preserving geometric consistency. Experiments show that GAA minimizes perceptual distortion and perturbs fewer points across datasets and models. The method provides a practical tool for vulnerability analysis and supports the development of more robust and secure optoelectronic perception systems for intelligent driving. Future work will examine robustness under adverse conditions and assess physical-world implications.
TTSPD: A Multimodal Traffic Scene Perception Dataset Integrating Tire Data
YING Zongchen, GUI Lin, YANG Jiahan, ZHANG Fangwei, WANG Junfan, DONG Zhekang
Available online  , doi: 10.11999/JEIT260022
Abstract:
  Objective  With the rapid development of Intelligent Transportation Systems (ITS) and autonomous driving technologies, accurate traffic environment perception is a fundamental prerequisite for vehicle safety and decision making. Current perception frameworks primarily rely on high-resolution cameras and LiDAR sensors. Although these sensors provide rich information, they create severe challenges across the Perception-Storage-Calculation pipeline. High acquisition costs limit large-scale deployment. In addition, the massive data volume produced by high-dimensional sensors places heavy pressure on onboard storage and computational resources, often exceeding the power and thermal budgets of vehicle-grade edge platforms. These constraints motivate the exploration of alternative sensing paradigms that are cost-effective, compact, and computationally efficient while maintaining reliable perception accuracy. In response, the present study shifts the perception perspective from conventional external sensors to the tire-road contact interface, where abundant physical interaction information naturally exists. The objective is to construct a novel multimodal dataset, termed the Tire-integrated Traffic Scene Perception Dataset (TTSPD), which combines internal tire dynamics with external visual observations. This dataset is used to examine whether low-dimensional tire sensing data can complement or partially substitute high-dimensional visual data for accurate road surface classification. The study also aims to establish a new data morphology that balances perception performance and system efficiency for future intelligent vehicles.  Methods  To construct a high-quality and practically usable multimodal dataset, an integrated hardware-software acquisition framework is developed. From a hardware perspective, a specialized sensing system is designed by coupling tire-mounted multi-parameter sensors with a vehicle-mounted camera. To ensure reliable operation under the harsh mechanical conditions of a rotating tire, sensing nodes are encapsulated using a rubber-based composite material that provides mechanical protection and long-term stability. Wireless transmission is implemented using Bluetooth Low Energy (BLE) 5.0 with an adaptive frequency-hopping mechanism, enabling low-power and reliable communication during high-speed rotation. During data acquisition, the system synchronously collects six types of internal tire signals, including radial acceleration, tire temperature, and tire pressure, producing approximately 1.8 million sampling points. In parallel, a dashboard-mounted camera records high-resolution traffic scene images totaling 309 GB across four representative road surface conditions. To address the heterogeneity between high-frequency one-dimensional tire signals and two-dimensional visual data, a timestamp-based association strategy is adopted to achieve scene-level temporal alignment rather than strict frame-by-frame correspondence. Sensor sequences and image segments are grouped according to shared temporal windows and driving scenarios. This approach ensures semantic and temporal consistency at the scene level. The alignment strategy reflects practical deployment conditions and forms the basis of the final TTSPD dataset for multimodal fusion research.  Results and Discussions  The effectiveness of the proposed TTSPD is evaluated through comprehensive road surface classification experiments using mainstream deep learning models. Initial experiments based solely on visual data demonstrate strong baseline performance, with classification accuracies ranging from 87.25% to 93.75% (Table 7). These results confirm the quality and diversity of the visual modality in the dataset. The primary contribution of this study is the quantification of efficiency gains enabled by tire-based sensing. Comparative experiments progressively reduce the amount of visual data while integrating low-dimensional tire signals, particularly radial acceleration (Table 9). The results show that the multimodal model achieves approximately 95% of the full-data baseline accuracy while using only about 38.75% of the original data volume. This reduction in data dependency produces significant system-level benefits. Storage requirements decrease by approximately 61.25%, and overall model training time decreases by about 54.10% (Fig. 8). These findings indicate that tire dynamics encode high-value physical features related to road texture and surface conditions that complement visual cues. The proposed dataset therefore supports the development of lighter perception pipelines without reducing recognition performance.  Conclusions  This study addresses the long-standing Perception-Storage-Calculation bottleneck in vision-dominated autonomous driving systems by proposing the TTSPD. Multi-parameter sensors are embedded within tires using rubber-based encapsulation, and stable wireless communication is achieved through BLE 5.0. A robust tire-camera data acquisition system is therefore established. The resulting dataset covers four common and safety-critical road surface types: cement, asphalt, damaged, and water-covered roads. It provides a comprehensive foundation for multimodal perception research. Experimental results show that combining low-dimensional tire sensing data with visual information significantly improves perception efficiency. Approximately 95% of peak classification accuracy is achieved using only about 38.75% of the original data volume. This result effectively reduces storage pressure and computational cost, reflected in a 61.25% reduction in data storage and a 54.10% reduction in training time. The TTSPD dataset therefore proposes a practical data morphology that supports efficient and high-performance perception under vehicle-grade computational constraints. It also provides valuable resources for the future development of ITS.
Blind Parameter Estimation Method for PSK Modulated Frequency-Hopping Signals Based on Improved Maximum Likelihood
ZHANG Tianhao, ZHANG Yushu, XU Zhongqiu, TANG Xinyi, DANG Wenhua, LI Guangzuo
Available online  , doi: 10.11999/JEIT260005
Abstract:
  Objective  Blind parameter estimation of non-cooperative Frequency-Hopping (FH) signals is a critical task in electronic reconnaissance and countermeasures. Estimation methods based on time-frequency analysis typically suffer from limited resolution or high computational complexity. Furthermore, methods based on compressive sensing rely heavily on the consistency between the predefined dictionary and the actual signal characteristics, and the estimation precision will be significantly compromised by grid mismatch or modulation-induced energy dispersion. Maximum Likelihood (ML)-based methods offer the advantage of high theoretical estimation accuracy with relatively low computational complexity. However, existing studies typically assume an ideal unmodulated signal model with a single frequency transition. Consequently, these ML-based methods suffer from severe model mismatch when processing FH signals with digital modulation, such as Phase Shift Keying (PSK), or multi-hop signals. Moreover, the conventional iterative solution of ML-based methods is prone to divergence or trapping in local optima. To address these limitations, this paper proposes an improved ML-based method for the blind parameter estimation of PSK-modulated FH signals.  Methods  To handle received multi-hop signals, a signal slicing technique based on the Short-Time Fourier Transform (STFT) is proposed to extract slices containing individual frequency transitions. Subsequently, to mitigate the model mismatch caused by digital modulation in conventional ML-based methods, a model-matching signal extraction approach based on the ML objective function is developed for PSK-modulated FH signals. Furthermore, a weighted iterative solving algorithm for ML estimation is designed to enhance convergence, thereby achieving robust and accurate estimation of frequency-hopping parameters.  Results and Discussions  To validate the effectiveness of the model-matching signal extraction approach, ablation experiments were carried out under various modulation schemes, including binary PSK (BPSK), quadrature PSK (QPSK), and 8-ary PSK (8PSK). The results indicate that the proposed approach (Group D) significantly reduces the Mean Square Error (MSE) of hopping frequency estimation compared to that without the proposed extraction (Group ND). These results demonstrate that the proposed method effectively mitigates the model mismatch (Fig. 5). Simulation results also illustrate that the designed weighted iterative algorithm achieves superior convergence performance compared with linear weighting and non-weighting schemes (Fig. 6). Moreover, the experiments verify the algorithm's insensitivity to initial frequency offsets, showing that it tolerates offsets of up to 2 MHz at SNR of -10 dB with little performance degradation (Fig. 7). Finally, comparative analysis with representative existing methods indicates that the proposed method outperforms the others in terms of estimation accuracy (Fig. 8).  Conclusions  To achieve blind parameter estimation for PSK-modulated FH signals, this paper proposes an improved ML-based method. By utilizing a signal slicing technique based on the STFT, the proposed method successfully extends the applicability of the ML-based estimator to continuous multi-hop signals. To mitigate the model mismatch induced by PSK modulation, a model-matching signal extraction approach is developed to isolate valid signal segments that conform to the ML model. Furthermore, a weighted iterative algorithm incorporating a dynamic weighting function is introduced to address the instability of the conventional iterative ML solver. Simulation results confirm that the proposed method effectively eliminates model mismatch and ensures superior convergence performance with insensitivity to initial frequency offsets. Moreover, it is shown to achieve high estimation precision for both hopping frequencies and hopping times.
A Semantic-Enhanced Cybersecurity Named Entity Recognition Approach Oriented to Lightweight Adaptation of Large Language Models
HU Ze, XU Tongwu, YANG Hongyu
Available online  , doi: 10.11999/JEIT251260
Abstract:
  Objective  Named Entity Recognition (NER) in the field of cybersecurity is a fundamental technology supporting threat intelligence analysis, vulnerability management, and security incident response. However, this field generally faces challenges such as dense technical terms, scarce labeled data, dynamic changes in entity categories, and highly complex semantic features, which make traditional deep learning models and existing Large Language Models (LLMs) significantly inadequate in terms of domain adaptability and semantic fusion capability. To address the aforementioned key issues while also considering the need for lightweight model deployment, this paper aims to construct a cybersecurity NER approach that can enhance domain semantic representation, improve the ability to identify rare entities, and apply to low-resource environments, providing a reliable technical path for intelligent threat analysis in cybersecurity scenarios.  Methods  To address the complex semantic features of cybersecurity texts, this paper proposes a semantically enhanced, lightweight, and LLMs-adaptable cybersecurity NER approach. The proposed approach uses LLM2Vec to achieve bidirectional semantic reconstruction of large model decoders and combines Low-Rank Adaptation (LoRA) for low-rank fine-tuning, so as to maintain deep semantic encoding capability while significantly reducing the amount of parameter updates. To address the challenges of sparse keywords and severe noise interference in cybersecurity texts, a sparse gated attention mechanism is introduced to strengthen keyword-focused feature extraction by dynamically selecting high-contribution cybersecurity terms through global gating and sparse inference. A SecRoBERTa-based semantic enhancement component is introduced, which utilizes a domain-pre-trained model to generate similar word embeddings, optimizes feature robustness in small-sample scenarios, and alleviates the challenges of identifying out-of-vocabulary words and low-frequency terms. Finally, a masked conditional random field is employed to constrain label transitions and guarantee BIO-compliant output sequences, achieving robust and consistent entity boundary prediction.  Results and Discussions  Extensive experiments were conducted on two public cybersecurity datasets, DNRTI and APTNER. The proposed approach achieved an F1 score of 91.91% on DNRTI, surpassing the previous state-of-the-art model by 2.14%. On APTNER, it reached an F1 score of 80.37%, outperforming the best baseline by 2.97%. Ablation studies confirmed the contribution of each key component: the Sparse Gated Attention mechanism improved F1 by 3.57% over standard Multi-Head Attention on DNRTI; the semantic enhancement module contributed a 2.32% F1 gain; and the MCRF (Masked Conditional Random Field) layer provided a 10.63% F1 improvement over traditional CRF (Conditional Random Field). The model also demonstrated efficient training and inference characteristics, aligning with its lightweight design goals.  Conclusions  This paper proposes a lightweight adaptation approach based on LLMs for NER in the cybersecurity domain, which effectively addresses the limitations of existing LLMs-based NER methods in domain adaptation and rare entity recognition. By integrating LLM2Vec and LoRA for lightweight fine-tuning, a sparse gated attention mechanism for domain feature fusion, and a SecRoBERTa-based semantic enhancement component for similar word precomputation, the proposed approach achieves high performance on DNRTI and APTNER datasets. The research provides an efficient technical path for NER tasks in low-resource cybersecurity scenarios and offers strong support for downstream tasks such as automated threat intelligence analysis.
Multi-projection plane InISAR 3D reconstruction method for complex moving ship targets
LI Ning, NIU Jinfa, WANG Weibin, HU Xingwang, WU Lin
Available online  , doi: 10.11999/JEIT251268
Abstract:
  Objective  Interferometric Inverse Synthetic Aperture Radar (InISAR) is a Three Dimensions (3D) reconstruction technique for non-cooperative target. However, the complex 3D rotational motion of the ship target causes unstable Doppler frequency changes, and Inverse Synthetic Aperture Radar (ISAR) imaging inevitably suffers from target overlap and occlusion problems, making high-precision complete 3D reconstruction difficult under a single projection plane. Thus, a multi-projection planes InISAR 3D reconstruction method of complex moving ship targets based on point cloud fusion is proposed. Through efficient and high-precision point clouds registration and fusion supplement target 3D information, significantly improving the 3D reconstruction quality.  Methods  This method fully leverages the advantages of multi-plane observation from the severe movement of ship targets, extracts the ship’s centerline and estimates the vertical rotation vector via Principal Component Analysis (PCA), to select the optimal imaging time corresponding to different Imaging Projection Planes, completes ISAR imaging and InISAR 3D reconstruction. Secondly, a point cloud fusion algorithm combining Weighted Random Sampling Consensus (RANSAC) and Hierarchical Iterative Closest Point (ICP) is proposed. The random sampling process is optimized through a feature stability weighting strategy, efficiently extracting and matching corresponding feature points in InISAR images, achieving high-precision multi- Imaging Projection Plane (IPP) point cloud fusion.  Results and Discussions  Experimental results demonstrate that the proposed method significantly enhances reconstruction accuracy and target completeness. For simulated ship point target data, Fig 7 shows excellent results, with a significant reduction in reconstruction error. Signal-to-noise ratio (SNR) analysis reveals that 3D fusion imaging quality improves continuously as SNR increases from –10 dB to 10 dB, maintaining robust fusion performance even under low SNR conditions. For simulated destroyer radar cross section data, this method achieved significant registration results, and the detail recovery and structural integrity of the fused image were significantly improved, effectively solving the problem of incomplete 3D information reconstruction caused by overlapping and occlusion of scattering points.  Conclusions  To address the issues of low reconstruction accuracy and information loss caused by target rotation, overlapping, and occlusion in traditional InISAR methods for 3D reconstruction of complex moving ship targets, this paper proposes a multi-IPP InISAR 3D reconstruction method based on point cloud fusion. This method employs a PCA optimal imaging time selection strategy, By employing weighted RANSAC and hierarchical ICP algorithms to achieve efficient and high-precision registration and fusion of InISAR point clouds under multiple IPPs, obtaining high-quality 3D reconstruction results. This paper conducts multi-scenario experiments by constructing a ship model with ideal scattering points and an electromagnetic simulation RCS model with occlusion effects, verifying the accuracy of the proposed method under ideal conditions and its applicability in complex real-world scenarios.
FPGA Hybrid PLB Architecture for Highly Efficient Resource Utilization
WANG Yanlin, GAO Lijiang, YANG Haigang
Available online  , doi: 10.11999/JEIT260108
Abstract:
6-input look-up tables (LUTs) are frequently used in commercial Field-Programmable Gate Arrays (FPGAs) to build programmable logic blocks, while related experiments reveal that their average application in circuits is less than 30%, resulting in a significant waste of programmable resources. In this paper, the 6-input LUTs are fractured based on fracturable factors and recombined with different granularities to construct several new Hybrid Basic Logic Elements (HBLE). Based on HBLE, several novel Hybrid Programmable Logic Block (HPLB) architectures are proposed. Then the Programmable Logic Blocks (PLB) of Xilinx is replaced by several innovative HPLB architectures. Concurrently, a statistical evaluation algorithm for the mapped netlist is proposed. Finally, several HPLB architectures are experimentally verified and evaluated as appropriate. Experimental evaluations of the three enhanced architectures show that the HPLBs achieve an average area reduction of more than 30% when compared to Xilinx’s PLBs without adding more input ports. The hybrid HPLB architectures constructed with a fracturable factor N=3 produces the best optimization results when taking into account both HPLB utilization and area optimization. Based on the MCNC and VTR benchmarks, resource consumption increased by an average of 8.27% and 27.64%, respectively, thereby improving FPGA logic efficiency.  Objective  Currently, modern commercial FPGA architectures employ 6-LUTs as the fundamental building blocks for Basic Logic Elements (BLEs). Only about 30% of the Logic Elements (LEs) in the circuit are ultimately translated to 6-LUTs when mapping 6-LUT BLEs, according to experimental results. Nevertheless, more than half of the logic resources are wasted when 6-LUTs implement functions with inputs smaller than 6. Programmable resources will unavoidably be significantly wasted as a result. A circuit design mapped to 100 4-LUTs can be mapped to 78 6-LUTs during 6-LUT mapping studies, according to experimental data, with the {6,5,4,3,2}-LUT function distribution being {23,32,17,9,13}. The findings indicate that only around 25% of the 6-LUTs are ultimately mapped to 6-input functions, with the remaining 6-LUTs being underutilized. This illustrates even more how inefficient technical mapping is for LUTs with large input K.Methods The fracturable factor N, which is the number of sub-LUTs that may be obtained from a single LUT, characterizes the fracturable and reconfigurable nature of LUT architectures in FPGAs. Motivated by this, we decompose a 6-LUT into several granularities according to the fracturable factor in order to address the previously described problem of low resource utilization. Three novel hybrid-granularity divisible logic (HBLE) structures are created by connecting and reconfiguring the resultant sub-LUTs with additional input ports and multiplexer modules. We shall now investigate how FPGA performance is optimized by these three HBLE topologies. We shall now investigate how FPGA performance is optimized by these three HBLE topologies. One undivided 6-LUT and one divisible 6-LUT, divided into two 5-LUTs with a divisibility factor N=2, make up the HBLE2 structure. One undivided 6-LUT and one divisible 6-LUT, divided into one 5-LUT and two 4-LUTs, with a divisibility factor N=3, are included in the HBLE3 structure. One undivided 6-LUT and one divisible 6-LUT, which divides into four 4-LUTs with a divisibility factor N=4, make up the HBLE4 structure. Adder units are supported by all three HBLE structures, allowing for both latched and direct combinational logic output. Additionally, they allow direct latched output by avoiding combinational logic. A Hybrid Programmable Logic Block (HPLB) is a novel structure created by merging several HBLEs. The MCNC circuit set and the VTR circuit set, the two most well-known academic circuit benchmarks (BMs), are chosen for experimental assessment. A Xilinx Virtex-7 FPGA is used to map each circuit set. The mapped netlist is then used to tally the kinds and numbers of LUTs that were utilized. The minimum number of CLBs needed is found once the data has been arranged using the corresponding greedy algorithms. Since each Xilinx CLB has eight 6-LUTs, the greedy approach uses # Total LUT Number / 8 to determine the smallest number of CLBs needed following BM mapping. In order to guarantee similar conditions, each structure also needs to be sorted using the greedy algorithm after Xilinx’s CLB structure is replaced with the HPLB structure suggested in this research. This results in the bare minimum of HPLBs needed. It is not possible to use every LUT in the mapped CLBs during actual packing owing to routing constraints. As a result, the smallest value that may be achieved in a theoretical optimization scenario is represented by the optimized result that is acquired following greedy algorithm restructuring.  Results and Discussions  The average number of HPLBs needed for both HPLB2 and HPLB3 structures drops by about 8% when CLB structures are swapped out for HPLBs in order to map the MCNC circuit set. However, the number of HPLBs needed increases by more than 30% on average as a result of the HPLB4 structure. The needed count is smaller when HPLBs are used in place of CLBs for mapping the VTR circuit set. On average, the HPLB2 and HPLB4 counts drop by less than 10%, whereas the HPLB3 count drops by around 30%. This enables SRAM scheduling and complete input pin use. On the other hand, because of resource waste, the uniform CLB structure results in higher CLB requirements when implementing functions with a tiny LUT input K. The HPLB4 structure performs worse than the HPLB3 structure, according to post-mapping HPLB counts. Both the MCNC and VTR circuit sets achieve average area reduction ratios over 30%, according to analysis of post-mapping area optimization. All three HPLB structures attained area optimization ratios of about 31% on the MCNC test set. Different optimization effects were seen in the VTR test circuit set: HPLB2 produced an average area reduction of 30.63%, whereas HPLB4 produced an average decrease of 51.21%. The HPLB2 structure produced a 45.22% area reduction, even though its optimization effect was marginally less than that of HPLB4. A thorough examination of the area optimization results showed that a higher divisibility factor N produces more noticeable benefits for integrating small-scale LUTs in circuits, resulting in higher area reduction ratios from the enhanced architectures.  Conclusions  In order to solve the issue of low resource utilization in 6-LUTs, this research proposes three split granularity-based HPLB enhancement architectures. In addition to establishing an assessment procedure and matching algorithms for the enhanced structures, these HPLBs take the place of Xilinx’s CLB structure in order to examine the new structure’s benefits in resource utilization. Based on the proportion differences of different LUTs in the post-mapping netlist, evaluation experiments using the MCNC and VTR circuit test suites show that, although HPLB4 achieves significant area optimization, it requires additional HPLBs, resulting in increased interconnect area. While both HPLB2 and HPLB3 structures obtain average area optimizations over 30%, HPLB3 produces a significantly greater HPLB count and area optimization than HPLB2 as the test circuit scale grows. Thus, after replacing the CLB structure, the HPLB3 structure provides a more balanced optimization impact, greatly improving the utilization of programmable resources when taking into account the combined aspects of HPLB usage count and area optimization.
Efficient and Verifiable Ciphertext Retrieval Scheme Based on Trusted Execution Environment
WU Axin, FENG Dengguo, ZHANG Min, CHI Jialin, YI Yuling
Available online  , doi: 10.11999/JEIT251358
Abstract:
The ciphertext retrieval mechanism enables retrieval functionality over encrypted data. Symmetric Searchable Encryption (SSE) is a critical branch of ciphertext retrieval. However, due to considerations such as saving computing power, cloud servers may return incorrect or incomplete results. Moreover, attackers can also exploit these leaked information from search and access patterns to reconstruct the keyword details. Therefore, it is necessary and meaningful to protect the privacy of search and access patterns while achieving result verifiability. Nevertheless, existing verifiable SSE schemes that support search and access pattern privacy typically rely on keyword traversal mechanisms and their verification mechanisms are inefficient, which impose high computational and communication overheads on users. To address the above performance bottlenecks, this paper introduces an efficient and verifiable ciphertext retrieval scheme based on Trusted Execution Environment (TEE). To improve the efficiency of ciphertext retrieval, this scheme employs the collaborative implementation of hardware-level security isolation and oblivious data rearrangement to achieve keyword trapdoor size independent of the size of the keyword dictionary. Meanwhile, the correctness of the returned results is verified by embedding random numbers and blinding polynomial constant terms. Thanks to these designs, the scheme achieves significant efficiency improvements. Specifically, firstly, this scheme ensures that the size of keyword trapdoors depends solely on the number of query keywords, not the global dictionary size, effectively minimizing communication and computational costs. Secondly, this scheme requires storing only two random numbers to enable verifiability, substantially minimizing local storage overhead for users. Thirdly, the adoption of techniques, such as enabling data users to retrieve results via single-server and single-round interaction and leveraging symmetric homomorphic encryption, further enhances operational efficiency. Additionally, confidential computing within TEE weakens the security assumptions and trust level towards TEE. After formally proving the security of the proposed scheme using simulation-based methods, this paper has conducted a comprehensive performance evaluation. The evaluation results confirm that this scheme is significantly more efficient than other schemes with the same functionalities.
Design and Verification of Robust Modulation Recognition Framework Under Blind Adversarial Attacks
ZHENG Qinghe, ZHOU Fuhui, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT260019
Abstract:
  Objective  Deep learning-based automatic modulation recognition (AMR) models have demonstrated superior performance in non-cooperative communication systems such as cognitive radio and spectrum monitoring. However, the inherent vulnerability of deep learning models to adversarial attacks, where imperceptible perturbations can cause catastrophic misclassification, poses the severe security threat. Existing defense methods, including adversarial training, often rely on prior knowledge of specific attacks, incur significant computational overhead, and face the trade-off between robustness and accuracy on clean samples. To address these limitations, this paper aims to design and validate a robust modulation recognition framework that can operate effectively under blind adversarial attack scenarios without prior knowledge of the attack type and strategy, thereby ensuring the reliable deployment of intelligent communication systems in adversarial environments.  Methods  The proposed framework integrates a novel feature-purifying autoencoder module with standard modulation classifiers (CNN and Transformer). The core innovation lies in the autoencoder’s bottleneck layer, which incorporates a dynamic purification mechanism. This mechanism first calculates an adaptive threshold based on the statistical properties of the encoded latent features to identify anomalies. Subsequently, the Top-K sparsification operation selectively preserves only the most significant feature activations, effectively suppressing noise and adversarial perturbations while retaining essential signal characteristics. Then the autoencoder is trained via a three-stage curriculum learning strategy that sequentially optimizes reconstruction fidelity, feature sparsity, and semantic consistency between the purified and original clean signals, ensuring the output aligns with the true modulation manifold. This model-agnostic module can be seamlessly prepended to any trained classifier without retraining.  Results and Discussions  Comprehensive experiments are conducted on a simulated dataset encompassing 12 digital modulation types under multipath fading channels. The framework demonstrated substantial performance improvements. For the CNN and Transformer, the recognition accuracies under challenging targeted white-box attacks increased to 82.1% and 83.2%, and under non-targeted black-box attacks reached 87.7% and 89.4%, respectively (Table 1). The attack success rate (ASR) and attack effectiveness index (AEI) remained at low levels, confirming strong defensive capability. Figure 4 shows that defense efficacy improves with higher SNR. Crucially, the ablation study in Figure 5 highlights the indispensable role of the autoencoder, whose removal caused accuracy to plummet by 4.02% and 2.36% on CNN and Transformer under strong attacks. Further analysis (Figure 6) indicates that the framework maintains robustness across a wide range of perturbation bounds (\begin{document}$ \epsilon \leq 0.1 $\end{document}). Moreover, parameter sensitivity studies (Figures 7 and 8) show stable performance for threshold coefficient \begin{document}$ \xi $\end{document} in [1.5, 1.9] and sparsity rate k around 0.7, confirming its practical deployability.  Conclusions  This paper presents a robust, blind defense framework for robust AMR based on the feature-purifying autoencoder. The key advantages are threefold: 1) It provides effective defense against diverse white-box and black-box attacks without requiring any prior knowledge of various attack methods, achieving true blind defense; 2) As a preprocessing module, it eliminates the need for computationally expensive retraining of the primary classifier and is compatible with various backbone networks; 3) The multi-stage training strategy successfully balances robustness against attacks with the preservation of high accuracy on clean samples. Finally, experimental results on the comprehensive dataset validate the framework’s superiority. Future work will focus on lightweight architectural designs to reduce inference latency and further investigate performance boundaries under extreme low-SNR conditions combined with complex nonlinear channel impairments.
Performance Analysis and Rapid Prediction of Long-range Underwater Acoustic Communications in Uncertain Deep-sea Environments
CHEN Xiangmei, TAI Yupeng, WANG Haibin, HU Chenghao, WANG Jun, WANG Diya
Available online  , doi: 10.11999/JEIT251244
Abstract:
  Objective  In complex and dynamically changing deep-sea environments, the performance of underwater acoustic communications shows substantial variability. Feedback-based channel estimation and parameter adaptation are impractical in long-range scenarios because platform constraints prevent reliable feedback channels and the slow propagation of sound introduces significant delay. In typical long-range systems, environmental dynamics are often ignored and communication parameters are selected heuristically, which frequently leads to mismatches with actual channel conditions and causes communication failures or reduced efficiency. Predictive methods able to assess performance in advance and support feed-forward parameter adjustment are therefore required. This study proposes a deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications under uncertain environmental conditions to enable efficient and reliable parameter–channel matching without feedback.  Methods  A feed-forward method for underwater acoustic communication performance analysis and rapid prediction is developed using deep-learning-based sound-field uncertainty estimation. A neural network is first used to estimate probability distributions of Transmission Loss (TL PDFs) at the receiver under dynamic environments. TL PDFs are then mapped to probability distributions of the Signal-to-Noise Ratio (SNR PDFs), enabling communication performance evaluation without real-time feedback. Statistical channel capacity and outage capacity are analyzed to characterize the theoretical upper limits of achievable rates in dynamic conditions. Finally, by integrating the SNR distribution with the bit-error-rate characteristics of a representative deep-sea single-carrier communication system under the corresponding channel, a rate–reliability prediction model is constructed. This model estimates the probability of reliable communication at different data rates and serves as a practical tool for forecasting link performance in highly dynamic and feedback-limited underwater acoustic environments.  Results and Discussions  The method is validated using simulation data and sea trial data. The TL PDFs predicted by the deep learning model show strong consistency with the traditional Monte Carlo (MC) method across multiple receiver locations (Fig. 6). Under identical computational settings, deep-learning-based TL PDF prediction reduces computation time by 2\begin{document}$ \sim $\end{document}3 orders of magnitude compared with the MC method. The chained mapping from TL PDFs to SNR PDFs and then to channel capacity metrics accurately represents the probabilistic features of communication performance under uncertain conditions (Fig. 7 and Fig. 8). The rate–reliability curves derived from the deep-learning-based TL PDFs are highly consistent with MC-based results. In the high sound-intensity region, prediction errors for reliable communication probabilities across data rates range from 0.1% to 3%, and in the low sound-intensity region errors are approximately 0.3% to 5% (Fig. 12). Sea trial results further indicate that predicted rate–reliability performance agrees well with measured data. In the convergence zone, deviations between predicted and measured reliability probabilities at each rate range from 0.9% to 4%, and in the shadow zone from 1% to 9% (Fig. 18). Under a 90% reliability requirement, the maximum achievable rates predicted by the method match the measurements in both the convergence and shadow zones, demonstrating accuracy and practical applicability in complex channel environments.  Conclusions  A deep-learning-based framework for performance analysis and rapid prediction of long-range underwater acoustic communications in uncertain deep-sea environments is developed and validated. The framework builds a chained mapping from environmental parameters to TL PDFs, SNR PDFs, and communication performance metrics, enabling quantitative capacity assessment under dynamic ocean conditions. Predictive “rate–reliability’’ profiles are obtained by integrating probabilistic propagation characteristics with the performance of a representative deep-sea single-carrier system under the corresponding channel, providing guidance for parameter selection without feedback. Sea trial results confirm strong agreement between predicted and measured performance. The proposed approach offers a technical pathway for feed-forward performance analysis and dynamic adaptation in long-range deep-sea communication systems, and can be extended to other communication scenarios in dynamic ocean environments.
Indoor Visible Light Positioning Based on CNN–MLP Multi-Feature Fusion under Random Receiver Tilt Conditions
JIA Kejun, WANG Jian, MAO Lifei, YOU Wei, HUANG Ziyang, PENG Duo
Available online  , doi: 10.11999/JEIT251021
Abstract:
  Objective  Traditional visible light positioning (VLP) methods based on received signal strength (RSS) suffer from instability when the receiver experiences orientation perturbations, which disrupt the correspondence between optical power and spatial position, making reliable three-dimensional (3D) positioning difficult to achieve. Existing approaches typically rely on inertial measurement units (IMUs) to obtain orientation information; however, sensor fusion increases system complexity and hardware cost and introduces cumulative errors. To address these issues, this paper proposes a positioning method that fuses cosine-of-incidence-angle estimation based on a photodiode (PD) array with RSS information, enabling high-accuracy 3D indoor positioning under receiver orientation perturbations.  Methods  In the proposed fusion-based positioning method, a multi-PD array structure is first adopted, and a local coordinate system (LCS) is established at the array center. Constraint equations are then constructed based on the differences in received optical power among PDs in the array. A Gauss–Newton iterative algorithm is employed to estimate the incident light direction vector. By exploiting the orthogonal rotation invariance between the LCS and the global coordinate system (GCS), the cosine of the incident angle is estimated without the need for orientation sensors. Subsequently, a serial CNN–MLP fusion network is constructed, in which the estimated incident-angle cosine is introduced as an additional positioning feature on top of RSS-based localization. The network jointly models the RSS and incident-angle cosine information received by the PD array and maps them to 3D spatial coordinates. Finally, training samples are generated using Latin hypercube sampling (LHS) to uniformly sample spatial positions and orientation dimensions, thereby improving the representativeness of the training dataset.  Results and Discussions  Simulation experiments are conducted in a 4 m × 4 m × 2.5 m indoor environment. First, the effects of different numbers of PDs and tilt angles on the accuracy of incident-angle cosine estimation and spatial coverage are evaluated (Fig. 6), and the cumulative distribution functions (CDFs) of positioning errors under different array configurations are compared (Fig. 7). The results show that a 3-PD array with a tilt angle of 40° achieves the best balance among cost, coverage, and positioning accuracy. Next, positioning performance under different receiver tilt angles is analyzed. When the tilt angle is small, more than 70% of positioning errors are below 5 cm; even when the receiver is tilted up to 55°, the average error remains within 11.7 cm (Fig. 8). Error component comparisons indicate that the error along the Z-axis is significantly smaller than those along the X and Y axes (Fig. 9). Further tests are conducted at a height of 0.0 m covered by the training data and at an unseen height of 0.6 m not included in the training set (Fig. 10). The results demonstrate that the proposed model does not exhibit strong dependence on a specific height plane and maintains stable 3D positioning performance at unseen heights. Finally, the proposed method is compared with related positioning schemes. It outperforms existing methods in terms of CDF convergence speed, RMSE, and standard deviation (Fig. 11), achieving an average error reduction of approximately 2.5 cm and an RMSE reduction of 31.58% compared with Ref. [12].  Conclusions  This paper estimates the cosine of the incident angle at the receiver by exploiting differences in the optical power received by different PDs in an array and introduces this cosine value as a joint positioning feature into conventional RSS-based localization, thereby alleviating the instability of position mapping caused by relying solely on RSS under random receiver perturbations. By further combining the spatial feature extraction capability of CNNs with the nonlinear modeling strength of MLPs, the proposed method effectively maps positioning features to 3D spatial coordinates. The approach reduces reliance on orientation sensors such as IMUs while overcoming the susceptibility of traditional geometric positioning methods to noise and high-dimensional nonlinear features. Under varying heights and receiver orientations, the proposed algorithm demonstrates significant advantages in both positioning accuracy and stability.
Inverse Design of a Silicon-Based Compact Polarization Splitter-Rotator
HUI Zhanqiang, ZHANG Xinglong, HAN Dongdong, LI Tiantian, GONG Jiamin
Available online  , doi: 10.11999/JEIT250858
Abstract:
  Objective  The Polarization Splitter-Rotator (PSR) is a key device used to control the polarization state of light in Photonic Integrated Circuits (PICs). Device size has become a major constraint on integration density in PICs. Traditional design methods are time-consuming and tend to yield larger device footprints. Inverse design, by contrast, determines structural parameters through optimization algorithms according to target performance and enables compact devices to be obtained while maintaining functionality. This strategy is now applied to wavelength and mode division multiplexers, all-optical logic gates, power splitters, and other integrated photonic components. The objective of this work is to use inverse design to address size limitations in silicon-based PSRs by combining the Momentum Optimization algorithm with the Adjoint Method. This combined approach improves the integration level of PICs and provides a feasible pathway for the miniaturization of other photonic devices.  Methods  The design region is defined on a 220 nm Silicon-on-Insulator (SOI) wafer and is discretized into 25×50 cylindrical elements. Each element has a 50 nm radius, a 150 nm height, and an initial relative permittivity of 6.55. The adjoint method is used to obtain gradient information across the design region, and this gradient is processed with the Momentum Optimization algorithm. The relative permittivity of each element is then updated according to the processed gradient. During optimization, the momentum factor is dynamically adjusted with the iteration number to accelerate convergence, and a linear bias is applied to guide the permittivity toward the values of silicon and air as the iterations progress. After optimization, the elements are binarized based on their final permittivity: values below 6.55 are assigned to air, whereas values above 6.55 are assigned to silicon. This results in a structure containing irregularly distributed air holes. To compensate for performance loss introduced during binarization, the etching depth of air holes with pre-binarization permittivity between 3 and 6.55 is optimized. Adjacent air holes are merged to reduce fabrication errors. The final device consists of air holes with five radii, among which three larger-radius types are selected for further refinement. Their etching radii and depths are optimized to recover remaining performance loss. Device performance is evaluated through numerical analysis. Calculated parameters include Insertion Loss (IL), Crosstalk (CT), Polarization Extinction Ratio (PER), and bandwidth. Tolerance analysis is also conducted to assess robustness under fabrication variations.  Results and Discussions   A compact PSR is designed on a 220 nm SOI wafer with dimensions of 5 μm in length and 2.5 μm in width. During optimization, the momentum factor in the Momentum Optimization algorithm is dynamically adjusted. A larger momentum factor is applied in the early stage to accelerate escape from local maxima or plateau regions, whereas a smaller momentum factor is used in later iterations to increase the weight of the current gradient. Compared with other optimization strategies, this algorithm requires only 20%~33% of the iteration count needed by alternative methods to reach a Figure of Merit (FOM) of 1.7, which improves optimization efficiency. Numerical analysis shows that the device achieves stable performance across the 1 520~1 575 nm wavelength range. The IL remains low (TM0 < 1 dB, TE0 < 0.68 dB), and the CT is effectively suppressed (TM0 < –23 dB, TE0 < –25.2 dB). The PER is high (TM0 > 17 dB, TE0 > 28.5 dB). Tolerance analysis indicates strong robustness to fabrication variations. Within the 1 520~1 540 nm range, performance remains stable under etching depth offsets of ±9 nm and etching radius offsets of ±5 nm, demonstrating reliable manufacturability.  Conclusions   Numerical analysis demonstrates that combining the adjoint method with the Momentum Optimization algorithm is a feasible strategy for designing an integrated PSR. The design principle relies on controlling light propagation through adjustments to the relative permittivity, which determine the distribution and placement of air holes to achieve polarization splitting and rotation. Compared with traditional design approaches, inverse design uses the design region more efficiently and enables a more compact device structure. The proposed PSR is markedly smaller and shows enhanced fabrication tolerance. It is suitable for future large-scale PICs and provides useful guidance for the miniaturization of other photonic devices.
Conditional Generative Adversarial Networks-based Channel Estimation for ISAC-RIS System
LIU Yu, ZHENG Zelin, LIU Gang
Available online  , doi: 10.11999/JEIT251168
Abstract:
  Objective  In RIS-assisted ISAC systems, accurate channel estimation is crucial to ensure reliable operation. Although traditional deep learning methods can partially address the channel estimation problem, their generalization ability and estimation accuracy remain limited in complex multi-user channel environments. To tackle these challenges, this paper proposes a two-stage channel estimation method based on Conditional Generative Adversarial Network(CGAN) for RIS-assisted multi-user ISAC systems, aiming to enhance both the accuracy and stability of channel estimation.  Methods  This paper proposes a two-stage channel estimation method based on CGAN for estimating the SAC channels in RIS-assisted multi-user ISAC systems. By adjusting the switching states of the RIS, the overall estimation problem is decomposed into subproblems, enabling sequential estimation of the direct and reflected channels. Within the proposed CGAN framework, the adversarial training between the generator and discriminator allows the model not only to learn the mapping relationship between the observed signals and the true channels but also to optimize the output according to the discriminator’s feedback, thereby effectively improving both training efficiency and estimation accuracy.  Results and Discussions  Extensive simulation experiments were conducted to verify the effectiveness of the proposed method. First, the estimation performance of the SAC channel under different SNR conditions was compared. The results demonstrate that the proposed CGAN-based method achieves significantly better NMSE performance than the LS benchmark and traditional models such as FNN and ELM (Fig. 4). Then, the impact of increasing the number of antennas and RIS elements on SAC channel estimation performance was investigated. Compared with the LS benchmark, the proposed CGAN method consistently maintains superior performance under various SNR conditions (Figs. 5 and 6).  Conclusions  This paper investigates the channel estimation problem in RIS-assisted multi-user ISAC systems and proposes a two-stage channel estimation method based on CGAN. By adjusting the switching states of the RIS and employing adversarial training between the generator and discriminator networks, the proposed method achieves accurate estimation of the SAC channel. Simulation results demonstrate that, under various SNR conditions and channel dimensions, the CGAN-based estimation method exhibits strong generalization capability and significantly outperforms the benchmark schemes in estimation accuracy. Therefore, it shows great potential as an effective solution for enhancing system stability and efficiency.
Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning in Wireless Networks
LIU Jingyuan, MA Ke, XU Runchen, CHANG Zheng
Available online  , doi: 10.11999/JEIT251221
Abstract:
  Objective  Multimodal Federated Learning (MFL) uses complementary information from multiple modalities, yet in wireless edge networks it is restricted by limited energy and frequent missing modalities because many clients store only images or only reports. This study presents Cross-modal Retrieval Enhanced Energy-efficient Multimodal Federated Learning (CREEMFL), which applies selective completion and joint communication–computation optimization to reduce training energy under latency and wireless constraints.  Methods  CREEMFL completes part of the incomplete samples by querying a public multimodal subset, and processes the remaining samples through zero padding. Each selected user downloads the global model, performs image-to-text or text-to-image retrieval, conducts local multimodal training, and uploads model updates for aggregation. An energy–delay model couples local computation and wireless communication and treats the required number of global rounds as a function of retrieval ratios. Based on this model, an energy minimization problem is formulated and solved using a two-layer algorithm with an outer search over retrieval ratios and an inner optimization of transmission time, Central Processing Unit (CPU) frequency, and transmit power.  Results and Discussions  Simulations on a single-cell wireless MFL system show that increasing the ratio of completing text from images improves test accuracy and reduces total energy. In contrast, a large ratio of completing images from text provides limited accuracy gain but increases energy consumption (Fig. 3, Fig. 4). Compared with four representative baselines, CREEMFL achieves shorter completion time and lower total energy across a wide range of maximum average transmit powers (Fig. 5, Fig. 6). For CREEMFL, increased system bandwidth further reduces completion time and energy consumption (Fig. 7, Fig. 8). Under different user modality compositions, CREEMFL also attains higher test accuracy than local training, zero padding, and cross-modal retrieval without energy optimization (Fig. 9).  Conclusions  CREEMFL integrates selective cross-modal retrieval and joint communication–computation optimization for energy-efficient MFL. By treating retrieval ratios as variables and modeling their effect on global convergence rounds, it captures the coupling between per-round costs and global training progress. Simulations verify that CREEMFL reduces training completion time and total energy while preserving classification accuracy in resource-constrained wireless edge networks.
Finite-time Adaptive Sliding Mode Control of Servo Motors Considering Frictional Nonlinearity and Unknown Loads
ZHANG Tianyu, GUO Qinxia, YANG Tingkai, GUO Xiangji, MING Ming
Available online  , doi: 10.11999/JEIT250521
Abstract:
  Objective  Ultra-fast laser processing with an infinite field of view requires servo motor systems with superior tracking accuracy and robustness. However, such systems are highly nonlinear and affected by coupled unknown load disturbances and complex friction, which constrain the performance of conventional controllers. Although Sliding Mode Control (SMC) exhibits inherent robustness, traditional SMC and observer designs cannot achieve accurate finite-time disturbance compensation under strong nonlinearities, thus limiting high-speed and high-precision trajectory tracking. To address this limitation, a novel finite-time adaptive SMC approach is proposed to ensure rapid and precise angular position tracking within a finite time, satisfying the stringent synchronization requirements of advanced laser processing systems.  Methods  A novel control strategy is developed by integrating an adaptive disturbance observer fused with a Radial Basis Function Neural Network (RBFNN) and finite-time Sliding Mode Control (SMC). First, the unknown load disturbance and complex frictional nonlinear dynamics are combined into a unified "lumped disturbance" term, improving model generality and the ability to represent real operating conditions. Second, a finite-time adaptive disturbance observer is constructed to estimate this lumped disturbance. The observer utilizes the universal approximation capability of the RBFNN to learn and approximate the dynamic characteristics of unknown disturbances online. Simultaneously, a finite-time adaptive law based on the error norm is introduced to update the neural network weights in real time, ensuring rapid and accurate finite-time estimation of the lumped disturbance while reducing dependence on precise model parameters. Based on this design, a finite-time SMC is developed. The controller uses the observer’s disturbance estimation as a feedforward compensation term, incorporates a carefully formulated finite-time sliding surface and equivalent control law, and introduces a saturation function to suppress control input chattering. A suitable Lyapunov function is then constructed, and the finite-time stability theory is rigorously applied to prove the practical finite-time convergence of both the adaptive observer and the closed-loop control system, guaranteeing that the system tracking error converges to a bounded neighborhood near the origin within finite time.  Results and Discussions  To verify the effectiveness and superiority of the proposed control strategy, a typical Permanent Magnet Synchronous Motor (PMSM) servo system model is constructed in the MATLAB environment, and a simulation scenario with desired trajectories of varying frequencies is established. The proposed method is comprehensively compared with the widely used Proportional–Integral (PI) control and the advanced method reported in reference [7]. Simulation results demonstrate the following: 1. Tracking performance: Under various reference trajectories, the proposed controller enables the system to accurately follow the target trajectory with a tracking error substantially smaller than that of the PI controller. Compared with the method in reference [7], it achieves smoother responses and smaller residual errors, effectively eliminating the chattering observed in some operating conditions of the latter. 2 Disturbance rejection and robustness: The adaptive disturbance observer based on the RBFNN rapidly and effectively learns and compensates for the lumped disturbance composed of unknown load variations and frictional nonlinearities. Even in the presence of these disturbances, the proposed controller maintains high-precision trajectory tracking, demonstrating strong disturbance rejection and robustness to system parameter variations. 3. Control input characteristics: Compared with the reference methods, the control signal of the proposed approach quickly stabilizes after the initial transient phase, effectively suppressing chattering caused by high-frequency switching. The amplitude range of the control input remains reasonable, facilitating practical actuator implementation. 4. Comprehensive evaluation: Based on multiple error performance indices, including Integral Squared Error (ISE), Integral Absolute Error (IAE), Time-weighted Integral Absolute Error (ITAE), and Time-weighted Integral Squared Error (ITSE), the proposed controller consistently outperforms both PI control and the method in reference [7]. It demonstrates comprehensive advantages in suppressing transient errors rapidly and reducing overall error accumulation. The method also improves steady-state accuracy and achieves a balanced response speed with effective noise attenuation. 5. Observer performance: The RBFNN weight norm estimation converges rapidly and stabilizes at a low level after initial adaptation, confirming the effectiveness of the proposed adaptive law and the learning efficiency of the observer.  Conclusions  A finite-time sliding mode control strategy with an adaptive disturbance observer is proposed for servo systems used in ultra-fast laser processing. The method models unknown load disturbances and frictional nonlinearities as a lumped disturbance term. An adaptive observer, integrating an RBF neural network with a finite-time mechanism, accurately estimates this disturbance for real-time compensation. Based on the observer, a finite-time SMC law is formulated, and the practical finite-time stability of the closed-loop system is theoretically proven. Simulations conducted on a permanent magnet synchronous motor platform confirm that the proposed approach achieves superior tracking accuracy, robustness, and control smoothness compared with conventional PI and existing advanced methods. This work offers an effective solution for achieving high-precision control in nonlinear systems subject to strong disturbances.
Breakthrough in Solving NP-Complete Problems Using Electronic Probe Computers
XU Jin, YU Le, YANG Huihui, JI Siyuan, ZHANG Yu, YANG Anqi, LI Quanyou, LI Haisheng, ZHU Enqiang, SHI Xiaolong, WU Pu, SHAO Zehui, LENG Huang, LIU Xiaoqing
Available online  , doi: 10.11999/JEIT250352
Abstract:
This study presents a breakthrough in addressing NP-complete problems using a newly developed Electronic Probe Computer (EPC60). The system employs a hybrid serial–parallel computational model and performs large-scale parallel operations through seven probe operators. In benchmark tests on 3-coloring problems in graphs with 2,000 vertices, EPC60 achieves 100% accuracy, outperforming the mainstream solver Gurobi, which succeeds in only 6% of cases. Computation time is reduced from 15 days to 54 seconds. The system demonstrates high scalability and offers a general-purpose solution for complex optimization problems in areas such as supply chain management, finance, and telecommunications.  Objective   NP-complete problems pose a fundamental challenge in computer science. As problem size increases, the required computational effort grows exponentially, making it infeasible for traditional electronic computers to provide timely solutions. Alternative computational models have been proposed, with biological approaches—particularly DNA computing—demonstrating notable theoretical advances. However, DNA computing systems continue to face major limitations in practical implementation.  Methods  Computational Model: EPC is based on a non-Turing computational model in which data are multidimensional and processed in parallel. Its database comprises four types of graphs, and the probe library includes seven operators, each designed for specific graph operations. By executing parallel probe operations, EPC efficiently addresses NP-complete problems.Structural Features:EPC consists of four subsystems: a conversion system, input system, computation system, and output system. The conversion system transforms the target problem into a graph coloring problem; the input system allocates tasks to the computation system; the computation system performs parallel operations via probe computation cards; and the output system maps the solution back to the original problem format.EPC60 features a three-tier hierarchical hardware architecture comprising a control layer, optical routing layer, and probe computation layer. The control layer manages data conversion, format transformation, and task scheduling. The optical routing layer supports high-throughput data transmission, while the probe computation layer conducts large-scale parallel operations using probe computation cards.  Results and Discussions  EPC60 successfully solved 100 instances of the 3-coloring problem for graphs with 2,000 vertices, achieving a 100% success rate. In comparison, the mainstream solver Gurobi succeeded in only 6% of cases. Additionally, EPC60 rapidly solved two 3-coloring problems for graphs with 1,500 and 2,000 vertices, which Gurobi failed to resolve after 15 days of continuous computation on a high-performance workstation.Using an open-source dataset, we identified 1,000 3-colorable graphs with 1,000 vertices and 100 3-colorable graphs with 2,000 vertices. These correspond to theoretical complexities of O(1.3289n) for both cases. The test results are summarized in Table 1.Currently, EPC60 can directly solve 3-coloring problems for graphs with up to n vertices, with theoretical complexity of at least O(1.3289n).On April 15, 2023, a scientific and technological achievement appraisal meeting organized by the Chinese Institute of Electronics was held at Beijing Technology and Business University. A panel of ten senior experts conducted a comprehensive technical evaluation and Q&A session. The committee reached the following unanimous conclusions:1. The probe computer represents an original breakthrough in computational models.2. The system architecture design demonstrates significant innovation.3. The technical complexity reaches internationally leading levels.4. It provides a novel approach to solving NP-complete problems.Experts at the appraisal meeting stated, “This is a major breakthrough in computational science achieved by our country, with not only theoretical value but also broad application prospects.” In cybersecurity, EPC60 has also demonstrated remarkable potential. Supported by the National Key R&D Program of China (2019YFA0706400), Professor Xu Jin’s team developed an automated binary vulnerability mining system based on a function call graph model. Evaluation of the system using the Modbus Slave software showed over 95% vulnerability coverage, far exceeding the 75 vulnerabilities detected by conventional depth-first search algorithms. The system also discovered a previously unknown flaw, the “Unauthorized Access Vulnerability in Changyuan Shenrui PRS-7910 Data Gateway” (CNVD-2020-31406), highlighting EPC60’s efficacy in cybersecurity applications.The high efficiency of EPC60 derives from its unique computational model and hardware architecture. Given that all NP-complete problems can be polynomially reduced to one another, EPC60 provides a general-purpose solution framework. It is therefore expected to be applicable in a wide range of domains, including supply chain management, financial services, telecommunications, energy, and manufacturing.  Conclusions   The successful development of EPC offers a novel approach to solving NP-complete problems. As technological capabilities continue to evolve, EPC is expected to demonstrate strong computational performance across a broader range of application domains. Its distinctive computational model and hardware architecture also provide important insights for the design of next-generation computing systems.
Personalized Federated Learning Method Based on Collation Game and Knowledge Distillation
SUN Yanhua, SHI Yahui, LI Meng, YANG Ruizhe, SI Pengbo
Available online  , doi: 10.11999/JEIT221203
Abstract:
To overcome the limitation of the Federated Learning (FL) when the data and model of each client are all heterogenous and improve the accuracy, a personalized Federated learning algorithm with Collation game and Knowledge distillation (pFedCK) is proposed. Firstly, each client uploads its soft-predict on public dataset and download the most correlative of the k soft-predict. Then, this method apply the shapley value from collation game to measure the multi-wise influences among clients and quantify their marginal contribution to others on personalized learning performance. Lastly, each client identify it’s optimal coalition and then distill the knowledge to local model and train on private dataset. The results show that compared with the state-of-the-art algorithm, this approach can achieve superior personalized accuracy and can improve by about 10%.
The Range-angle Estimation of Target Based on Time-invariant and Spot Beam Optimization
Wei CHU, Yunqing LIU, Wenyug LIU, Xiaolong LI
Available online  , doi: 10.11999/JEIT210265
Abstract:
The application of Frequency Diverse Array and Multiple Input Multiple Output (FDA-MIMO) radar to achieve range-angle estimation of target has attracted more and more attention. The FDA can simultaneously obtain the degree of freedom of transmitting beam pattern in angle and range. However, its performance is degraded due to the periodicity and time-varying of the beam pattern. Therefore, an improved Estimating Signal Parameter via Rotational Invariance Techniques (ESPRIT) algorithm to estimate the target’s parameters based on a new waveform synthesis model of the Time Modulation and Range Compensation FDA-MIMO (TMRC-FDA-MIMO) radar is proposed. Finally, the proposed method is compared with identical frequency increment FDA-MIMO radar system, logarithmically increased frequency offset FDA-MIMO radar system and MUltiple SIgnal Classification (MUSIC) algorithm through the Cramer Rao lower bound and root mean square error of range and angle estimation, and the excellent performance of the proposed method is verified.
Image and Intelligent Information Processing
Spatial Information-guided Diffusion for Domain Adaptation Semantic Segmentation of Remote Sensing Images
LIANG Yan, LI Jun-Fan, SHAO Kai, HU Lin
Available online  , doi: 10.11999/JEIT260031
Abstract:
  Objective  Domain Adaptation Semantic Segmentation (DASS) is critical for remote sensing applications, including land-cover mapping, urban planning, and environmental monitoring. However, deep learning models often show severe performance degradation under domain shifts caused by imaging variation, geographic differences, and label-semantic heterogeneity. Conventional feature-alignment and generative adversarial network-based methods often fail to preserve semantic consistency. They are also sensitive to noisy supervision, especially when cross-domain gaps are large. This work aims to construct a robust DASS framework for semantically consistent image translation and reliable knowledge transfer.  Methods  A two-stage framework, termed Co-training Spatial-Guided DASS (CoSG-DASS), is proposed by integrating image translation and co-training. In the image-translation stage, a spatial information-guided latent diffusion model enhanced by ControlNet is designed. Semantic pseudo-labels and depth estimates are used as horizontal semantic and vertical spatial conditions to guide target-style image generation. To reduce the effect of noisy pseudo-labels, an Entropy-based Adaptive Guidance Intensity Module (EAGIM) is introduced. EAGIM estimates pixel-level confidence using information entropy and suppresses unreliable features. In the co-training stage, translated target-style images and unlabeled real target-domain images are used to train a segmentation model with a depth-guided segmentation head. Cross-entropy loss and adversarial loss are jointly used for optimization.  Results and Discussions  Extensive experiments are conducted on three cross-domain tasks. CoSG-DASS generates images that better match target-domain distributions. Quantitative results based on Fréchet Inception Distance (FID) show that the proposed method outperforms CycleGAN, UNI-Diff, and CRS-Diff in most settings (Table 1). Visual comparisons (Fig. 6) show that the method reduces edge blurring and category confusion. It also improves the separation of roads and vegetation and preserves small objects, such as vehicles. In the semantic segmentation stage, CoSG-DASS outperforms state-of-the-art domain adaptation methods. It improves mean Intersection over Union (mIoU) by 1.14%, 3.78%, and 2.49% on the cross-geographic task (Vaihingen IRRG→Potsdam IRRG), cross-imaging-mode task (Vaihingen IRRG→Potsdam RGB), and bidirectional label-semantic-heterogeneity tasks between DFC25 and LoveDA, respectively (Tables 24). Visual segmentation results (Fig. 7) confirm its strong boundary preservation and high accuracy in complex scenes. Ablation studies (Table 5) verify the contribution of the core components, including depth control, pseudo-label guidance, EAGIM, and the co-training strategy. Feature-distribution visualization based on Uniform Manifold Approximation and Projection (UMAP) further shows that CoSG-DASS reduces intra-class variation and increases inter-class separation after adaptation (Fig. 8).  Conclusions  CoSG-DASS alleviates domain shifts in remote sensing images through semantic-preserving diffusion-based translation and depth-guided co-training. It improves both image-translation quality and segmentation accuracy over existing methods. The proposed framework provides an effective solution for multi-source remote sensing interpretation. Future work will focus on extreme label-semantic heterogeneity and lightweight diffusion architectures.
Multi-dimensional Spatio-temporal Feature Enhancement for Lip Reading
MA Jinlin, ZHONG Yaowei, MA Ruishi
Available online  , doi: 10.11999/JEIT251111
Abstract:
  Objective  Lip reading is a challenging yet important task in computer vision that aims to decode spoken language solely from visual lip movements. The task is difficult mainly because of inherent ambiguity in visual speech signals. On one hand, articulatory movements for different visemes can be highly subtle. For instance, lip displacement differences for confusable pairs such as /p/–/b/ and /m/–/n/ may be as small as 0.3~0.7 mm. These fine-grained spatial variations often fall below the effective resolution limits of conventional 3D convolutional neural networks. On the other hand, natural coarticulation in speech introduces temporal ambiguity, as mouth shapes may transiently blend multiple phonemes and make distinct visual units difficult to separate. These challenges are further aggravated by real-world factors such as uneven lighting and substantial inter-speaker differences in articulation. Current lip-reading models therefore often show limited ability to capture discriminative spatiotemporal features, which leads to suboptimal performance, particularly for phonemes with minimal visual differences. To address these issues, a robust lip-reading framework is developed to capture and exploit fine-grained spatiotemporal dependencies and improve recognition accuracy under diverse and realistic conditions.  Methods  To address the above limitations, a novel lip-reading framework, termed the Multi-dimensional Spatio-Temporal Enhancement Network (MSTEN), is proposed. The framework is designed to strengthen spatial and temporal representations through integrated attention mechanisms and advanced residual learning. It contains three core components that collaboratively model dependencies between spatial and temporal features, which are often insufficiently used in conventional architectures. The first component, the Self-adjusting Spatio-temporal Attention (SaSTA) module, adopts a self-adjusting mechanism that operates simultaneously across the height, width, and temporal dimensions. Query, key, and value tensors are generated through 1×1×1 3D convolutions, flattened across spatial and temporal dimensions, and used to compute attention weights by multiplying the query tensor with the transposed key tensor, followed by softmax normalization. The resulting attention map is multiplied by the value tensor and then combined with the original input through learnable parameters and a residual connection, thereby preserving contextual information and producing globally enhanced features. The second component, the Three-dimensional Enhanced Residual Block (TE-ResBlock), improves spatiotemporal feature extraction through temporal shift, multi-scale convolution, and channel shuffle. The temporal shift operation moves one quarter of the feature channels along the time axis to fuse information from adjacent frames without adding parameters. Multi-scale convolution adopts parallel branches with kernel sizes of 3×3, 3×1, 1×3, and 1×1 to capture features at different receptive fields. The outputs are concatenated and processed by channel shuffle to improve information exchange across groups, and four TE-ResBlocks are stacked to achieve progressive feature refinement. The third component, the Multi-dimensional Adaptive Fusion (MDAF) module, integrates spatial, temporal, and channel information through three submodules. These are a Channel Enhancement Module (CEM), which recalibrates features through max pooling, temporal convolution, and sigmoid activation, a Spatial Enhancement Module (SEM), which expands the receptive field through identity mapping and both standard and dilated convolution, and an Adaptive Temporal Capture Module (ATCM), which emphasizes dynamic movements through frame-difference features and temporal weight maps. The MDAF modules are inserted between TE-ResBlock stacks for iterative refinement. Finally, features extracted by the MSTEN front end are fed into a Densely Connected Temporal Convolutional Network (DC-TCN) back end, which consists of four blocks, each containing three temporal convolutional layers with dense connections, to model long-range phonological dependencies effectively.  Results and Discussions  The proposed framework is evaluated comprehensively on the widely used LRW and GRID datasets. The LRW dataset contains more than 500,000 video clips from over 1,000 speakers. The GRID dataset contains video clips from 34 speakers, each providing 1,000 utterances, with a total duration of 28 h. The proposed model achieves an accuracy of 91.18%, which is an absolute improvement of 2.82 percentage points over a strong ResNet18 baseline, demonstrating its substantial effectiveness. Ablation studies are further conducted to analyze the contribution of each key component. The results show clearly that each proposed module yields a meaningful performance gain. Specifically, the SaSTA module alone improves accuracy by 2.09%, which confirms the key role of global spatiotemporal attention. The TE-ResBlock increases accuracy by 1.73%, which verifies its effectiveness in multi-scale local feature extraction and inter-frame information fusion. The MDAF module provides a further 1.74% improvement, which highlights the benefit of adaptive multi-dimensional feature fusion, as shown in Table 2.  Conclusions  This study advances lip reading through the proposed MSTEN front-end network. The framework is built on three main contributions. First, the SaSTA module proposes an effective mechanism for global context aggregation and performs multi-dimensional feature weighting across height, width, and temporal sequences. Second, the TE-ResBlock addresses central challenges in spatiotemporal modeling through the combined use of temporal displacement, multi-scale convolution, and enhanced channel interaction. Third, the MDAF module enables deep and coordinated integration of spatial, temporal, and channel information. Together, these components improve model performance substantially and achieve accuracies of 91.18% on the challenging LRW dataset and 97.82% on the GRID dataset. Ablation studies further confirm the individual and combined effectiveness of the proposed components. Future work will examine the extension of this framework to audio-visual speech recognition under noisy conditions and the development of domain adaptation strategies to improve robustness in low-resolution or resource-constrained scenarios.
Drug Response Prediction Based on Graph Topology Attention Network
XU Peng, XU Hao, BAO Zhenshen, ZHOU Chi, LIU Wenbin
Available online  , doi: 10.11999/JEIT251099
Abstract:
  Objective  A central goal in modern cancer research is to determine why patients respond differently to the same therapy. This requires computational tools that combine genetic information with drug properties to predict treatment results, which is essential for advancing personalized oncology. Although existing methods have improved cancer drug response prediction, effective drug feature extraction and integration of multi-omics data from cell lines remain challenging. To address these issues, Graph Neural Networks (GNNs) have been increasingly used to process drug molecular graphs. In this study, a model based on a graph topology attention network is proposed to extract features from drug molecular graphs, and an attention mechanism is used to integrate multi-omics data.  Methods  In this study, a drug response prediction method based on Graph Topology Attention Network (GTAT) is proposed. The model integrates topological graph information to predict drug responses in cell lines. Drug SMILES strings are used to generate two different drug representations, and multi-omics data are incorporated to characterize cell lines (Fig. 1). For drug feature extraction, SMILES strings are first parsed to construct molecular graphs, which are then processed by GTAT. This network captures both topological information at the molecular graph level and atom-level features, thereby generating structured molecular representations. At the same time, Extended Connectivity Fingerprints are computed from the same SMILES strings and transformed into continuous feature vectors through a Multi-Layer Perceptron (MLP). The graph-based drug representation and the fingerprint-based representation are then concatenated to form a comprehensive drug feature vector. For cell line representation, multi-omics data are processed through omics-specific neural networks. The resulting features are fused through multi-head self-attention mechanisms, which enable the model to capture contextual interactions across omics modalities and generate an integrated cell line representation. Finally, the drug and cell line features are combined and fed into an MLP classifier to predict drug response results. The proposed model effectively integrates heterogeneous biological data sources and significantly improves prediction accuracy through multimodal learning and attention-based feature fusion.  Results and Discussions  The proposed method achieves competitive performance on both the GDSC and CCLE benchmark datasets (Table 2). Specifically, on the GDSC dataset, the proposed approach outperforms all competing methods across all four metrics, including AUC, AUPR, F1-score, and Accuracy. In particular, the AUPR is improved by approximately 1.92% compared with that of the second-best method, MOFGCN, which demonstrates an advantage in handling class imbalance. On the CCLE dataset, the proposed method still achieves the best performance in terms of AUC and Accuracy. Although its AUPR and F1-score are slightly lower than those of GADRP, the differences are minimal, and the method shows stronger overall discriminative ability, as reflected by AUC. These results validate the effectiveness and strong generalizability of the proposed method in drug sensitivity prediction tasks. The variation in AUPR and F1-score across datasets may be attributed to inherent differences in sample size and class distribution. The limited size of the CCLE dataset, combined with its specific class imbalance, with an approximately 4:1 ratio of resistant to sensitive samples, may restrict the model’s ability to fully learn the underlying data distribution, particularly for minority classes. In contrast, the GDSC dataset shows greater heterogeneity and a more pronounced class imbalance, approximately 8:1, which increases prediction difficulty and leads to lower performance on certain metrics.  Conclusions  Accurate prediction of drug response in cell lines remains a central challenge in precision medicine and has important implications for accelerating drug development and advancing personalized treatment. However, construction of a highly accurate predictive model that effectively integrates multi-source biological information remains difficult because of the complexity of drug molecular structures and the inherent heterogeneity of cell lines. To address this issue, a cell line drug response prediction model based on GTAT is proposed. In this model, GTAT is used to extract molecular graph features of drugs, which are then fused with molecular fingerprint features. Meanwhile, multi-omics features of cell lines are integrated through an attention mechanism. Experimental results demonstrate that the proposed model achieves superior performance compared with existing state-of-the-art benchmark methods on the employed datasets. This study provides a new perspective for cell line drug response prediction. Certain limitations remain, including the use of only three types of omics features for cell line representation and the effect of sample size on predictive performance. Future work will focus on integrating more diverse omics features, applying pre-trained large-scale models, and promoting clinical translation for personalized medicine.
A High-Performance Eye Tracking Method Based on Event Camera and Dual-Channel Differential Illumination
SONG Sishun, FENG Junchi, PU Chengyu, GUO Yu, LIU Shijie, HE Xin, CHENG Yuwei
Available online  , doi: 10.11999/JEIT251162
Abstract:
  Objective  Eye tracking has become an essential technology in human-computer interaction, medical diagnostics, cognitive neuroscience, and augmented and virtual reality applications. Traditional eye tracking systems, however, often suffer from two major limitations: low spatial accuracy and limited temporal resolution, especially during high-speed eye movements. These limitations hinder precise gaze estimation and reduce the reliability of real-time interactive systems. To address these challenges, an event camera is integrated with a dual-channel differential illumination strategy to improve the signal-to-noise ratio of corneal reflection events. The Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is introduced to achieve accurate localization of corneal reflection points. On this basis, the coordinates of the corneal reflection points are combined with Singular Value Decomposition (SVD) and the least-squares method to determine the center of corneal curvature, thereby significantly improving gaze direction estimation accuracy. This study provides a practical technical route for next-generation eye tracking systems and offers theoretical support for their use in complex interactive environments.  Methods  An event-camera-based gaze tracking method is proposed that integrates asynchronous eye-movement event data through a dual-channel differential illumination framework, thereby improving gaze direction estimation accuracy under high-speed and dynamic conditions. First, the event camera asynchronously captures brightness-change events with microsecond-level temporal resolution, which enables precise tracking of rapid eye movements. At the same time, the dual-channel differential illumination mechanism suppresses redundant reflections and improves the contrast of corneal reflection points. Second, the DBSCAN algorithm is used to process the event data, effectively removing noise and improving the spatial localization accuracy of corneal reflection features. Finally, a ray-tracing model is reconstructed using SVD and least-squares fitting to determine the center of corneal curvature, thereby enabling robust and high-precision gaze direction estimation. Experimental results on a biomimetic eye-movement dataset show that the proposed method achieves high temporal resolution, localization accuracy, and robustness in dynamic tracking scenarios.  Results and Discussions  Experiments show that the proposed method achieves a temporal resolution of 25 kHz (Fig. 6), which far exceeds that of conventional cameras. Differential illumination significantly improves the signal-to-noise ratio of corneal reflection events. The DBSCAN algorithm localizes corneal reflection points more efficiently than K-means, agglomerative clustering, mean shift, and OPTICS, and achieves accurate results within 10 ms without requiring predefined clusters (Fig. 8, Table 3). For gaze estimation, the proposed method maintains stable accuracy across sampling frequencies from 2 kHz to 25 kHz. At a 15° cone angle, the Mean Error (ME) and Root Mean Square Error (RMSE) are approximately 0.66° and 0.67°, respectively. At 25°, these values increase slightly to 0.87° and 0.90°, respectively (Table 4). Compared with existing state-of-the-art gaze tracking methods, the proposed approach shows better overall performance in both temporal resolution and accuracy (Table 5). Trajectory results (Fig. 9) show close agreement between the estimated and ground-truth gaze paths, and distribution analyses (Fig. 10) confirm that most errors remain below 1°.  Conclusions  A novel eye tracking method is presented that integrates an event camera with dual-channel differential illumination. The method achieves high temporal resolution (25 kHz), improves event signal quality, and reduces localization errors, yielding gaze estimation errors of less than 1°. The proposed approach provides a reliable technical route for next-generation high-performance eye tracking systems. Future work should address sensor noise modeling and computational optimization to further improve real-world applicability.
Cryption and Network Information Security
A Lightweight and High-Reliability Challenge Generation Strategy for APUF
LAN Guohao, ZHANG Hui, DUO Bin, WANG Zibin, ZHOU Rang, LI Dongfen
Available online  , doi: 10.11999/JEIT251073
Abstract:
  Objective  The Arbiter Physical Unclonable Function (APUF) is a lightweight security primitive widely used for identity authentication and key generation in resource-constrained devices. However, its response consistency is highly sensitive to environmental perturbations. The same challenge may therefore produce inconsistent responses under different conditions, which reduces the reliability of APUF-based security systems. Existing reliability improvement schemes mainly rely on hardware modification or challenge screening. These schemes often require high resource overhead and have low efficiency. To address these limitations, a Delay-constrained Challenge Generation Strategy (DCGS) is proposed to improve APUF reliability without additional hardware overhead or inefficient candidate screening.  Methods  DCGS models APUF path-delay characteristics and constructs challenges with constrained delay differences to ensure response stability. First, a Logistic Regression (LR) model is established to characterize the relationship between challenge bits and path delays. A delay-weight vector is then derived from the trained LR model to quantify the contribution of each challenge bit to the overall path delay. Second, a two-stage challenge generation mechanism is designed for delay-constraint control. In the first stage, prefix-bit initialization generates different prefix sequences to establish a delay baseline for subsequent bitwise extension. In the second stage, bitwise extension dynamically determines each remaining challenge bit according to the delay-weight vector. During this process, the cumulative delay difference of each challenge is monitored in real time and maintained within a preset delay-difference threshold range. Unlike conventional screening methods that post-process candidate challenges, DCGS directly generates stable challenges by design. This design removes the need for candidate challenge pools and improves generation efficiency.  Results and Discussions  DCGS is evaluated under different noise intensities. At a noise intensity of 0.3, which represents the maximum practical noise level, the reliability of DCGS-generated challenges remains 100% (Fig. 2). For generation efficiency, DCGS requires only 0.017 s to generate 10,000 challenges (Table 4). The response uniformity reaches 50.02% (Table 4), and the uniqueness reaches 50.46% (Table 4). Both metrics are close to the ideal theoretical value of 50%. The security analysis shows that the average bit entropy of DCGS-generated challenges is 0.980 7 (Fig. 3). The conditional entropy is 0.987 8, only 0.002 3 lower than that of random challenges (0.990 1).  Conclusions  This paper proposes DCGS for APUF to address inconsistent responses, low generation efficiency, and high hardware resource consumption in traditional schemes under high-noise conditions. By modeling path-delay characteristics with LR and combining prefix-bit initialization with bitwise extension, the proposed strategy ensures that the generated challenges satisfy the preset delay-difference threshold range. DCGS achieves high reliability, high efficiency, and good response uniformity without increasing hardware overhead. Experimental results show that DCGS improves APUF reliability in complex environments and supports secure applications in resource-constrained devices.
Wireless Communication and Internet of Things
Modulation Recognition Method for High-Speed Mobile Communication Based on Attention Dynamic Fusion and Hybrid Pruning Transformer
ZHENG Qinghe, CHEN Bin, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
Available online  , doi: 10.11999/JEIT251211
Abstract:
  Objective  Automatic modulation recognition is a critical preprocessing step in dynamic spectrum access and anti-jamming communication systems. It directly affects the robustness and spectrum efficiency of noncooperative communication. In high-speed mobile communication scenarios, such as satellite communication, high-speed rail communication, and drone swarm communication, signal modulation features experience severe distortion due to Doppler shifts, time-varying channels, and non-stationary interference. These factors challenge traditional modulation recognition methods that rely on static assumptions, leading to feature mismatch and higher misclassification rates. To address the limited robustness and real-time performance of existing deep learning–based modulation recognition models in high-speed mobile environments, a lightweight dynamic fusion Transformer-based method is proposed.  Methods  The proposed method contains three components: a signal representation fusion block, a Transformer architecture, and a model pruning strategy for lightweight inference. First, a RollingQ mechanism dynamically adjusts the direction of the attention query matrix according to the quality of each signal representation. This design prevents attention fixation and enables balanced use of multiple signal representations. Next, a Multi-head Attention Frequency Enhancement Transformer (MAFE-Transformer) is designed. The model integrates local and global spatiotemporal features through lightweight convolutional enhancement, multi-attention feature extraction, and frequency learning and selection modules. Finally, an attention-based dynamic hybrid pruning strategy removes structural redundancy and accelerates inference, which supports real-time modulation recognition.  Results and Discussions  Experiments are conducted on two public datasets, RadioML 2016.10a and RML22, to evaluate the proposed method. The MAFE-Transformer achieves average classification accuracies of 65.14% and 78.40% on the two datasets. Under low Signal-to-Noise Ratio (SNR) conditions of –20~0 dB, the model maintains strong robustness, particularly on the RML22 dataset with the dynamic channel model ETU70 (Fig. 5). The confusion matrix indicates that classification errors are relatively evenly distributed across different modulation schemes, which reflects balanced classification performance (Fig. 6). Ablation experiments show that the RollingQ-based dynamic fusion mechanism improves accuracy by 7.2% on RadioML 2016.10a and 9.5% on RML22 compared with a single signal representation (Fig. 7). The hybrid pruning strategy reduces inference latency to 2.2 ms per signal while maintaining high accuracy (Fig. 8). Comparative experiments indicate that the proposed model outperforms several advanced deep learning models, including Ms-RaT, MobileViT, MobileRaT, and KA-CNN, by 4%~10% in recognition accuracy. These results indicate strong performance in high-speed mobile communication scenarios (Fig. 9).  Conclusions  A lightweight dynamic fusion Transformer-based automatic modulation recognition method is proposed for high-speed mobile communication environments. The RollingQ mechanism and the MAFE-Transformer architecture, combined with a dynamic hybrid pruning strategy, improve the balance between recognition accuracy and inference efficiency. Experimental results on public datasets confirm the effectiveness and robustness of the method under complex channel conditions with Doppler shifts and time-varying interference. However, the method has not been systematically evaluated under more complex interference conditions, such as impulsive noise or frequency-selective fading. Future work will examine adaptability to non-stationary noise, cross-device generalization, and optimization for edge deployment.
Satellite Navigation
Research on GRI Combination Design of eLORAN System
LIU Shiyao, ZHANG Shougang, HUA Yu
Available online  , doi: 10.11999/JEIT201066
Abstract:
To solve the problem of Group Repetition Interval (GRI) selection in the construction of the enhanced LORAN (eLORAN) system supplementary transmission station, a screening algorithm based on cross interference rate is proposed mainly from the mathematical point of view. Firstly, this method considers the requirement of second information, and on this basis, conducts a first screening by comparing the mutual Cross Rate Interference (CRI) with the adjacent Loran-C stations in the neighboring countries. Secondly, a second screening is conducted through permutation and pairwise comparison. Finally, the optimal GRI combination scheme is given by considering the requirements of data rate and system specification. Then, in view of the high-precision timing requirements for the new eLORAN system, an optimized selection is made in multiple optimal combinations. The analysis results show that the average interference rate of the optimal combination scheme obtained by this algorithm is comparable to that between the current navigation chains and can take into account the timing requirements, which can provide referential suggestions and theoretical basis for the construction of high-precision ground-based timing system.