Advanced Search
Articles in press have been peer-reviewed and accepted, which are not yet assigned to volumes /issues, but are citable by Digital Object Identifier (DOI).
Display Method:
Impact of Wireless Priors on the Computation and Energy Cost of MU-MIMO Precoding Learning
CONG Pengyu, HAN Shengqian, DENG Mingyu, LIU Shengjie, YANG Chenyang, SHEN Songhui
 doi: 10.11999/JEIT260388
[Abstract](247) [FullText HTML](98) [PDF 712KB](7)
Abstract:
  Objective  This paper investigates downlink MultiUser Multi-Input Multi-Output (MU-MIMO) precoding policy learning from the perspectives of computational complexity and energy consumption. Traditional numerical optimization algorithms achieve high performance but exhibit rapidly increasing computational complexity as the numbers of base station antennas and served users increase, leading to high inference latency and energy consumption. In recent years, deep learning has been widely adopted to reduce online computational cost; however, existing evaluations generally rely on training or inference time and FLoating-point OPerations (FLOPs), without direct measurements of energy consumption and power. More importantly, the computational cost of a deep learning model is closely related to network architecture design, which should effectively exploit the wireless prior knowledge of the precoding policy. Therefore, this paper develops a network architecture that matches the multidimensional permutation properties of the precoding policy and systematically investigates how wireless priors affect computational complexity and energy consumption through comprehensive hardware-based measurements and simulations.  Methods  The MU-MIMO precoding policy is formulated as a mapping from multiuser channel information to the optimal precoding matrix under a transmit power constraint. The optimal policy satisfies multidimensional joint permutation equivariance and invariance with respect to user indices, receive antenna indices, and base station antenna indices. To exploit these wireless priors, an Attention-based Graph Neural Network (AGNN) is proposed based on a hypergraph structure, in which the update and aggregation operations satisfy the required equivariance and invariance properties. An attention mechanism is incorporated to model inter-user interference and improve generalization across different numbers of users. For broadband precoding, multi-subcarrier channel information is aggregated at the input layer to construct an expanded feature representation. To quantify computational energy cost, a hardware measurement platform is developed to collect energy consumption and power for the GPU, CPU, and DRAM during both training and inference. Simulations are conducted using 3GPP TR 38.901 Urban Macrocell (UMa) channel datasets with different antenna array sizes and bandwidth configurations. The proposed AGNN is compared with a traditional numerical optimization algorithm based on Zero-Forcing Block Diagonalization (ZFBD) with greedy user pairing and two Transformer-based architectures that satisfy only one-dimensional permutation equivariance.  Results and Discussions  Two major findings are obtained. First, incomplete exploitation of wireless priors results in inferior performance and higher computational cost. In the MU-MISO scenario, the Transformer-based architectures achieve lower spectral efficiency than the ZFBD+Greedy baseline while requiring substantially larger models and higher inference FLOPs than AGNN. By matching the multidimensional permutation properties of the precoding policy, AGNN achieves higher spectral efficiency while reducing inference FLOPs by approximately one order of magnitude. Hardware measurements further demonstrate that AGNN reduces inference energy consumption and power on both the CPU and GPU. Second, in small- and large-scale MU-MIMO scenarios, ZFBD+Greedy increases the system sum rate by 10.9×, whereas inference FLOPs, inference time, and inference energy increase by 79.0×, 21.3×, and 38.6×, respectively. In contrast, AGNN increases the system sum rate by 11.5×, while inference FLOPs increase by only 1.89×. Meanwhile, inference time and inference energy are reduced to 0.03× and 0.20×, respectively. These results demonstrate that exploiting the multidimensional permutation properties of the precoding policy provides an effective approach for reducing computational complexity, inference latency, and energy consumption in large-scale 6G MU-MIMO systems.  Conclusions  This paper investigates the effect of wireless priors on the computational complexity and energy consumption of MU-MIMO precoding learning. By exploiting the multidimensional joint permutation equivariance and invariance of the optimal precoding policy, an AGNN is developed that is well matched to these properties. A hardware-aware measurement platform is established to obtain direct measurements of energy consumption and power for the GPU, CPU, and DRAM during training and inference. Simulations based on 3GPP TR 38.901 UMa channel datasets demonstrate that Transformer-based architectures satisfying only one-dimensional permutation equivariance achieve lower spectral efficiency while incurring substantially higher computational and energy costs. In contrast, AGNN achieves higher spectral efficiency while substantially reducing inference FLOPs, inference time, inference energy consumption, and training complexity. As system size increases, traditional numerical optimization algorithms exhibit much faster growth in computational and energy costs than in system sum rate, whereas the proposed learning method based on wireless priors maintains low inference latency and energy consumption. Overall, exploiting the wireless prior knowledge of the MU-MIMO precoding policy in network architecture design provides an effective solution for computationally and energy-efficient high-dimensional precoding optimization in future 6G networks.
A Joint Source-Channel Coding Modulation Scheme for the Transmission of Gaussian Sources
LÜ Yaping, MA Xiao
 doi: 10.11999/JEIT251224
[Abstract](248) [FullText HTML](118) [PDF 1946KB](26)
Abstract:
  Objective  The Separated Source-Channel Coding (SSCC) scheme has been proven to incur no performance loss when the source block length tends to infinity. However, SSCC usually requires a large buffer and causes long delay. It may also lead to error propagation when a single symbol error occurs in the communication channel. To alleviate these issues, Joint Source-Channel Coding (JSCC) schemes have been studied for Gaussian source transmission. In this paper, a Joint Source-Channel Coding Modulation (JSCCM) scheme is proposed for Gaussian sources. A Gaussian source reconstruction scheme and its reconstruction expression are also provided.  Methods  The Gaussian source sequence is quantized into an M-ary symbol sequence by a Lloyd-Max quantizer. For the M-ary quantized symbol sequence, a matching M-ary Fourier Transform Pair (FTP) code is constructed. The corresponding M-ary Pulse Amplitude Modulation (M-PAM) scheme is adopted for modulation. The modulated M-ary symbol sequence is transmitted using Block Markov Superposition Transmission (BMST), forming a BMST-FTP code. In addition, a Geometric Shaping (GS) scheme is proposed to obtain shaping gain. In the proposed source reconstruction scheme, the system output is the weighted average of the representative elements of the Lloyd-Max quantizer, rather than a single representative element.  Results and Discussions  Simulations are conducted over Additive White Gaussian Noise (AWGN) channels with M-PAM modulation and BMST-FTP codes over Galois Field (GF) orders 3 and 5, denoted GF(3) and GF(5). For FTP codes with random mapping, the Word Error Rate (WER) approaches the Union Bound (UB) at high Signal-to-Noise Ratio (SNR). Similarly, FTP codes with m repeated transmissions show WER performance close to the corresponding UBs. The WER performance of BMST-FTP codes with memory m also approaches the UBs in the high SNR region (Fig. 6). In terms of Symbol Error Rate (SER), the GF(3) BMST-FTP code outperforms the GF(5) BMST-FTP code (Fig. 7(a)). For the GF(5) BMST-FTP code, GS provides an SER performance gain of approximately 0.3 dB (Fig. 8(a)). In terms of distortion performance, the GF(3) BMST-FTP code performs better in the low SNR region, whereas the GF(5) BMST-FTP code performs better in the high SNR region (Fig. 7(b)). Compared with other work, the GF(3) BMST-FTP code with m = 1 achieves similar performance, whereas the GF(5) BMST-FTP code with m = 1 achieves better performance (Fig. 7(b)).  Conclusions  This work proposes a JSCCM scheme for Gaussian source transmission. In the proposed scheme, two types of BMST-FTP codes are constructed. Each code is matched with a corresponding Lloyd-Max quantizer and M-PAM modulator. A Gaussian source reconstruction scheme and its reconstruction expression are also provided. Simulation results show that an appropriate transmission scheme can be selected according to the target performance. The proposed GS scheme provides an SER gain of approximately 0.3 dB and improves distortion performance in the waterfall region.
Cross-Domain Collaborative Enhancement for Tiny Object Detection in Remote Sensing Images
ZHANG Tianyang, ZHANG Xiangrong, WANG Guanchun, TANG Xu
 doi: 10.11999/JEIT260317
[Abstract](271) [FullText HTML](128) [PDF 8181KB](24)
Abstract:
  Objective  Deep learning has substantially advanced object detection in Remote Sensing Images (RSIs). However, because of imaging conditions and the inherently small size of many objects, a large proportion of targets in RSIs occupy fewer than 16 × 16 pixels. Therefore, current object detection methods achieve substantially lower detection accuracy for tiny objects than for normal-scale objects. This limitation primarily arises from two critical factors: insufficient positive sample assignment and weak feature representation. To address these challenges, a Cross-Domain Collaborative Enhancement Detector (CDCEDet) is proposed. CDCEDet jointly optimizes label assignment in the spatial domain and enhances feature representation in the frequency domain, thereby improving the accuracy and robustness of tiny object detection in RSIs.  Methods  The overall framework of CDCEDet is illustrated in Fig. 2 and consists of three major components. First, a Scale-Adaptive Anchor Generator (SAAG) is designed to dynamically generate anchors that match the scales of ground-truth (GT) objects, thereby effectively alleviating the scale mismatch between anchors and tiny objects that has been largely overlooked in previous studies. Compared with conventional uniformly distributed anchor generators, SAAG substantially increases the number of positive samples assigned to tiny objects, even under Intersection over Union (IoU)-based label assignment. Second, a Quantile-based Adaptive Label Assignment (QALA) mechanism is developed to replace the conventional fixed IoU threshold-based label assignment. QALA models the IoU distribution between each GT object and its matched anchors to generate an adaptive label assignment threshold for each object, thereby further increasing the number of positive samples assigned to tiny objects. Third, a Frequency-Adaptive Fusion (FAF) module is developed to enhance feature representation from a frequency-domain perspective. An adaptive high-pass filter is used to strengthen high-frequency details and compensate for information loss caused by channel compression, whereas an adaptive low-pass filter preserves semantic consistency during feature upsampling, thereby reducing semantic inconsistency within upsampled objects.  Results and Discussions  Extensive experiments are conducted on two public remote sensing tiny object detection datasets, AI-TODv2 and AI-TOD-R. The proposed method is compared with several state-of-the-art methods, including RFLA, DCNet, and DCFL. On the AI-TODv2 dataset (Table 1), CDCEDet improves AP50 and AP50–95 by 1.8% and 0.7%, respectively, compared with the best existing method. On the AI-TOD-R dataset (Table 2), AP50 and AP50–95 are improved by 2.6% and 0.7%, respectively. These results demonstrate that CDCEDet achieves superior detection performance and strong generalization capability for tiny object detection in RSIs. Ablation studies and parameter analyses of the proposed modules (Tables 36) further verify the effectiveness of each component and their complementary contributions. Qualitative results on both datasets (Fig. 3) show that the proposed method accurately detects tiny objects in both sparse and dense scenes. As illustrated in Fig. 4, SAAG generates scale-matched anchors for individual objects and assigns substantially more positive samples to tiny objects than the conventional uniformly distributed anchor generator. Furthermore, visual comparisons with RFLA and DCNet (Fig. 5) demonstrate that CDCEDet achieves higher detection accuracy while substantially reducing missed detections.  Conclusions  A CDCEDet is proposed to address insufficient positive sample assignment and weak feature representation in remote sensing tiny object detection. Specifically, SAAG dynamically generates anchors that match the scales of GT objects, substantially increasing the number of positive samples assigned to tiny objects. QALA further improves label assignment by modeling the IoU distribution between GT objects and their matched anchors to adaptively determine the label assignment threshold, thereby effectively reducing the scale bias introduced by fixed IoU thresholds. In addition, FAF enhances feature representation from a frequency-domain perspective through an adaptive high-pass filter and an adaptive low-pass filter. Experimental results on two benchmark datasets demonstrate the superior detection performance and strong generalization capability of CDCEDet. Future work will focus on improving model efficiency and real-time performance to facilitate practical deployment in remote sensing applications.
A Cross-Precision Motion Compensation Technique for Security Surveillance Video Coding
JIANG Wei, MA Wei, LU Jinghui, ZHANG Yue, ZHANG Yundong
 doi: 10.11999/JEIT251301
[Abstract](259) [FullText HTML](138) [PDF 4757KB](14)
Abstract:
  Objective  High-altitude dome cameras are widely used in modern security surveillance. They are often deployed at critical locations, such as bridges and tower tops, where they are vulnerable to external interference. Such interference can cause jitter, blur, and distortion in captured videos, creating major challenges for video coding. In video compression, high-precision motion compensation is essential for improving coding efficiency. However, the existing Ultimate Motion Vector Expression (UMVE) technique has limited motion-vector precision and insufficient flexibility in adaptive adjustment. High-precision motion compensation tools, such as Registration Coding Mode (RCM) and Affine Motion Compensation Prediction (AFFINE), can improve compensation accuracy, but they require high computational complexity and hardware cost. These limitations make it difficult to meet the requirements for coding efficiency, power consumption, and real-time processing in high-altitude surveillance scenarios. Therefore, this study aims to design an optimized UMVE scheme that integrates high-precision motion compensation, low computational complexity, and scene adaptability to improve coding efficiency while balancing resource consumption.  Methods  This study proposes UMVE_CPMC, an Ultimate Motion Vector Expression technique supporting Cross-Precision Motion Compensation. The proposed method improves motion compensation accuracy by constructing an extended Up-Precision Motion Vector (UPMV), expressed as UPMV = BaseMV + MMV(p, angle). Here, Base Motion Vector (BaseMV) denotes the base vector obtained by the existing UMVE method, and Micro-Motion Vector (MMV) denotes the fine-adjustment vector defined by a specific precision p and angle. Incremental candidates are provided only at the 1/8 precision level to balance computational complexity and compression efficiency. For step-size adaptive adjustment, a six-mode improved scheme is proposed. It covers enhanced UMVE, conventional UMVE, and four precision-improved modes, allowing the encoder to switch flexibly according to scene characteristics. The average image gradient is used as an objective evaluation index. Test scenes are divided into Class A, representing high-clarity motion scenes, and Class B, representing low-clarity scenes. Different coding configurations, sequences, and parameters are used to compare coding gains and computational efficiency under different modes.  Results and Discussions  Experiments show that UMVE_CPMC improves performance under different scenes and modes. In Class A high-clarity motion scenes, with both the adaptive strategy and RCM disabled, the average gains of the Y, U, and V components in Fusion Mode 1 are –2.912%, –1.656%, and –1.654%, respectively. The average coding time is reduced to 94.55% of the baseline. In Independent Mode 1, the average Y-component gain reaches –2.925%, and the coding time is reduced to 91.91% of the baseline. Compared with conventional UMVE, when CPMC Independent Mode 1 is enabled with RCM and other tools working together, the gain improves from –0.276% to –1.310%, indicating higher cost effectiveness. In Class B low-clarity scenes, adaptive adjustment significantly reduces the losses of coding gain in Fusion Mode 1 and Fusion Mode 0. The average losses of coding gain are limited to 0.071% and 0.108%, respectively, which maintains the original coding gain. In multi-scene tests with RCM and AFFINE disabled, 9 of 10 test sequences in adaptive Fusion Mode 1 show positive gains. The Y-component gain reaches –10.691% for the yuxuedaolu sequence and –11.400% for the BQTerrace sequence. When all existing coding tools are enabled, the Y-component gains of the dianjing, yuxuedaolu, and BQTerrace sequences reach –1.29%, –2.05%, and –1.21%, respectively. The coding time is reduced to 94%~96% of the baseline. In addition, correlation analysis shows a clear positive relationship between the average image gradient and the coding gain. Images with a high average gradient, corresponding to high clarity, gain more from UMVE_CPMC, whereas images with a low average gradient, corresponding to low clarity, benefit little. Principle analysis shows that pixel changes in low-clarity images are smooth, so high-precision interpolation cannot generate effective new pixel values. The compensation effect is therefore limited. The performance differences among modes are consistent with their computational complexity. The fusion mode balances gain and stability, whereas the independent mode further reduces computation. The six step-size adaptive modes can meet the real-time and precision requirements of different scenes.  Conclusions  The proposed UMVE_CPMC technique integrates Cross-Precision Motion Compensation with the UMVE algorithm. It addresses the limited precision of conventional UMVE and the high computational complexity of high-precision motion compensation tools. It also achieves a favorable balance among coding efficiency, computational complexity, and scene adaptability. In Class A high-clarity motion scenes, UMVE_CPMC achieves notable coding gains. The gain exceeds 10% for some sequences when other high-precision motion compensation tools are disabled and reaches 1%~2% when used with other tools. In Class B low-clarity scenes, the original coding gain is maintained through a frame-level adaptive adjustment interface. In addition, the fusion mode does not increase hardware complexity, whereas the independent mode significantly reduces coding time. These features make the proposed method suitable for encoder designs with limited resources or simplified requirements. UMVE_CPMC provides an effective approach for improving the coding efficiency of high-altitude dome camera videos affected by jitter and blur. It also enriches the video coding toolset and provides practical guidance for optimizing video coding technologies in security surveillance. Future work will further optimize the adaptive strategy, explore integration with other advanced coding tools, develop scenario-specific coding schemes, and improve performance in complex scenes.
Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement
ZHANG Chunxiang, ZHANG Huibin, GAO Xueyao
 doi: 10.11999/JEIT260061
[Abstract](130) [PDF 1126KB](11)
Abstract:
  Objective  Biomedical Word Sense Disambiguation (WSD) is important for biomedical text mining and clinical data analysis. However, biomedical terms are susceptible to semantic noise, fine-grained semantic classes are difficult to distinguish, and model generalization is limited in small-sample settings. A three-branch parallel WSD framework with contrastive learning is proposed to address these problems. Electra, mDeBERTa, and Flan-T5 (FT5) are integrated to obtain complementary semantic features. Chi-square attention, a Focal Loss and Margin Loss hybrid loss, and hard sample mining are further incorporated to improve feature discrimination and model robustness.  Methods  The proposed framework uses Electra, mDeBERTa, and FT5 to extract complementary semantic features from biomedical terms. A chi-square attention module is designed to identify representative biomedical terms and guide token-level attention allocation. Focal Loss and Margin Loss are combined to improve the learning of hard samples and class boundaries under class imbalance. A two-stage hard sample mining strategy is developed by jointly considering training loss and prediction uncertainty. In addition, a constrained contrastive learning mechanism based on core disambiguation tokens is introduced. Random cropping is used to generate semantically equivalent augmented views, and the NT-Xent loss is applied to optimize the representation space.  Results and Discussions  Experiments are conducted on the MSH WSD dataset, which contains 203 ambiguous biomedical terms. The proposed model achieves an accuracy of 95.27%, a precision of 90.33%, a recall of 85.49%, and an F1 score of 87.84%. It outperforms the Neural Concept Embeddings model, which achieves an accuracy of 94.34%, by 0.93 percentage points. Ablation experiments show that the successive addition of Word2Vec, FT5, hard sample mining, chi-square attention, the hybrid loss, and contrastive learning improves model performance. Among the evaluated contrastive learning settings, NT-Xent with a temperature of 0.10 and a contrastive weight of 1.00 achieves the best performance. The proposed hard sample mining strategy, which combines training loss and prediction uncertainty, also outperforms alternative sample selection methods.  Conclusions  A three-branch parallel contrastive learning framework is proposed for biomedical WSD. Electra, mDeBERTa, and FT5 are integrated to capture complementary semantic features. Chi-square attention is used to strengthen representative feature extraction, whereas the Focal Loss and Margin Loss hybrid loss improves learning of hard samples and class boundaries. Two-stage hard sample mining and constrained contrastive learning further enhance the discrimination of fine-grained semantic classes. Experimental results demonstrate that the proposed method improves biomedical WSD performance and reduces confusion among semantically similar biomedical concepts. The framework is evaluated on English biomedical texts from the MSH WSD dataset and can be further extended to multilingual biomedical corpora and domain-specific knowledge graphs.
Finite-time Adaptive Sliding Mode Control of Servo Motors Considering Frictional Nonlinearity and Unknown Loads
ZHANG Tianyu, GUO Qinxia, YANG Tingkai, GUO Xiangji, MING Ming
 doi: 10.11999/JEIT250521
[Abstract](598) [FullText HTML](376) [PDF 3507KB](33)
Abstract:
  Objective  Ultra-fast laser processing with an infinite field of view requires servo motor systems with superior tracking accuracy and robustness. However, such systems are highly nonlinear and affected by coupled unknown load disturbances and complex friction, which constrain the performance of conventional controllers. Although Sliding Mode Control (SMC) exhibits inherent robustness, traditional SMC and observer designs cannot achieve accurate finite-time disturbance compensation under strong nonlinearities, thus limiting high-speed and high-precision trajectory tracking. To address this limitation, a novel finite-time adaptive SMC approach is proposed to ensure rapid and precise angular position tracking within a finite time, satisfying the stringent synchronization requirements of advanced laser processing systems.  Methods  A novel control strategy is developed by integrating an adaptive disturbance observer fused with a Radial Basis Function Neural Network (RBFNN) and finite-time SMC. First, the unknown load disturbance and complex frictional nonlinear dynamics are combined into a unified "lumped disturbance" term, improving model generality and the ability to represent real operating conditions. Second, a finite-time adaptive disturbance observer is constructed to estimate this lumped disturbance. The observer utilizes the universal approximation capability of the RBFNN to learn and approximate the dynamic characteristics of unknown disturbances online. Simultaneously, a finite-time adaptive law based on the error norm is introduced to update the neural network weights in real time, ensuring rapid and accurate finite-time estimation of the lumped disturbance while reducing dependence on precise model parameters. Based on this design, a finite-time SMC is developed. The controller uses the observer’s disturbance estimation as a feedforward compensation term, incorporates a carefully formulated finite-time sliding surface and equivalent control law, and introduces a saturation function to suppress control input chattering. A suitable Lyapunov function is then constructed, and the finite-time stability theory is rigorously applied to prove the practical finite-time convergence of both the adaptive observer and the closed-loop control system, guaranteeing that the system tracking error converges to a bounded neighborhood near the origin within finite time.  Results and Discussions  To verify the effectiveness and superiority of the proposed control strategy, a typical Permanent Magnet Synchronous Motor (PMSM) servo system model is constructed in the MATLAB environment, and a simulation scenario with desired trajectories of varying frequencies is established. The proposed method is comprehensively compared with the widely used Proportional–Integral (PI) control and the advanced method reported in reference[7]. Simulation results demonstrate the following: 1. Tracking performance: Under various reference trajectories, the proposed controller enables the system to accurately follow the target trajectory with a tracking error substantially smaller than that of the PI controller. Compared with the method in reference[7], it achieves smoother responses and smaller residual errors, effectively eliminating the chattering observed in some operating conditions of the latter. 2 Disturbance rejection and robustness: The adaptive disturbance observer based on the RBFNN rapidly and effectively learns and compensates for the lumped disturbance composed of unknown load variations and frictional nonlinearities. Even in the presence of these disturbances, the proposed controller maintains high-precision trajectory tracking, demonstrating strong disturbance rejection and robustness to system parameter variations. 3. Control input characteristics: Compared with the reference methods, the control signal of the proposed approach quickly stabilizes after the initial transient phase, effectively suppressing chattering caused by high-frequency switching. The amplitude range of the control input remains reasonable, facilitating practical actuator implementation. 4. Comprehensive evaluation: Based on multiple error performance indices, including Integral Squared Error (ISE), Integral Absolute Error (IAE), Time-weighted Integral Absolute Error (ITAE), and Time-weighted Integral Squared Error (ITSE), the proposed controller consistently outperforms both PI control and the method in reference[7]. It demonstrates comprehensive advantages in suppressing transient errors rapidly and reducing overall error accumulation. The method also improves steady-state accuracy and achieves a balanced response speed with effective noise attenuation. 5. Observer performance: The RBFNN weight norm estimation converges rapidly and stabilizes at a low level after initial adaptation, confirming the effectiveness of the proposed adaptive law and the learning efficiency of the observer.  Conclusions  A finite-time sliding mode control strategy with an adaptive disturbance observer is proposed for servo systems used in ultra-fast laser processing. The method models unknown load disturbances and frictional nonlinearities as a lumped disturbance term. An adaptive observer, integrating an RBF neural network with a finite-time mechanism, accurately estimates this disturbance for real-time compensation. Based on the observer, a finite-time SMC law is formulated, and the practical finite-time stability of the closed-loop system is theoretically proven. Simulations conducted on a permanent magnet synchronous motor platform confirm that the proposed approach achieves superior tracking accuracy, robustness, and control smoothness compared with conventional PI and existing advanced methods. This work offers an effective solution for achieving high-precision control in nonlinear systems subject to strong disturbances.
Reynolds Decomposition Motion-Guided Texture Learning for Scarred Myocardium Phenotyping
RUAN Dongsheng, YANG Daiguo, ZHANG Xiaolin, MA Jianhua, JIANG Mingfeng, WANG Yaming
 doi: 10.11999/JEIT260330
[Abstract](67) [PDF 4272KB](6)
Abstract:
  Objective  Scarred myocardium is a key imaging marker of multiple cardiovascular diseases, and accurate phenotyping is clinically valuable for treatment planning and prognosis. Cine-MRI noninvasively provides both cardiac motion and anatomical texture information. However, existing classification methods have two major limitations. First, motion representation is often overly discretized and insensitive to subtle abnormalities. Second, multimodal fusion commonly relies on simple feature concatenation, which limits the ability to capture the deep pathological association between motion impairment and texture alteration. A motion-guided texture learning framework is therefore developed to improve the accuracy of non-invasive scarred myocardium classification.  Methods  A Motion-Guided Texture fusion Network (MGTNet) is proposed to enable deep interaction between motion and texture features. First, inspired by Reynolds decomposition in fluid dynamics, a Reynolds Decomposition Motion Network (RDMNet) is designed within a diffeomorphic registration framework to decompose the myocardial motion field into a regular mean motion component and an abnormal pulsatile motion component. This decomposition provides a refined motion representation that is sensitive to subtle abnormalities. Second, a Motion-Guided Cross-Attention (MGCA) module is designed, in which refined motion features serve as query features to dynamically modulate texture features and enhance the perception of suspected lesion regions. In addition, an inter-frame motion interaction module is used to capture temporal dependencies across cardiac frames, while an intra-frame texture extraction module learns multi-scale spatial texture patterns within each frame. Through a serial pipeline of motion field estimation, motion-guided texture enhancement, and feature fusion, end-to-end scarred myocardium classification is achieved.  Results and Discussions  Experiments on the CMRD and ACDC datasets show that MGTNet consistently outperforms existing single-modality and multimodal methods (Tables 1 and 2). It achieves accuracies of 95.6% on CMRD and 97.3% on ACDC, with AUC values of 96.6% and 96.0%, respectively. Compared with the baseline MTNet, MGTNet improves accuracy by up to 1.4 percentage points and the F1-score by up to 2.3 percentage points. Further comparisons with different motion field estimation methods (Tables 3 and 4) show that RDMNet provides more discriminative motion priors and achieves the best overall performance on both datasets. These results indicate that fine-grained motion modeling and motion-guided texture enhancement effectively capture complementary pathological information from Cine-MRI. The ROC curves of the nine methods on both datasets further support the superior classification performance of the proposed method.  Conclusions  A scarred myocardium classification method based on motion-guided texture learning is presented. Reynolds decomposition is incorporated into motion field estimation to separate regular mean motion from abnormal pulsatile motion, and cross-attention is used to guide texture extraction with motion priors. This strategy addresses the limitations of discrete motion representation and shallow multimodal fusion. The results confirm that deep motion-texture interaction improves the accuracy and robustness of non-invasive scarred myocardium classification and provides an effective approach for Cine-MRI-based myocardial phenotyping.
Labeled Multi-Bernoulli Sensor Management Strategy Based on Twin-Delayed Deep Deterministic Policy Gradient Learning Mechanism
ZHANG Xindi, CHEN Hui, ZHANG Hongyun, LIAN Feng, ZHANG Guanghua, YIN Zhipeng
 doi: 10.11999/JEIT260045
[Abstract](265) [FullText HTML](109) [PDF 1633KB](25)
Abstract:
  Objective  Multi-target tracking requires sensor management to adapt the observation process to clutter, missed detections, target-number variations, and target-motion changes. Conventional methods typically search over a finite set of sensor actions, which increases computational cost and limits control resolution. Furthermore, reward functions constructed from multiple single-target metrics may not adequately characterize the joint multi-target posterior. To address these limitations, a continuous-action sensor management method that integrates Twin-Delayed Deep Deterministic Policy Gradient (TD3) with the Labeled Multi-Bernoulli (LMB) filter is proposed to optimize the mobile-sensor heading angle according to the multi-target belief state.  Methods  The LMB posterior, including target existence probabilities and state densities, is used to construct the belief state. At each filtering step, the mobile sensor selects a continuous heading angle that determines the sensor-target geometry, detection probability, and LMB update. Predicted target states and candidate heading actions are used to generate pseudo measurements and obtain pseudo-updated LMB densities. The Cauchy-Schwarz (CS) divergence between the predicted and pseudo-updated LMB densities is adopted to construct an information-gain reward. TD3 employs twin critics, target policy smoothing, and delayed policy updates to reduce value-estimation bias. Random control, Policy Gradient (PG)-based sensor management, CS divergence-based sensor management, and Deep Deterministic Policy Gradient (DDPG)-based sensor management are used for comparison.  Results and Discussions  DDPG-LMB and TD3-LMB produce smoother sensor trajectories than the discrete-action methods (Fig. 2). TD3-LMB achieves the highest or near-highest detection probabilities for most targets (Fig. 3) and yields larger CS divergence values during most time steps, while random control consistently produces lower values (Fig. 4). TD3-LMB also achieves the lowest overall Optimal Subpattern Assignment (OSPA) distance in the evaluated scenario, while DDPG-LMB generally outperforms the discrete baseline methods (Fig. 5). These results demonstrate that continuous heading-angle control improves observation quality and enhances the overall tracking performance of the LMB filter.  Conclusions  A TD3-based continuous-action sensor management framework for the LMB filter is presented. Candidate heading actions are evaluated using pseudo-updated LMB densities and CS divergence, directly associating action selection with the joint multi-target posterior. Simulation results demonstrate smoother sensor trajectories, higher detection probabilities for most targets, greater information gain, and lower OSPA distances in the evaluated scenario. Future work will consider higher-dimensional action spaces and cooperative multi-sensor management.
Multi-projection Plane InISAR 3D Reconstruction Method for Complex Moving Ship Targets
LI Ning, NIU Jinfa, WANG Weibin, HU Xingwang, WU Lin
 doi: 10.11999/JEIT251268
[Abstract](485) [FullText HTML](218) [PDF 3554KB](50)
Abstract:
  Objective  Interferometric Inverse Synthetic Aperture Radar (InISAR) is a three-dimensional (3D) reconstruction technique for non-cooperative targets. However, the complex 3D rotational motion of ship targets causes unstable Doppler frequency variation. Inverse Synthetic Aperture Radar (ISAR) imaging also inevitably suffers from target overlap and occlusion. These factors make high-precision and complete 3D reconstruction under a single projection plane difficult. Therefore, a multi-projection plane InISAR 3D reconstruction method for complex moving ship targets based on point cloud fusion is proposed. The method supplements target 3D information through efficient and high-precision point cloud registration and fusion, thereby significantly improving 3D reconstruction quality.  Methods  This method fully exploits the advantage of multi-plane observation enabled by the severe motion of ship targets. The ship centerline is extracted, and the vertical rotation vector is estimated by Principal Component Analysis (PCA) to select the optimal imaging times corresponding to different Imaging Projection Planes (IPPs). ISAR imaging and InISAR 3D reconstruction are then completed. In addition, a point cloud fusion algorithm that combines Weighted Random Sample Consensus (RANSAC) and hierarchical Iterative Closest Point (ICP) is proposed. The random sampling process is optimized through a feature stability weighting strategy, which enables efficient extraction and matching of corresponding feature points in InISAR images and achieves high-precision point cloud fusion under multiple IPPs.  Results and Discussions  Experimental results show that the proposed method significantly improves reconstruction accuracy and target completeness. For simulated ship point-target data, Fig. 7 shows excellent results, with a significant reduction in reconstruction error. Signal-to-Noise Ratio (SNR) analysis shows that the quality of 3D fusion imaging improves steadily as the SNR increases from –10 dB to 10 dB, and robust fusion performance is maintained even under low-SNR conditions. For simulated destroyer Radar Cross Section (RCS) data, the method achieves strong registration performance. The detail recovery and structural integrity of the fused image are also significantly improved, effectively addressing the incomplete reconstruction of 3D information caused by scattering-point overlap and occlusion.  Conclusions  To address the low reconstruction accuracy and information loss caused by target rotation, overlap, and occlusion in traditional InISAR methods for 3D reconstruction of complex moving ship targets, a multi-IPP InISAR 3D reconstruction method based on point cloud fusion is proposed. The method uses a PCA-based optimal imaging time selection strategy. Weighted RANSAC and hierarchical ICP algorithms are then applied to achieve efficient and high-precision registration and fusion of InISAR point clouds under multiple IPPs, thereby producing high-quality 3D reconstruction results. Multi-scenario experiments are conducted by constructing both a ship model with ideal scattering points and an electromagnetic simulation RCS model with occlusion effects. The results verify the accuracy of the proposed method under ideal conditions and demonstrate its applicability in complex real-world scenarios.
DroneRFc-MM: Anti-UAV Multimodal Detection Measured Dataset
YU Taosong, YANG Qianqian, HU Zhuo, LI Mingkai, WU Jiajun, SU Yufan, PAN Junyu, SHI Zhiguo, CHEN Jiming
 doi: 10.11999/JEIT260889
[Abstract](935) [PDF 1254KB](192)
Abstract:
Objective: A comprehensive multimodal benchmark is developed for Anti-Unmanned Aerial Vehicle (UAV) detection in low-altitude urban environments. Existing datasets generally provide limited sensing modalities and UAV models, with relatively coarse annotations that constrain tasks requiring spatial, motion, and cross-modal information. DroneRFc-MM addresses these limitations by providing synchronized multimodal data, broader coverage of consumer-grade DJI UAV models, and fine-grained annotations for target detection, UAV model recognition, trajectory analysis, flight-direction reasoning, and multimodal fusion evaluation. Methods: DroneRFc-MM is synchronously collected using six heterogeneous sensor types: a Pan-Tilt-Zoom (PTZ) camera, a fisheye camera, a Radio Frequency (RF) antenna, LiDAR, millimeter-wave radar, and a microphone array. Data are acquired on an open rooftop at a university in Zhejiang Province, representing a typical urban low-altitude environment. The dataset contains recordings of six consumer-grade DJI UAV models. All devices are synchronized using a common network time reference, with inter-device timestamp discrepancies of approximately 0.3 s. The UAVs fly in “H”-shaped and vertical reciprocating trajectories at distances of 20–60 m from the sensor array. Fine-grained annotations, including UAV model, position, attitude, and velocity, are derived from flight logs. For the flight-direction reasoning task, approximately 5-s multimodal clips are generated, including camera videos, RF spectrogram videos, microphone audio, and coordinate-based text representations of radar point-cloud data. Zero-shot inference is conducted using Qwen 3.6-Plus and Qwen 3.5-Omni-Plus with unified prompts. Prediction accuracy and inference time are evaluated by comparing predicted directions with ground-truth directions calculated from UAV positioning data. Results and Discussions: The DroneRFc-MM dataset provides multimodal data from six sensor types and six consumer-grade DJI UAV models, together with fine-grained annotations and sample extraction tools. In the flight-direction reasoning task, the Qwen-series multimodal large language models (MLLMs) achieve accuracies ranging from 20% to 30% across the different input modalities. The inference time is also relatively long, with the mean response time exceeding 40 s for most sensor inputs. These results indicate that current general-purpose MLLMs can capture weak motion-related information from UAV videos, audio, RF spectrograms, and point-cloud data, but their accuracy and response speed remain insufficient for practical real-time Anti-UAV detection. Conclusions: DroneRFc-MM provides a multimodal benchmark for Anti-UAV detection, UAV model recognition, flight-direction reasoning, and multimodal model evaluation. The dataset integrates six sensor types, six consumer-grade DJI UAV models, and fine-grained annotations within a common measurement framework. The experimental results show that current general-purpose MLLMs remain limited in flight-direction reasoning and real-time inference in Anti-UAV scenarios. Domain-specific pre-training, supervised fine-tuning, knowledge augmentation, and lightweight inference are therefore needed to improve their practical utility. Future work will expand the dataset scale and application scenarios to support intelligent and efficient low-altitude airspace management systems.
Resilience-Aware Cooperative Mission Planning Algorithm for Multiple UAV Systems in Complex Dynamic Environments
ZHAO Xuejian, XIE Lulu, WANG Enliang
 doi: 10.11999/JEIT260138
[Abstract](303) [FullText HTML](132) [PDF 3042KB](26)
Abstract:
  Objective  This paper addresses the strongly coupled problem of task allocation and route planning in cooperative task and route planning for multiple UAV systems operating in complex dynamic environments, where dynamic task arrivals, UAV failures, no-fly-zone constraints, and link quality degradation occur simultaneously.  Methods  A Resilience-Aware Hybrid Swarm Optimization (RAHSO) algorithm is proposed. First, an integrated task-route planning model is established by jointly considering task value, route cost, energy consumption, interference penalties, time-window constraints, platform capability constraints, conflict resolution, and link quality within a unified optimization framework. High-quality initial solutions are generated through clustering-based and genetic initialization. A hybrid optimization framework that integrates the Dung Beetle Optimizer (DBO), Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Variable Neighborhood Search (VNS) is then employed to perform global exploration and local refinement. In addition, Tarjan-based deadlock detection and repair are incorporated to guarantee feasible task assignments. Finally, an event-driven Proximal Policy Optimization (PPO) online replanning module is designed to rapidly update affected task subsets in response to emergent tasks, UAV failures, and network topology changes.  Results and Discussions  Comparative and ablation experiments are conducted under static, large-scale, dynamic-event, and interruption scenarios. The results demonstrate that the proposed method consistently outperforms representative baseline algorithms in task completion rate, accumulated task value, average energy consumption, recovery time, and resilience index while maintaining satisfactory online replanning latency.  Conclusions  The proposed method provides an effective solution for resilient cooperative task and route planning for multiple UAV systems operating in complex dynamic environments.
Dynamic Data Mapping and Co-Optimization Method for TSVs in 3D-Integrated MoE Accelerators
YANG Jialin, XIA Chenjie, WU Huiming, LI Ningyuan, SONG Yuan, LIU Bo
 doi: 10.11999/JEIT260565
[Abstract](52) [PDF 2124KB](7)
Abstract:
  Objective  The rapid progress of large-scale intelligent computing, especially Mixture-of-Experts (MoE), has positioned Three-Dimensional Integrated Circuit (3D IC) based on Through-Silicon Via (TSV) as a key solution to memory-wall bottlenecks via high bandwidth and density. As a core 3D IC technology, TSVs enable vertical inter-chip connections, reducing path length, parasitic delays, power, and boosting data rates. MoE-specific accelerators, characterized by high data density and strong fault tolerance, introduce new challenges and opportunities for TSV layout. These include aggravated signal integrity and reliability issues in dense arrays, and the inadequacy of static TSV allocation for dynamic, bursty MoE traffic. Conversely, their inherent fault tolerance permits optimization design spaces for employing fault-tolerance mechanisms. This paper exploits MoE dataflow characteristics and hardware fault tolerance to devise a data mapping strategy for high-density TSV arrays based on fault-tolerance mechanisms, targeting improved performance and reliability.  Methods  This paper investigates cluster partitioning schemes and data mapping strategies to enhance the reliability of TSV data transmission. To address the high complexity of global optimization in large-scale TSV arrays, a cluster size partitioning scheme is proposed. By structurally partitioning a large-scale TSV array into several small-scale TSV clusters, the global optimization problem is decomposed into local, scalable subproblems, thereby improving optimization efficiency and flexibility while ensuring optimization moderation. Through comprehensive consideration of multiple metrics and simulation-based evaluation, the cluster size is finally determined to be 6×6. In response to the varying dataflow characteristics and load distribution across different computational stages, this paper proposes a Phase- and Load-Aware Dynamic Data Mapping (PLDM) strategy. The strategy pre-partitions the TSV array into multiple fixed-size clusters and classifies them into critical clusters and general clusters based on metrics such as coupling strength, bandwidth, and latency. At runtime, the PLDM strategy dynamically adjusts data mapping according to the characteristics of different computational stages. Furthermore, this paper achieves a co-optimization design of PLDM with the encoding circuit. The load monitoring module and the error monitoring module share certain data buffers and control status registers, enabling hardware resource reuse. Meanwhile, the error monitoring results provide real-time feedback on the reliability level of each TSV cluster, based on which the mapping controller preferentially allocates data transmission to clusters with lighter loads and lower bit error rates. This approach realizes resource sharing and load balancing, thereby improving data transmission reliability and link utilization efficiency for high-density TSV arrays.  Results and Discussions  This paper analyzes the bandwidth utilization and load balancing performance of three mapping schemes: random, static, and dynamic. The results show that both the dynamic and random mapping schemes achieve average bandwidth utilization close to the theoretical maximum. However, the random mapping scheme maps approximately 37.52% of critical data into general clusters with relatively high bit error rates, thereby increasing unreliability. Compared with static mapping, the dynamic mapping scheme improves average bandwidth utilization from 0.7982 to 0.8984, a relative increase of about 12.6%, reduces inter-cluster load fluctuation by 54.6%, and correspondingly improves load balancing by a factor of 2.2 (Fig. 4). Compared with random mapping, the dynamic mapping scheme reduces inter-cluster load fluctuation by about 8.6%, and reduces the latency of critical data and non-critical data by 36.6% and 34.7%, respectively (Table 2). To further evaluate the optimization effects of the proposed PLDM strategy on metrics such as load balancing and bandwidth utilization, four comparative schemes are configured: (1) Baseline scheme; (2) Static mapping scheme; (3) Sparse TSV layout using TSV-Aware Adaptive Fault-Tolerant Coding (TSV-AFTC) and PLDM; (4) High-density TSV layout based on scheme (3). Taking the Qwen3-30B-A3B model as an example, the normalized loads of 16 clusters in the Multi-Head Attention (MHA) and Feed-Forward Network (FFN) stages are compared across the four schemes. The results indicate that the proposed dynamic data mapping scheme achieves balanced load distribution across clusters in both the MHA and FFN stages, ranging from 0.48 to 0.52, while ensuring that all critical data are mapped to critical clusters. The high-density TSV scheme further reduces the load per cluster to approximately 0.34–0.37, demonstrating that dynamic mapping can effectively suppress stage-wise hot spots and improve load balancing (Fig. 7). Subsequently, system-level fault injection is applied to the transmitted data to simulate data reliability under extreme conditions for different schemes. The results show that for the proposed scheme (sparse), the degradation in perplexity (PPL) compared to the ideal case is controlled within 0.02, while the average bandwidth utilization is improved by approximately 15% and cluster load balancing is enhanced by a factor of 3.4. Under the high-density scheme, the PPL increase is controlled within 0.05, the average bandwidth utilization reaches about 71.7%, and the cluster load balancing is improved by a factor of 2 (Table 4).  Conclusions  This paper investigates TSV data mapping for 3D MoE accelerators and proposes a PLDM strategy based on TSV-AFTC, which allocates data from different computational stages to reliable and lightly loaded TSV clusters according to cluster-level bit error rates and load conditions. Through circuit co-design, approximately 12% of hardware resources can be saved. Compared with static mapping, the proposed scheme improves average bandwidth utilization by about 12.6% and enhances cluster load balancing by a factor of 2.2. Under system-level fault injection, the scheme limits the degradation of model inference perplexity to within 0.02, while achieving approximately 15% improvement in bandwidth utilization and a 3.4× enhancement in cluster load balancing.
A Dual-Trellis Message-Passing Decoding for Non-Binary LDPC Codes
XX XX
 doi: 10.11999/JEIT260958
[Abstract](57) [PDF 3254KB](4)
Abstract:
  Objective  Due to their capacity approaching performance, Low Density Parity-Check (LDPC) codes have been widely applied to wireless communication and data storage systems. Compared to their binary counterparts, Non-Binary LDPC (NB-LDPC) codes with short or moderate code lengths have been demonstrated to achieve superior error performance under non-binary Belief Propagation (BP) decoding. However, the computational complexity of Check Node (CN) update of the optimal BP decoding is too complex for practical applications. Recently, many works have been presented to perform updates of CNs based on truncated messages, rather than full-length reliability messages, to significantly reduce the computational complexity of CN updates. Most of them construct the trellis of a CN based on the truncated input vectors, called truncated-trellis, such that CN updates are efficiently processed in parallel based on the selected candidate paths. These paths generally contain only a small number of deviation nodes, and such deviation nodes usually have high reliability. However, the Variable Node (VN) update in most decoding algorithms based on CN truncated-trellis still sequentially processes each element in the input vectors of each VN by the elementary steps. When the CN update is simplified, the complexity of the VN update may primarily determine the overall computational complexity. To address the above issues, this paper proposes the Dual-Trellis Min-Sum (DTMS) decoding algorithm. By further introducing truncated-trellises for VNs and updating the output messages of CNs and VNs in parallel, respectively, it further improves the decoding efficiency, while maintaining the similar decoding performance.  Methods  The different contributions of nodes in the CN truncated-trellis of the Pruning path Min-Sum (PMS) decoding algorithm on the selected highly reliable candidate paths are first analyzed, and it reveals that the selected highly reliable paths are primarily determined by the deviation nodes from the first few rows of the trellis of a CN, especially the second row. Thereby, it is not critical to update and sort every element of each output vector of one VN during the VN update. Next, a new trellis of one VN is constructed, and highly reliable elements over this trellis shared by all the output vectors of this VN are searched using a row-wise pruning strategy, such that the conventional element-wise VN updating procedure is transformed into a trellis-based parallel updating process based on an extra column in the trellis. In this basis, the unequal protection for the reliability values of each VN output vector is conducted, e.g., only the first few elements in each output vector of VN are updated and arranged, and the rest elements of each output vector are directly set to a compensation value. As a result, the computational complexity required for less reliable elements during each VN update can be significantly reduced, while retaining the crucial messages.  Results and Discussions  Experimental results show that compared with the PMS decoding algorithm using the original VN updating procedure, the proposed DTMS decoding algorithm maintains almost the same Bit Error Rate (BER) performance and convergence speed for decoding NB-LDPC codes under different finite fields, code lengths, and code construction methods (Figs. 37). Meanwhile, the number of real-domain operations required for the proposed simplified VN updates is reduced by approximately 71.8% on average (Table 2). In addition, the error-correction performance and convergence speed of the proposed DTMS decoding algorithm are close to those of the sub-optimal BP decoding algorithms (Figs. 37) with relatively low computational complexity (Table 3). The average performance gap of the DTMS decoding algorithm from the optimal BP decoding algorithm is only about 0.11 dB (Figs. 37). Thus, optimizing the VN updating is an effective way to further reduce the decoding complexity of truncated-trellis-based message-passing decoding algorithms.  Conclusions  This paper proposes a DTMS decoding algorithm to reduce the computational complexity of VN update in truncated-trellis-based decoding algorithms for NB-LDPC codes. Based on the CN updating process of the PMS decoding algorithm, the proposed algorithm further constructs a truncated-trellis and introduces the unequal protection scheme for VN update, such that the output vectors of each VN can be efficiently updated in parallel. Experimental results show that, under the same CN trellis-based update, the proposed parallel VN updating method significantly reduces the computational complexity compared with the original VN updating method, while maintaining similar decoding performance. Moreover, the proposed DTMS decoding algorithm performs closely to the sub-optimal BP decoding algorithms with similar convergence speed and lower complexity. In future studies, it will be interesting to further exploit the adaptive pruning strategies for the VN parallel updates. Based on the distribution of field elements from different iterations, less reliable field elements can be adaptively eliminated to reduce the set of candidate field elements, which may further reduce the complexity of VN update with negligible performance loss.
Study on Deployment Optimization of Reconfigurable Intelligent Surface for Troposcatter Communications
ZHAO Ziyan, SONG Zhiqun, LIU Lizhe, LI Yong, LI Xingjian, WANG Bin
 doi: 10.11999/JEIT260922
[Abstract](51) [PDF 3106KB](4)
Abstract:
  Objective   Troposcatter communication serves as a valuable complement to satellite communication and thus is still quite promising in scenarios such as military long-distance communication. However, when a troposcatter communication system is deployed in mountainous environments, it is often faced with a prevalent and challenging engineering problem known as the “Line-of-Sight (LoS) obstruction”. Traditional solutions to this issue are still confronted with engineering difficulties. Increasing the antenna elevation angle to cross obstacles makes the scattering angle increase sharply and consequently lead to transmission loss surging beyond acceptable link budget limits; alternatively, building tall towers to raise antenna height preserves low-angle transmission but introduces construction difficulties and sacrifices the advantage of terrain concealment. Reconfigurable Intelligent Surface (RIS) has emerged as a disruptive technology in wireless communications, with the capability of reconstructing the wireless environment and artificially altering channel characteristics. It has been successfully applied in various civilian mobile communication systems. Obviously, it also provides an alternative to address the problem of LoS obstruction in troposcatter communications. Unfortunately, it has never been reported that RIS had been applied in such scenarios. Herein, to solve the LoS obstruction problem in troposcatter communications, RIS is involved for the first time in this field, a conceptual architecture of RIS-assisted troposcatter communication is set up, and then the problem of optimal RIS deployment is systematically investigated.  Methods   Based on the proposed framework of RIS-assisted troposcatter communication system, the deployment optimization of RIS is addressed step by step:Firstly, a three-dimensional model of feasible deployment region is established under four types of practical engineering constraints, i.e., the intrinsic constraint of obstacle-crossing, the optimal constraint of engineering upper bound, the antenna radiation constraint of Fresnel near-field region, and the hardware constraint of RIS effective angle.Secondly, the optimization problem of RIS deployment is formulated as minimizing the comprehensive system gain loss. The overall loss consists of three major components, namely, troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles. The first two parts are easily computed according to corresponding engineering knowledge of troposcatter communication and typical antenna theory, respectively. Then to calculate the dynamic gain attenuation of Cassegrain antennas, a quantitative model is developed based on cantilever beam bending theory. The model quantifies the pointing errors of a Cassegrain antenna caused by dynamic over-compensation with its elevation angle adjustment, and then the gain loss is calculated with the assistance of Taylor radiation pattern.Thirdly, through theoretical analysis and numerical verification via sectional slicing heatmaps, a dimensionality reduction property of the objective function is observed and validated. Within the feasible region, the first-order partial derivative of the objective function with respect to deployment height is always negative, implying that the global optimal deployment position necessarily lies on the upper boundary surface of the feasible region. The dimensionality reduction property is rigorously validated through slice analysis across the entire feasible domain, with more than 46,100 verification points confirming that the optimal position always resides on the upper boundary surface. This finding reduces the intractable three-dimensional constrained optimization problem to a much simpler two-dimensional manifold optimization, which significantly reduces the computational complexity of the optimization.Finally, based on this dimensionality reduction property, an improved gradient descent algorithm with momentum and adaptive backtracking line search (IGD-M&ABLS) is put forward. The algorithm introduces momentum gradient updates to suppress zigzag oscillations and accelerate convergence; it also incorporates an adaptive backtracking line search strategy to dynamically adjust step sizes, balancing iterative stability with computational efficiency.  Results and Discussions   A series of simulation experiments are conducted under typical engineering parameters, i.e., a 3-meter aperture Cassegrain antenna; 5 GHz signal frequency; mountain heights of 150 m, 200 m and 250 m, representing medium-high hills, the dividing line between hills and mountains, and relatively mountainous terrain, respectively. The proposed IGD-M&ABLS algorithm is benchmarked against Grid Search (GS, a classic deterministic exhaustive-search method) and Particle Swarm Optimization (PSO, a typical efficient heuristic algorithm). The results demonstrate that IGD-M&ABLS consistently converges to the global optimal solution with less gain losses than both benchmarks. Specifically, for the typical case with a 200 m mountain height, in 50 independent runs, IGD-M&ABLS always achieves the best objective function value of 11.6094 dB, much more steadily than PSO does, and it also outperforms GS's 11.6127 dB. As far as time consumption is concerned, IGD-M&ABLS exhibits remarkable advantages of computational efficiency. Its average runtime is approximately 0.005 seconds, compared with 0.07 seconds of PSO and more than one hour of GS. This order-of-magnitude improvement in computational speed is consistent with the theoretical complexity analysis. IGD-M&ABLS optimizes two independent variables on a 2D manifold, whereas PSO and GS handle three independent variables in the 3D feasible domain. Robustness tests under varying terrain conditions (H = 150 m and H = 250 m) confirm that IGD-M&ABLS reliably obtains the best results across different scenarios. In all simulation tests, IGD-M&ABLS demonstrates excellent stability and reproducibility, producing consistent results across multiple independent runs, while PSO exhibits randomness-induced variations and GS remains limited by its discretization step size.  Conclusions   This paper pioneers the application of RIS technology in troposcatter communication, providing a new technical solution to address the LoS obstruction problem in mountainous environments. A conceptual framework of RIS-assisted troposcatter communication system is established, incorporating a three-dimensional feasible RIS deployment region model with four practical engineering constraints. Then the optimization problem of RIS deployment is formulated as minimizing the overall system gain loss including troposcatter transmission loss, free space path loss, and the dynamic gain attenuation of Cassegrain antennas with respect to their elevation angles, and the computation method of its third term is also developed for the first time based on cantilever beam bending theory. More interestingly, the objective function is found and verified with dimensionality reduction property, i.e., its minimum value always resides on the upper boundary manifold surface. That property effectively transforms the complex 3D optimization into an equivalent 2D manifold problem. Finally, a new algorithm IGD-M&ABLS is proposed by introducing momentum and adaptive backtracking line search into the traditional gradient descent framework. Simulation results show that compared with benchmarks, IGD-M&ABLS algorithm achieves the best deployment positions with order-of-magnitude faster computation, while maintaining excellent stability and reproducibility.
Non-Orthogonal PSWFs Signal Detection Method Based on Adaptive Temporal-Spatial Feature Fusion
CHEN Wenhua, MAO Zhongyang, LU Faping, SUN Ye, GAO Yixuan
 doi: 10.11999/JEIT260024
[Abstract](344) [FullText HTML](148) [PDF 6145KB](20)
Abstract:
  Objective   To address the demands of B5G/6G systems for high spectral efficiency and transmission reliability, Prolate Spheroidal Wave Functions (PSWFs)-based non-orthogonal modulation has attracted extensive research interest because of its strong time-frequency energy concentration. However, severe mutual interference among multiplexed PSWF signals degrades the performance of conventional detection methods in complex channel environments. Existing methods are limited by ideal channel assumptions or single-modal feature extraction and therefore cannot fully exploit the temporal and spatial information of PSWF signals or adapt to dynamic interference. An Adaptive Temporal-Spatial Feature Fusion (ATSFF) architecture is proposed for accurate and robust detection of non-orthogonal PSWF signals.  Method   A dual-path parallel framework is constructed to extract complementary temporal and spatial features. A Gated Recurrent Unit (GRU) network extracts deep temporal features and captures long-term dependencies from one-dimensional received signals. In the other path, one-dimensional signals are transformed into two-dimensional representations using the Gramian Angular Difference Field (GADF), and hierarchical spatial features are extracted using ResNet50. An adaptive probability-weighted fusion mechanism dynamically adjusts the contributions of the two feature branches according to their prediction uncertainty, thereby integrating complementary temporal and spatial information and improving detection robustness.  Results and Discussion   Simulations on a 32-class non-orthogonal PSWF signal dataset (Fig. 2) show that the proposed ATSFF method outperforms coherent detection, cross-term detection, Approximate Message Passing-Interleave Division Multiple Access (AMP-IDMA), and Temporal Multiple Sparse Bayesian Learning-Least Squares (TMSBL-LS) over the full Signal-to-Noise Ratio (SNR) range. t-SNE visualization (Fig. 4) shows that the fused features achieve better inter-class separation and greater intra-class compactness. At a bit error rate of 4 × 10–5, the proposed method achieves a gain of approximately 0.2 dB over cross-term detection (Fig. 6). Although ATSFF has higher computational overhead and lower real-time performance than conventional methods, its single-sample inference cost remains fixed after the network architecture is established, and GPU-based batch processing is supported. The method is therefore suitable for communication scenarios with high detection-accuracy requirements.  Conclusions   An adaptive temporal-spatial feature fusion method is proposed for non-orthogonal PSWF signal detection under severe mutual interference. Dual-path feature extraction is achieved using GRU and ResNet50, and a prediction-uncertainty-based adaptive probability-weighted fusion mechanism is used to integrate complementary temporal and spatial features. The simulation results demonstrate improved detection accuracy and robustness under complex channel conditions. The proposed method provides a feasible approach for high-accuracy detection of non-orthogonal PSWF signals.
2026, 48(7): 1-1.  
[Abstract](66) [FullText HTML](32) [PDF 3236KB](13)
Abstract:
2026, 48(7): 1-4.  
[Abstract](55) [FullText HTML](23) [PDF 287KB](9)
Abstract:
Excellence Action Plan Leading Column
A Survey of Processor Security
CHEN Congcong, GU Zhiyang, ZHANG Jiliang
2026, 48(7): 2765-2780.   doi: 10.11999/JEIT260026
[Abstract](1057) [FullText HTML](313) [PDF 1331KB](62)
Abstract:
  Significance   Processor security is a cornerstone of modern information security. Cryptographic algorithms, operating systems, and applications have long relied on processors as trusted computing bases. However, as Moore’s Law slows, modern processors increasingly adopt aggressive microarchitectural optimization techniques to improve performance and energy efficiency, often without sufficient security consideration. This trend has led to frequent security vulnerabilities in recent years. In particular, microarchitectural timing channels, exemplified by Meltdown and Spectre, exploit timing differences caused by microarchitectural state changes to break fundamental hardware and software isolation, affecting billions of devices worldwide. At the same time, the boundary between architectural and microarchitectural behavior has become less clear, giving rise to new attack paradigms and turning timing channels from isolated hardware flaws into cross-layer system security problems.  Progress   Although substantial progress has been made in the study of timing channels, existing surveys still have several limitations. First, the mechanisms of timing channels are highly diverse, and the set of exploitable components continues to grow. Hardware-centric classification schemes are therefore insufficient to capture emerging and previously unknown attacks, and they often obscure the common features shared across different techniques. Second, as traditional microarchitectural channels become better understood and partially mitigated, leakage increasingly shifts to higher-level shared resources, including operating system policies and software-managed shared resources. However, previous studies have often treated software mainly as an execution context rather than a direct source of timing leakage. In addition, current discussions of defenses tend to emphasize individual techniques, with limited analysis of their scope and failure modes.  Contributions   This survey systematically reviews timing channels from a cross-layer perspective and unifies hardware- and software-based timing channels under a common abstraction. Four necessary conditions for timing channel exploitation are identified, and a unified classification framework is established based on the nature of shared mutable state and the mechanisms that make timing differences observable. Within this framework, representative attacks from the past decade are comprehensively reviewed, their attack procedures are systematically analyzed, and their common features are clarified. In addition, existing defense mechanisms are classified according to the leakage conditions they are intended to disrupt, and their scope and possible failure modes are examined. This survey also reviews current automated vulnerability detection methods.  Prospects   Future research on timing channels faces several emerging challenges. New microarchitectural optimization techniques continue to create new attack surfaces, while resource sharing at the software level may produce additional forms of timing leakage. Moreover, emerging platforms, including chiplet-based architectures, cloud computing environments, hardware accelerators, and heterogeneous systems, are likely to expose new types of timing channels that require systematic study.
Review of Non-invasive Brain-Computer Interfaces for Continuous Motor Control
XU Minpeng, JIA Leyi, ZHOU Xiaoyu, CHEN Enze, WANG Junyang, XIAO Xiaolin, MING Dong
2026, 48(7): 2781-2791.   doi: 10.11999/JEIT260011
[Abstract](928) [FullText HTML](664) [PDF 1029KB](46)
Abstract:
  Significance   Continuous motor control is a core capability of Brain-Computer Interface (BCI) systems for natural and efficient interaction with external devices. Compared with discrete command-based control, continuous control supports real-time and smooth regulation of motion parameters such as position, velocity, and trajectory. This capability is required for applications in assistive mobility, neurorehabilitation, robotic manipulation, and immersive human-machine interaction. Although invasive BCIs have achieved high-performance continuous control through high-quality neural recordings, their dependence on surgical implantation limits long-term use and large-scale deployment. A systematic review of non-invasive continuous motor control BCI technologies is therefore needed to clarify research progress, methodological features, and remaining challenges.  Progress   Advances in non-invasive continuous motor control BCIs are reviewed from four closely related aspects: control paradigms, decoding algorithms, applications, and performance evaluation. At the paradigm level, motor imagery, steady-state visual evoked potentials, P300, and hybrid paradigms have been studied to support continuous control through sustained intention modulation, dynamic stimulus encoding, and hierarchical or shared-control strategies. For decoding algorithms, two main frameworks are identified: motion parameter mapping and motion parameter regression. Motion parameter mapping generates continuous output by temporally integrating discrete classification results or mapping them to velocity or state variables, whereas motion parameter regression directly establishes relationships between Electroencephalogram (EEG) features and continuous kinematic parameters. Recent studies increasingly incorporate nonlinear models and deep learning methods to improve robustness under the non-stationary nature of EEG signals. At the application level, non-invasive continuous control has progressed from two-dimensional cursor tasks to more practical scenarios, including wheelchair navigation, robotic arm manipulation, unmanned systems, and virtual or augmented reality environments. Existing studies also assess continuous control performance using both objective and subjective indicators, including trajectory error, task success rate, information transfer rate, workload, and user experience, reflecting varied experimental designs and control aims.  Conclusions  Existing studies show that non-invasive BCIs can support continuous motor control. However, current research remains at a stage in which multiple methods coexist without a unified framework. At the paradigm level, available approaches differ in their ability to elicit and sustain continuous motor intention reliably. For decoding algorithms, both motion parameter mapping and motion parameter regression are limited by the non-stationary nature of EEG signals, which affects robustness, generalization, and long-term stability. At the application level, many studies remain restricted to specific tasks and controlled environments, and the transfer of continuous control strategies to complex real-world scenarios still requires further validation. Moreover, the lack of standardized evaluation protocols hinders direct comparison and systematic optimization across studies.  Prospects   Future research should improve the stability and reliability of continuous control paradigms, enhance decoding robustness under realistic EEG conditions, and strengthen the match between control strategies and application requirements. Unified evaluation frameworks that integrate objective and subjective indicators should also be established to support methodological convergence and fair comparison. With continued progress, non-invasive continuous motor control BCIs are expected to play a growing role in assistive technologies, rehabilitation systems, and advanced human-machine interaction.
Datasets
TTSPD: A Multimodal Traffic Scene Perception Dataset Integrating Tire Data
YING Zongchen, GUI Lin, YANG Jiahan, ZHANG Fangwei, WANG Junfan, DONG Zhekang
2026, 48(7): 2792-2804.   doi: 10.11999/JEIT260022
[Abstract](607) [FullText HTML](425) [PDF 3907KB](77)
Abstract:
  Objective  With the rapid development of Intelligent Transportation Systems (ITS) and autonomous driving technologies, accurate traffic environment perception is a fundamental prerequisite for vehicle safety and decision making. Current perception frameworks primarily rely on high-resolution cameras and LiDAR sensors. Although these sensors provide rich information, they create severe challenges across the Perception-Storage-Calculation pipeline. High acquisition costs limit large-scale deployment. In addition, the massive data volume produced by high-dimensional sensors places heavy pressure on onboard storage and computational resources, often exceeding the power and thermal budgets of vehicle-grade edge platforms. These constraints motivate the exploration of alternative sensing paradigms that are cost-effective, compact, and computationally efficient while maintaining reliable perception accuracy. In response, the present study shifts the perception perspective from conventional external sensors to the tire-road contact interface, where abundant physical interaction information naturally exists. The objective is to construct a novel multimodal dataset, termed the Tire-integrated Traffic Scene Perception Dataset (TTSPD), which combines internal tire dynamics with external visual observations. This dataset is used to examine whether low-dimensional tire sensing data can complement or partially substitute high-dimensional visual data for accurate road surface classification. The study also aims to establish a new data morphology that balances perception performance and system efficiency for future intelligent vehicles.  Methods  To construct a high-quality and practically usable multimodal dataset, an integrated hardware-software acquisition framework is developed. From a hardware perspective, a specialized sensing system is designed by coupling tire-mounted multi-parameter sensors with a vehicle-mounted camera. To ensure reliable operation under the harsh mechanical conditions of a rotating tire, sensing nodes are encapsulated using a rubber-based composite material that provides mechanical protection and long-term stability. Wireless transmission is implemented using Bluetooth Low Energy (BLE) 5.0 with an adaptive frequency-hopping mechanism, enabling low-power and reliable communication during high-speed rotation. During data acquisition, the system synchronously collects six types of internal tire signals, including radial acceleration, tire temperature, and tire pressure, producing approximately 1.8 million sampling points. In parallel, a dashboard-mounted camera records high-resolution traffic scene images totaling 309 GB across four representative road surface conditions. To address the heterogeneity between high-frequency one-dimensional tire signals and two-dimensional visual data, a timestamp-based association strategy is adopted to achieve scene-level temporal alignment rather than strict frame-by-frame correspondence. Sensor sequences and image segments are grouped according to shared temporal windows and driving scenarios. This approach ensures semantic and temporal consistency at the scene level. The alignment strategy reflects practical deployment conditions and forms the basis of the final TTSPD dataset for multimodal fusion research.  Results and Discussions  The effectiveness of the proposed TTSPD is evaluated through comprehensive road surface classification experiments using mainstream deep learning models. Initial experiments based solely on visual data demonstrate strong baseline performance, with classification accuracies ranging from 88.50% to 93.75% (Table 7). These results confirm the quality and diversity of the visual modality in the dataset. The primary contribution of this study is the quantification of efficiency gains enabled by tire-based sensing. Comparative experiments progressively reduce the amount of visual data while integrating low-dimensional tire signals, particularly radial acceleration (Table 9). The results show that the multimodal model achieves approximately 95% of the full-data baseline accuracy while using only about 38.75% of the original data volume. This reduction in data dependency produces significant system-level benefits. Storage requirements decrease by approximately 61.25%, and overall model training time decreases by about 54.10% (Fig. 8). These findings indicate that tire dynamics encode high-value physical features related to road texture and surface conditions that complement visual cues. The proposed dataset therefore supports the development of lighter perception pipelines without reducing recognition performance.  Conclusions  This study addresses the long-standing Perception-Storage-Calculation bottleneck in vision-dominated autonomous driving systems by proposing the TTSPD. Multi-parameter sensors are embedded within tires using rubber-based encapsulation, and stable wireless communication is achieved through BLE 5.0. A robust tire-camera data acquisition system is therefore established. The resulting dataset covers four common and safety-critical road surface types: cement, asphalt, damaged, and water-covered roads. It provides a comprehensive foundation for multimodal perception research. Experimental results show that combining low-dimensional tire sensing data with visual information significantly improves perception efficiency. Approximately 95% of peak classification accuracy is achieved using only about 38.75% of the original data volume. This result effectively reduces storage pressure and computational cost, reflected in a 61.25% reduction in data storage and a 54.10% reduction in training time. The TTSPD dataset therefore proposes a practical data morphology that supports efficient and high-performance perception under vehicle-grade computational constraints. It also provides valuable resources for the future development of ITS.
Image and Intelligent Information Processing
A Processing-In-Memory Neural Network Inference System Design for Infrared Gesture Recognition
SHI Xiangyang, LIU Jinchang, LU Qiulin, WANG Ziang, HAN Yongkang, GAO Gen, SUN Haoran, JIANG Xiaoyong, SHI Tuo, LI Qing, MIAO Jinshui
2026, 48(7): 2805-2814.   doi: 10.11999/JEIT260122
[Abstract](100) [FullText HTML](50) [PDF 3978KB](18)
Abstract:
  Objective  Infrared (IR) image recognition is widely used in security monitoring, autonomous driving, and human-machine interaction, where stable perception under low illumination and noisy backgrounds is required. However, IR images often exhibit low contrast, blurred edges, and weak texture representation, which reduce recognition accuracy. Conventional Deep Neural Networks (DNNs) further aggravate this issue in edge environments because the von Neumann architecture requires intensive computation, frequent memory-processor data transfer, and high power consumption. Memristor devices provide an alternative because they support in-memory computing. However, most existing implementations still depend on CPUs or GPUs to execute nonlinear operations such as activation layers, which reintroduces data transfer overhead. To address this issue, a nearly all-memristor neural network framework for IR image recognition is proposed. In this framework, computationally dominant linear operations are executed entirely on memristor arrays, whereas CPU participation is limited to the final activation step.  Methods  The proposed system maps two fully connected layers of a neural network directly onto memristor crossbar arrays. This mapping enables large-scale matrix-vector multiplications to be executed in memory with high parallelism and low energy consumption. Network weights are encoded as memristor conductance states, and inference is performed by applying voltage inputs and summing output currents. Intermediate nonlinear activations and final classification are computed on the host computer. Because these operations require minimal computation, the CPU overhead in both latency and energy remains negligible. Based on this design, an IR gesture recognition system is constructed to distinguish two hand gestures and a no-gesture state. System evaluation considers recognition accuracy, inference latency, energy consumption, and the amount of data transferred between the memristor arrays and the CPU. A CPU-only neural network is used as the baseline. The framework provides a practical approach for near all-hardware neural network computing while maintaining recognition accuracy and reducing energy consumption and data transfer for edge-scale IR applications.  Results and Discussions  The system is evaluated on a three-class IR gesture recognition task that includes two gestures and a no-gesture state. The memristor-based network achieves approximately 95% accuracy (Table 2), which is close to the 97% obtained with a GPU implementation (Table 2). This result indicates that conductance variation has limited effect on recognition performance. The average inference latency per frame is substantially lower than that of the GPU baseline, and the achieved frame rate satisfies real-time requirements. Power measurements indicate that the memristor array consumes only milliwatts, whereas the GPU requires approximately 25 W (Table 2). Although peripheral circuit consumption is not fully included, the results demonstrate the inherent energy efficiency of in-memory computing. Intermediate and final activation outputs are transferred to the CPU, which removes most memory-processor interactions. This reduction in data movement, combined with array-level parallelism, accounts for the observed improvements in latency and energy consumption. Overall, the framework maintains high recognition accuracy while significantly improving computational efficiency, which indicates strong potential for edge-scale IR recognition.  Conclusions  This study presents a nearly all-memristor infrared neural network framework in which two fully connected layers are executed on memristor arrays and intermediate nonlinear activations are processed on the host computer. When applied to a three-class IR gesture recognition task, the system achieves recognition accuracy comparable to that of a GPU platform while significantly reducing inference latency and energy consumption. By minimizing memory-processor data transfer and exploiting in-memory computing, the framework provides clear advantages for edge applications. The results confirm the feasibility of deploying memristor-based neural networks in practical infrared recognition systems. Future research will focus on integrating memristor-based activation functions, scaling the system to larger circuits, and extending the approach to more complex network architectures and datasets.
Explicit Discrimination-driven Automatic Unknown Class Clustering for Open-World Semi-Supervised Learning
SONG Jialun, DU Lan, CHEN Jian
2026, 48(7): 2815-2828.   doi: 10.11999/JEIT251291
[Abstract](643) [FullText HTML](339) [PDF 10457KB](91)
Abstract:
  Objective   Traditional target recognition is generally developed under the closed-set assumption, in which all test classes are assumed to be included in the training set. In real-world scenarios, however, unlabeled unknown classes are commonly encountered. Therefore, recognition systems should simultaneously recognize known classes and discover and cluster unknown classes. To address this challenge, a transductive Open-World Semi-Supervised Learning (OWSSL) method driven by explicit discrimination and automatic unknown-class clustering is proposed.   Methods   The proposed method is trained using a small set of labeled known-class samples and a large set of unlabeled test samples containing both known and unknown classes. It consists of two complementary modules. The Dynamic Known-Unknown Class Discrimination (DKUCD) module models the known-class boundary distribution using Extreme Value Theory (EVT) and progressively refines the distribution with high-confidence known-class samples during semi-supervised learning to improve known-unknown discrimination. The Neighbor Intersection-Over-Union Cluster Merging (NIOUCM) module automatically clusters high-confidence unknown-class samples by merging neighboring clusters according to their intersection-over-union relationships. The DKUCD and NIOUCM modules are optimized iteratively to improve discrimination and unknown-class clustering jointly.   Results and Discussions   Experiments conducted on the optical CIFAR-10 dataset and measured radar datasets demonstrate that the proposed method achieves accurate known-class recognition while effectively clustering unknown classes.   Conclusions   By explicitly discriminating between known and unknown classes and automatically estimating unknown-class clusters, the proposed method improves both known-class recognition and unknown-class clustering under open-world conditions.
Spatial Information-guided Diffusion for Domain Adaptation Semantic Segmentation of Remote Sensing Images
LIANG Yan, LI Junfan, SHAO Kai, HU Lin
2026, 48(7): 2829-2842.   doi: 10.11999/JEIT260031
[Abstract](728) [FullText HTML](304) [PDF 8177KB](68)
Abstract:
  Objective  Domain Adaptation Semantic Segmentation (DASS) is critical for remote sensing applications, including land-cover mapping, urban planning, and environmental monitoring. However, deep learning models often show severe performance degradation under domain shifts caused by imaging variation, geographic differences, and label-semantic heterogeneity. Conventional feature-alignment and generative adversarial network-based methods often fail to preserve semantic consistency. They are also sensitive to noisy supervision, especially when cross-domain gaps are large. This work aims to construct a robust DASS framework for semantically consistent image translation and reliable knowledge transfer.  Methods  A two-stage framework, termed Co-training Spatial-Guided DASS (CoSG-DASS), is proposed by integrating image translation and co-training. In the image-translation stage, a spatial information-guided latent diffusion model enhanced by ControlNet is designed. Semantic pseudo-labels and depth estimates are used as horizontal semantic and vertical spatial conditions to guide target-style image generation. To reduce the effect of noisy pseudo-labels, an Entropy-based Adaptive Guidance Intensity Module (EAGIM) is introduced. EAGIM estimates pixel-level confidence using information entropy and suppresses unreliable features. In the co-training stage, translated target-style images and unlabeled real target-domain images are used to train a segmentation model with a depth-guided segmentation head. Cross-entropy loss and adversarial loss are jointly used for optimization.  Results and Discussions  Extensive experiments are conducted on three cross-domain tasks. CoSG-DASS generates images that better match target-domain distributions. Quantitative results based on Fréchet Inception Distance (FID) show that the proposed method outperforms CycleGAN, UNI-Diff, and CRS-Diff in most settings (Table 1). Visual comparisons (Fig. 6) show that the method reduces edge blurring and category confusion. It also improves the separation of roads and vegetation and preserves small objects, such as vehicles. In the semantic segmentation stage, CoSG-DASS outperforms state-of-the-art domain adaptation methods. It improves mean Intersection over Union (mIoU) by 1.14%, 3.78%, and 2.49% on the cross-geographic task (Vaihingen IRRG→Potsdam IRRG), cross-imaging-mode task (Vaihingen IRRG→Potsdam RGB), and bidirectional label-semantic-heterogeneity tasks between DFC25 and LoveDA, respectively (Tables 24). Visual segmentation results (Fig. 7) confirm its strong boundary preservation and high accuracy in complex scenes. Ablation studies (Table 5) verify the contribution of the core components, including depth control, pseudo-label guidance, EAGIM, and the co-training strategy. Feature-distribution visualization based on Uniform Manifold Approximation and Projection (UMAP) further shows that CoSG-DASS reduces intra-class variation and increases inter-class separation after adaptation (Fig. 8).  Conclusions  CoSG-DASS alleviates domain shifts in remote sensing images through semantic-preserving diffusion-based translation and depth-guided co-training. It improves both image-translation quality and segmentation accuracy over existing methods. The proposed framework provides an effective solution for multi-source remote sensing interpretation. Future work will focus on extreme label-semantic heterogeneity and lightweight diffusion architectures.
A Social-Aware Ant Colony Optimization Algorithm with Reproductive Division of Labor for MCS Task Allocation
SHEN Xiaoning, SHE Juan, WANG Zhilong, LI Jiayuan
2026, 48(7): 2843-2853.   doi: 10.11999/JEIT260018
[Abstract](282) [FullText HTML](156) [PDF 1777KB](11)
Abstract:
  Objective  With the rapid development of handheld and wearable smart devices, Mobile Crowd Sensing (MCS) has become an efficient data collection paradigm. Effective task allocation can improve system efficiency, requester and participant satisfaction, and platform sustainability. Existing models often neglect task skill requirements, do not use participants’ social networks as auxiliary execution resources in emergencies, and overlook the effect of collaboration efficiency on team-task quality. To address these issues, this paper proposes a Social-Aware MCS Task Allocation model (SAMCSTA) with two objectives: maximizing total platform revenue and total task sensing quality. Social networks are used to build a two-layer collaboration framework of platform participants and social-network friends, which expands available execution resources and improves allocation flexibility. For complex tasks, participant sensing capability is quantified, and collaboration efficiency is introduced to optimize team composition.  Methods  This paper proposes a Multi-objective Ant Colony Optimization based on Reproductive Division of Labor (MACORDL) algorithm. The main innovations are as follows. First, the ant colony is divided into four collaborative subpopulations: queen ants, male ants, scout ants, and worker ants. Local enhancement, memetic crossover, knowledge transfer, and other search strategies are designed for these subpopulations to form a hierarchical collaborative search framework. Second, a statistical-learning-based mating selection strategy is designed to support intelligent transfer of elite genes. Third, the short-term contribution of each subpopulation is predicted from historical performance, which enables dynamic and adaptive allocation of computational resources. Fourth, a cooperative update mechanism for node pheromones and participant pheromones is designed to establish a dual-layer search guidance system.  Results and Discussions  The evaluation uses 8 synthetic instances and 4 real-world instances. Performance is measured by HyperVolume Ratio (HVR) and Inverted Generational Distance (IGD). The Wilcoxon rank-sum test at a significance level of 0.05 is used for statistical comparison. The results show that MACORDL achieves the best HVR and IGD on most instances (Table 2, Table 3). On average, MACORDL improves HVR and IGD by 16.41% and 18.04%, respectively, compared with the second-best algorithm. Visual comparisons further show that the Pareto front obtained by MACORDL has better convergence, distribution uniformity, and breadth (Fig. 4). Although its fine-grained local search can still be improved for a few large-scale instances, MACORDL shows stable performance and good scalability across different problem scales. It helps the platform obtain task allocation schemes with higher revenue and better sensing quality.  Conclusions  This paper studies the task allocation problem in MCS systems by considering interactions among platform participants and between participants and their social-network friends. A social-aware MCS task allocation model is established, and MACORDL is proposed to solve it. Comparative experiments on 8 synthetic instances and 4 real-world instances with different scales show that MACORDL outperforms six representative algorithms on most instances. It obtains allocation schemes and paths that yield higher total platform revenue and better task sensing quality, indicating good scalability. MACORDL uses multiple strategies to balance local exploitation and global exploration. However, the current model assumes that all tasks are released at the initial stage and that complete information is available. Participant privacy protection is also not considered. Future work will focus on MCS task allocation models in dynamic and uncertain environments and on privacy-preserving distributed optimization.
Multi-dimensional Spatio-temporal Feature Enhancement for Lip Reading
MA Jinlin, ZHONG Yaowei, MA Ruishi
2026, 48(7): 2854-2864.   doi: 10.11999/JEIT251111
[Abstract](568) [FullText HTML](339) [PDF 3147KB](44)
Abstract:
  Objective  Lip reading is a challenging yet important task in computer vision that aims to decode spoken language solely from visual lip movements. The task is difficult mainly because of inherent ambiguity in visual speech signals. On one hand, articulatory movements for different visemes can be highly subtle. For instance, lip displacement differences for confusable pairs such as /p/-/b/ and /m/-/n/ may be as small as 0.3~0.7 mm. These fine-grained spatial variations often fall below the effective resolution limits of conventional 3D convolutional neural networks. On the other hand, natural coarticulation in speech introduces temporal ambiguity, as mouth shapes may transiently blend multiple phonemes and make distinct visual units difficult to separate. These challenges are further aggravated by real-world factors such as uneven lighting and substantial inter-speaker differences in articulation. Current lip-reading models therefore often show limited ability to capture discriminative spatiotemporal features, which leads to suboptimal performance, particularly for phonemes with minimal visual differences. To address these issues, a robust lip-reading framework is developed to capture and exploit fine-grained spatiotemporal dependencies and improve recognition accuracy under diverse and realistic conditions.  Methods  To address the above limitations, a novel lip-reading framework, termed the Multi-dimensional Spatio-Temporal Enhancement Network (MSTEN), is proposed. The framework is designed to strengthen spatial and temporal representations through integrated attention mechanisms and advanced residual learning. It contains three core components that collaboratively model dependencies between spatial and temporal features, which are often insufficiently used in conventional architectures. The first component, the Self-adjusting Spatio-temporal Attention (SaSTA) module, adopts a self-adjusting mechanism that operates simultaneously across the height, width, and temporal dimensions. Query, key, and value tensors are generated through 1×1×1 3D convolutions, flattened across spatial and temporal dimensions, and used to compute attention weights by multiplying the query tensor with the transposed key tensor, followed by softmax normalization. The resulting attention map is multiplied by the value tensor and then combined with the original input through learnable parameters and a residual connection, thereby preserving contextual information and producing globally enhanced features. The second component, the Three-dimensional Enhanced Residual Block (TE-ResBlock), improves spatiotemporal feature extraction through temporal shift, multi-scale convolution, and channel shuffle. The temporal shift operation moves one quarter of the feature channels along the time axis to fuse information from adjacent frames without adding parameters. Multi-scale convolution adopts parallel branches with kernel sizes of 3×3, 3×1, 1×3, and 1×1 to capture features at different receptive fields. The outputs are concatenated and processed by channel shuffle to improve information exchange across groups, and four TE-ResBlocks are stacked to achieve progressive feature refinement. The third component, the Multi-dimensional Adaptive Fusion (MDAF) module, integrates spatial, temporal, and channel information through three submodules. These are a Channel Enhancement Module (CEM), which recalibrates features through max pooling, temporal convolution, and sigmoid activation, a Spatial Enhancement Module (SEM), which expands the receptive field through identity mapping and both standard and dilated convolution, and an Adaptive Temporal Capture Module (ATCM), which emphasizes dynamic movements through frame-difference features and temporal weight maps. The MDAF modules are inserted between TE-ResBlock stacks for iterative refinement. Finally, features extracted by the MSTEN front end are fed into a Densely Connected Temporal Convolutional Network (DC-TCN) back end, which consists of four blocks, each containing three temporal convolutional layers with dense connections, to model long-range phonological dependencies effectively.  Results and Discussions  The proposed framework is evaluated comprehensively on the widely used LRW and GRID datasets. The LRW dataset contains more than 500 000 video clips from over 1 000 speakers. The GRID dataset contains video clips from 34 speakers, each providing 1 000 utterances, with a total duration of 28 h. The proposed model achieves an accuracy of 91.18%, which is an absolute improvement of 2.82 percentage points over a strong ResNet18 baseline, demonstrating its substantial effectiveness. Ablation studies are further conducted to analyze the contribution of each key component. The results show clearly that each proposed module yields a meaningful performance gain. Specifically, the SaSTA module alone improves accuracy by 2.09%, which confirms the key role of global spatiotemporal attention. The TE-ResBlock increases accuracy by 1.73%, which verifies its effectiveness in multi-scale local feature extraction and inter-frame information fusion. The MDAF module provides a further 1.74% improvement, which highlights the benefit of adaptive multi-dimensional feature fusion, as shown in Table 2.  Conclusions  This study advances lip reading through the proposed MSTEN front-end network. The framework is built on three main contributions. First, the SaSTA module proposes an effective mechanism for global context aggregation and performs multi-dimensional feature weighting across height, width, and temporal sequences. Second, the TE-ResBlock addresses central challenges in spatiotemporal modeling through the combined use of temporal displacement, multi-scale convolution, and enhanced channel interaction. Third, the MDAF module enables deep and coordinated integration of spatial, temporal, and channel information. Together, these components improve model performance substantially and achieve accuracies of 91.18% on the challenging LRW dataset and 97.82% on the GRID dataset. Ablation studies further confirm the individual and combined effectiveness of the proposed components. Future work will examine the extension of this framework to audio-visual speech recognition under noisy conditions and the development of domain adaptation strategies to improve robustness in low-resolution or resource-constrained scenarios.
Network Metric System and Scenario-Differentiated Analysis Driven by LLM Literature Mining
XU Qikun, LIU Yaxi, HAN Shuxian, ZHANG Huifeng, HUANGFU Wei
2026, 48(7): 2865-2875.   doi: 10.11999/JEIT251120
[Abstract](506) [FullText HTML](242) [PDF 4177KB](54)
Abstract:
  Objective  Network metrics provide the foundation for network design, operation, and optimization. Existing studies primarily focus on individual scenarios or representative metrics and lack unified extraction rules and reproducible workflows for large-scale, cross-scenario metric analysis. To address terminology ambiguity, scenario heterogeneity, and the quantification of complex metric relationships, this study proposes a reproducible domain-specific literature mining framework based on a Large Language Model (LLM). The framework automatically extracts and standardizes network metrics, annotates application scenarios, quantifies inter-metric relationships, and establishes a Service-Multiplexing-Versatility (SMV) analytical framework. Rather than providing a complete set of metric calculation methods, the SMV framework serves as a conceptual model for guiding multi-objective tradeoffs in network architecture design and lifecycle management.  Methods  An automated literature mining framework based on a multi-agent LLM architecture is developed (Fig. 1). A dataset comprising 583 articles published in IEEE/ACM Transactions on Networking during 2023~2024 is analyzed. The framework consists of three specialized agents. A terminology normalization agent maps aliases and synonymous expressions to standardized metric names. A scenario annotation agent assigns primary application scenario labels using high-information-density sections of each article. A correlation mining agent identifies the semantic direction and strength of relationships between metric pairs and quantifies these relationships as signed correlation coefficients ranging from –1 to +1. The reliability of the mining results is evaluated through dual-LLM cross-validation and manual sampling review (Fig. 2).  Results and Discussions  The proposed framework extracts 3 978 independent network metrics, of which 138 appear in more than 1% of the analyzed articles (Fig. 3). The metric frequency distribution exhibits a pronounced heavy-tailed distribution, with throughput (79.1%), end-to-end delay (74.6%), and packet error rate (59.5%) representing the most frequently studied metrics (Table 1). The core metric sets show strong scenario dependence (Fig. 4). For example, data center networks primarily emphasize throughput, end-to-end delay, and flow completion time, whereas Internet of Things (IoT) applications additionally prioritize energy consumption and network lifetime (Table 2). Furthermore, scenario-specific correlation matrices reveal markedly different coupling patterns among metrics (Figs. 5 and 6). In data center networks, throughput is strongly negatively correlated with flow completion time and queueing delay, reflecting the fundamental tradeoff associated with congestion control. In edge computing networks, end-to-end delay is negatively correlated with resource utilization, indicating the balance between real-time task offloading and resource utilization.  Conclusions  The strong coupling between network metrics and application scenarios indicates that future network architectures should be evaluated from a multidimensional perspective. Based on the extracted scenario-specific metric relationships, this study proposes the SMV analytical framework (Fig. 7). By jointly considering differentiated service quality requirements (Service), physical infrastructure cost and resource reuse (Multiplexing), and adaptive reconfiguration capability for emerging services and application scenarios (Versatility), the framework provides a theoretical basis for adaptive resource orchestration in AI-native networks. Future work will extend the current static literature mining pipeline into a continuously updated network metric knowledge base and further validate the engineering applicability of the SMV framework in programmable and multimodal networks.
UWF-YOLO: A Lightweight Framework for Underwater Object Detection via Redundant Information Optimization
HOU Guojia, MA Jiaqi, WANG Yuechuan, HUANG Baoxiang, LI Kunqian
2026, 48(7): 2876-2886.   doi: 10.11999/JEIT251129
[Abstract](879) [FullText HTML](682) [PDF 3089KB](102)
Abstract:
  Objective  The rapid development of underwater imaging technology has increased the significance of underwater object detection for resource exploration and environmental monitoring. Complex underwater environments often degrade image quality through color casts, haze-like effects, and non-uniform illumination. These factors reduce the performance of existing vision-based object detection algorithms, particularly for small objects, and often lead to missed detections and false positives. In addition, current deep learning-based underwater detection models face difficulty balancing detection accuracy and lightweight design under limited computational resources. Therefore, efficient underwater object detection methods are required for water-related vision tasks. Such methods support marine resource exploration, ecological monitoring, underwater robotics, and perception systems for autonomous underwater vehicles.  Methods  A lightweight framework based on redundant information optimization is proposed for underwater object detection. Specifically, a lightweight underwater object detection network, termed UWF-YOLO, is designed based on redundant information optimization. First, the C2f module is reconstructed using the FasterNet Block to optimize both the backbone and neck networks. A feature channel selection mechanism is integrated to reduce redundant feature representations. Furthermore, redundant convolutional features in the conventional YOLO neck limit adaptation to underwater environments. Therefore, Ghost Convolution is introduced to generate Ghost feature maps and improve the multi-scale feature fusion capability of the neck network. Next, parameter sharing is achieved by replacing the original detection head with a redundant optimization group detection head (RRG-Head) based on group convolution, which reduces computational cost. Finally, a structured channel pruning strategy is applied to identify inter-layer dependencies in the computational graph and bind pruning units. Combined with LAMP weight magnitude score normalization to evaluate channel importance, low-contributing groups are pruned and subsequently fine-tuned to compress the network size. In addition, existing underwater detection datasets usually contain monotonous scenes, and the objects are typically small and densely distributed. To address this limitation, an underwater object detection dataset with complex scenes, termed CSUOD, is constructed by collecting real-world underwater images from various websites and platforms. Manual annotation and resolution normalization are then performed to ensure dataset consistency. CSUOD is designed for challenging underwater environments characterized by color casts, haze-like effects, and non-uniform illumination. A total of 1 135 images containing six object categories are manually selected and annotated.  Results and Discussions  Extensive experiments are conducted on three public underwater object detection datasets, namely DUO, RUOD, and TrashCan, and several widely used detection methods are compared. The proposed model is evaluated against mainstream detectors, including YOLOv5s, YOLOv7-tiny, YOLOv8s, YOLOv9-tiny, and Deformable DETR. In terms of computational complexity, the proposed method reduces FLOPs, model size, and parameters by 60.4%, 77.3%, and 78.4%, respectively, compared with the baseline model. Furthermore, the proposed method outperforms YOLOv9-tiny with comparable parameters by 0.3%, 2.3%, and 3.4% in mAP on the three datasets. Additional comparative experiments on the constructed CSUOD dataset also demonstrate improved performance and stable detection capability in complex underwater environments. Qualitative visualization results further demonstrate the robustness and detection stability of the model under various underwater degradations, including haze-like effects and non-uniform illumination.  Conclusions  Quantitative and qualitative experiments on multiple datasets validate the effectiveness and robustness of the proposed method. The proposed framework achieves superior detection performance in complex underwater environments and reduces missed detections and false positives caused by background interference. Experimental results indicate that the proposed UWF-YOLO achieves significant model lightweighting while maintaining detection accuracy comparable to benchmark models. This balance between detection accuracy and low computational cost makes the framework suitable for underwater devices with limited resources. The proposed method also shows strong potential for practical applications such as marine ecological monitoring, underwater resource exploration, and perception systems for autonomous underwater vehicles. It provides a reliable technical foundation for real-time applications, supports integration into embedded platforms, and enables real-time perception and decision-making under different underwater conditions. In addition, the constructed CSUOD dataset helps address the limitations of existing underwater detection datasets and supports further research in underwater object detection. Future work will extend this framework to multi-modal perception systems and larger-scale datasets, enabling adaptive models for dynamic underwater scenarios and supporting broader applications in intelligent ocean observation and autonomous navigation.
A Multimodal Sentiment Analysis Model with Multi-source Knowledge guided Visual Confidence Perception
PENG Juhong, ZHANG Zhi, LIU Peng, GE Wenhui, LIU Chen, LIAO Lingxin, ZHANG Kai
2026, 48(7): 2887-2897.   doi: 10.11999/JEIT260063
[Abstract](368) [FullText HTML](226) [PDF 2881KB](28)
Abstract:
  Objective  Multimodal sentiment analysis is often affected by visual noise from complex environments, image-text sentiment inconsistency, and imbalanced modality contributions. When all modalities are treated without distinction, visual noise can degrade model performance. A robust mechanism is therefore needed to evaluate visual confidence and filter redundant visual information.  Methods  A Multimodal Sentiment Analysis Model with Multi-source Knowledge-guided Visual confidence Perception (MKVP) is proposed (Fig. 1). A multi-source knowledge guidance matrix is constructed using syntactic-dependency, sentiment-intensity, and aspect-focused operators (Fig. 2). Guided by this matrix, the Visual Confidence Perception (VCP) module measures semantic affinity and dynamically suppresses irrelevant visual noise (Fig. 3). A dual-stream parallel interaction module is then used to support deep cross-modal alignment, and a global gated fusion mechanism further adjusts the fusion weights of different modalities.  Results and Discussions  Extensive experiments are conducted on the MVSA-Single, MVSA-Multiple, and HFM datasets. The proposed MKVP model achieves accuracy and F1 scores of 77.56% and 76.70%, 72.72% and 70.66%, and 87.26% and 86.78%, respectively. Compared with the baseline models, the accuracy and F1 score are improved by 2.45% and 3.68%, 2.19% and 2.21%, and 1.83% and 1.91%, respectively (Table 3). Ablation studies show that each component contributes to performance, especially the VCP module, which filters visual noise and improves feature quality (Table 5). Feature-space visualization further confirms that the VCP module refines semantic representations by promoting clearer clustering of samples with the same sentiment polarity (Fig. 4). Case studies on mismatched image-text samples also verify the ability of the model to resolve cross-modal semantic conflicts (Table 6). Model-complexity analysis shows that MKVP maintains high computational efficiency and low inference latency (Table 8).  Conclusions  The proposed MKVP framework reduces the effects of visual noise and image-text sentiment inconsistency in multimodal sentiment analysis. By using multi-source knowledge to guide visual confidence perception and combining dual-stream interaction with dynamic gated fusion, the model learns robust sentiment representations from noisy multimodal data. This method provides an efficient and reliable solution for complex social media scenarios.
Multipath Scheduling Algorithm for UAV Video Streaming
CAO Changlong, LI Lingzhi, SHI Lianmin, ZHAO Qingyue
2026, 48(7): 2898-2908.   doi: 10.11999/JEIT260002
[Abstract](501) [FullText HTML](340) [PDF 2955KB](40)
Abstract:
  Objective   With the rapid growth of the low-altitude economy, Unmanned Aerial Vehicle (UAV) technology has been widely used in emergency rescue, disaster monitoring, urban security, and other applications. In these scenarios, stable, low-latency, and high-fidelity video backhaul is critical for task execution. Multipath transport protocols can improve Quality of Experience (QoE) through bandwidth aggregation, providing an effective basis for UAV video streaming. However, under dynamic and heterogeneous network conditions, the performance of multipath transport protocols depends strongly on the design of multipath scheduling algorithms. Existing heuristic schedulers use predefined rules to reduce head-of-line blocking and inter-path load imbalance, but their adaptability remains limited in highly dynamic environments. Learning-based schedulers can learn the mapping between network states and scheduling rewards from real-time feedback, enabling adaptive performance optimization. However, most existing learning-based schedulers are designed for general network scenarios. They are not optimized for UAV networks, and their ability to guarantee QoE has not been fully validated. A multipath scheduling algorithm tailored to UAV video streaming is therefore needed to better exploit the performance potential of multipath transport protocols.  Methods   To address the dynamic and heterogeneous challenges of UAV video streaming, this paper proposes NeuroFly, a multipath scheduling framework based on the NeuralUCB algorithm. In NeuroFly, multipath traffic scheduling is formulated as a Contextual Multi-Armed Bandit (CMAB) problem. The context space is constructed by integrating path state information, video encoding features, and UAV mobility parameters, which jointly characterize the current transmission environment. In the action space, a frame-priority-driven redundant transmission mechanism is proposed. Video frames are assigned different frame priorities according to decoding dependencies, and differentiated redundancy strategies are used to improve the probability of successful video-frame delivery. A multi-objective reward function is further designed to guide policy learning and support adaptive optimization under dynamic and heterogeneous network conditions. In addition, a context monitoring mechanism is integrated into NeuroFly to handle abrupt environmental changes caused by high UAV mobility. This mechanism detects context distribution shifts and triggers a two-stage restart strategy. A soft restart is activated when gradual context drift is detected, removing outdated historical experience. A hard restart is performed under abrupt context changes by clearing the experience replay buffer and reinitializing model parameters, allowing learning to restart under a new distribution.  Results and Discussions   The proposed NeuroFly framework is evaluated in both simulation and field environments. First, Mininet-WiFi is used to simulate realistic UAV network environments and evaluate overall QoE performance. The results (Fig. 4) show that, compared with state-of-the-art heuristic and learning-based schedulers, NeuroFly achieves broad performance gains by fully using aggregated multipath bandwidth. Specifically, the 99th-percentile latency is reduced by 19.9%~51.0%, the average video frame rate is increased by up to 24.6%, image structural similarity is improved by up to 49.2%, and the buffering time ratio is reduced by 13.4%~77.6%. These results demonstrate the strong ability of NeuroFly to guarantee QoE. Field experiments (Fig. 6) further confirm that NeuroFly provides favorable optimization in real UAV operation scenarios. Compared with mainstream transport solutions widely deployed in production environments, NeuroFly achieves better real-time transmission performance and shows strong practical applicability for future large-scale UAV deployment.  Conclusions   This paper addresses network dynamics, path heterogeneity, and time-varying transmission conditions in UAV video streaming over multipath transport protocols. An intelligent multipath scheduling framework, NeuroFly, is proposed based on the NeuralUCB algorithm. In this framework, multipath traffic scheduling is modeled as a CMAB problem. Through the design of the context space, action space, and multi-objective reward function, online learning and adaptive optimization of traffic allocation policies are achieved. To further improve robustness under severe environmental changes, a lightweight context monitoring mechanism is introduced to detect context distribution drift and restart the learning process when needed. Systematic evaluations are conducted on both simulation platforms and real UAV operation environments. The simulation results show that NeuroFly achieves consistent improvements across QoE metrics compared with state-of-the-art heuristic and learning-based schedulers. The field results further indicate that NeuroFly provides reliable guarantees in actual UAV operation scenarios when compared with mature solutions that have been widely deployed in production environments. These results validate the practicality, robustness, and engineering feasibility of NeuroFly, and suggest its potential for large-scale deployment in UAV applications that are sensitive to real-time video quality, including emergency response, power inspection, agricultural monitoring, and logistics delivery.
Box Particle Filter δ-GLMB Algorithm for Multiple Maneuvering Group Targets Tracking
GAN Linhai, WANG Gang, LI Zhihui, SUN Wen, WANG Baotang
2026, 48(7): 2909-2918.   doi: 10.11999/JEIT251273
[Abstract](364) [FullText HTML](184) [PDF 1377KB](27)
Abstract:
  Objective  Targets that move in a coordinated manner or show similar motion patterns are commonly referred to as group targets. Dense group targets contain many closely spaced individuals and often suffer from poor measurement resolvability, severe measurement overlap, and frequent target disappearance and reappearance. These factors make it difficult to establish stable tracks for individual targets within the group. Such groups are therefore usually treated as a whole to jointly estimate the kinematic state of the centroid and the extended shape. To improve tracking accuracy and computational efficiency for multiple maneuvering group targets under nonlinear measurements, an Interacting Multiple Model Gamma Box Particle δ-Generalized Labeled Multi-Bernoulli (IMM-GBP-δ-GLMB) algorithm is proposed. Tracking efficiency under nonlinear measurements is improved using the Box Particle Filter (BPF). The likelihood function of the GBP algorithm is improved, and the IMM algorithm is introduced to enhance tracking of the extended shape and centroid kinematic state of group targets. Finally, the method is integrated with the GLMB filter to track an unknown number of multiple maneuvering group targets.  Methods  Existing algorithms mainly describe the area-based overlap between the predicted extended state of group targets and the measurement distribution, but they do not fully capture shape similarity. To address this limitation, the likelihood function of the BPF is modified. Geometric parameters, including the semi-major axis, semi-minor axis, and inclination angle, are incorporated into the likelihood function. This improves the modeling of similarity between the predicted extended state and the measurement distribution. The modification is particularly useful for maneuvering group targets, because the inclination angle of the extended shape changes frequently during maneuvering. Based on IMM modeling of group motion, a model index is added to the centroid kinematic state of each box particle. The model index and centroid kinematic state are jointly estimated in each iteration, allowing mode transitions of individual box particles to be tracked and further improving tracking accuracy. The improved IMM-GBP filter is then embedded into the labeled random finite set framework, and the IMM-GBP-δ-GLMB algorithm is derived for effective tracking of multiple maneuvering group targets.  Results and Discussions  Simulation experiments are conducted to compare the proposed IMM-GBP-δ-GLMB algorithm with the IMM Sequential Monte Carlo δ-GLMB (IMM-SMC-δ-GLMB) filter. The proposed algorithm maintains comparable estimation accuracy for the centroid state, extended state, measurement rate, and number of targets, while improving computational efficiency. In the given simulation scenario, the proposed algorithm achieves a 3.8-fold improvement in timeliness, with an approximately 8.5% reduction in tracking accuracy. In scenarios with two and three group targets, the average tracking time growth rate of the proposed algorithm is 96% of that of the IMM-SMC-δ-GLMB filter. This result indicates good temporal robustness as the number of group targets increases. Therefore, the proposed algorithm has strong practical value.  Conclusions  This paper addresses the tracking of multiple maneuvering group targets under nonlinear measurement conditions by proposing the IMM-GBP-δ-GLMB algorithm. The main contributions are as follows: (1) The likelihood function of the BPF is improved to strengthen the measurement of similarity between the target extended shape and the measurement distribution, improving the tracking accuracy of group target states. (2) A motion model label is assigned to each box particle, and transitions in the target motion state are tracked during filtering. This allows the filter to achieve higher tracking accuracy with fewer box particles and improves computational efficiency. (3) The IMM-GBP method is integrated into the δ-GLMB framework to obtain the final IMM-GBP-δ-GLMB filter, which realizes effective tracking of multiple maneuvering group targets.
A High-Performance Eye Tracking Method Based on Event Camera and Dual-Channel Differential Illumination
SONG Sishun, FENG Junchi, PU Chengyu, GUO Yu, LIU Shijie, HE Xin, CHEN Yuwei
2026, 48(7): 2919-2929.   doi: 10.11999/JEIT251162
[Abstract](560) [FullText HTML](293) [PDF 2584KB](31)
Abstract:
  Objective  Eye tracking has become an essential technology in human-computer interaction, medical diagnostics, cognitive neuroscience, and augmented and virtual reality applications. Traditional eye tracking systems, however, often suffer from two major limitations: low spatial accuracy and limited temporal resolution, especially during high-speed eye movements. These limitations hinder precise gaze estimation and reduce the reliability of real-time interactive systems. To address these challenges, an event camera is integrated with a dual-channel differential illumination strategy to improve the signal-to-noise ratio of corneal reflection events. The Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is introduced to achieve accurate localization of corneal reflection points. On this basis, the coordinates of the corneal reflection points are combined with Singular Value Decomposition (SVD) and the least-squares method to determine the center of corneal curvature, thereby significantly improving gaze direction estimation accuracy. This study provides a practical technical route for next-generation eye tracking systems and offers theoretical support for their use in complex interactive environments.  Methods  An event-camera-based gaze tracking method is proposed that integrates asynchronous eye-movement event data through a dual-channel differential illumination framework, thereby improving gaze direction estimation accuracy under high-speed and dynamic conditions. First, the event camera asynchronously captures brightness-change events with microsecond-level temporal resolution, which enables precise tracking of rapid eye movements. At the same time, the dual-channel differential illumination mechanism suppresses redundant reflections and improves the contrast of corneal reflection points. Second, the DBSCAN algorithm is used to process the event data, effectively removing noise and improving the spatial localization accuracy of corneal reflection features. Finally, a ray-tracing model is reconstructed using SVD and least-squares fitting to determine the center of corneal curvature, thereby enabling robust and high-precision gaze direction estimation. Experimental results on a biomimetic eye-movement dataset show that the proposed method achieves high temporal resolution, localization accuracy, and robustness in dynamic tracking scenarios.  Results and Discussions  Experiments show that the proposed method achieves a temporal resolution of 25 kHz (Fig. 6), which far exceeds that of conventional cameras. Differential illumination significantly improves the signal-to-noise ratio of corneal reflection events. The DBSCAN algorithm localizes corneal reflection points more efficiently than K-means, agglomerative clustering, mean shift, and OPTICS, and achieves accurate results within 10 ms without requiring predefined clusters (Fig. 8, Table 3). For gaze estimation, the proposed method maintains stable accuracy across sampling frequencies from 2 kHz to 25 kHz. At a 15° cone angle, the Mean Error (ME) and Root Mean Square Error (RMSE) are approximately 0.66° and 0.67°, respectively. At 25°, these values increase slightly to 0.87° and 0.90°, respectively (Table 4). Compared with existing state-of-the-art gaze tracking methods, the proposed approach shows better overall performance in both temporal resolution and accuracy (Table 5). Trajectory results (Fig. 9) show close agreement between the estimated and ground-truth gaze paths, and distribution analyses (Fig. 10) confirm that most errors remain below 1°.  Conclusions  A novel eye tracking method is presented that integrates an event camera with dual-channel differential illumination. The method achieves high temporal resolution (25 kHz), improves event signal quality, and reduces localization errors, yielding gaze estimation errors of less than 1°. The proposed approach provides a reliable technical route for next-generation high-performance eye tracking systems. Future work should address sensor noise modeling and computational optimization to further improve real-world applicability.
Defending against Deepfakes by Attribute-Aware Attack
GAO Fan, YAN Weidan, SHAO Wenze, ZHANG Dengyin
2026, 48(7): 2930-2942.   doi: 10.11999/JEIT260043
[Abstract](510) [FullText HTML](347) [PDF 5062KB](43)
Abstract:
  Objective  Deepfakes can cause serious personal and property damage when misused. To prevent forged images from spreading, existing methods often use adversarial examples to protect facial images from deepfake manipulation. However, traditional gradient-based attacks show limited generalization and low generation efficiency in black-box attack scenarios. Their performance is also weaker than that of current methods based on Generative Adversarial Networks (GANs), which are used to train cross-model adversarial examples. Although GAN-based methods support fast inference, their lack of perceptual constraints often makes the generated adversarial perturbations visually noticeable. The rapid development of deepfake models also raises higher requirements for the generalization ability of adversarial examples. Therefore, imperceptible and generalizable adversarial attack methods are needed for proactive deepfake defense.  Methods  To further improve the transferability and imperceptibility of adversarial examples generated by existing methods, this paper proposes an attribute-aware adversarial example generation method for deepfake defense. The proposed method generates imperceptible perturbations and improves cross-model generalization through a frequency-domain identity fusion mechanism. Specifically, it focuses on the foreground regions of facial images, uses attribute-aware salient segmentation masks to separate facial and hairstyle regions, and combines these masks with adaptive spatial-frequency attention-based perturbation generators to generate region-specific adversarial perturbations. This strategy improves the imperceptibility of adversarial examples and reduces the additional computational cost caused by global processing. From the perspective of data augmentation, this paper further uses phase swapping in the frequency domain to fuse identity-related features from reference face images. This design reduces perturbation overfitting and improves generalization performance.  Results and Discussions  The proposed method is trained and tested on the CelebA-HQ dataset using proxy models. Compared with existing proactive defense methods, the experimental results show that the proposed method generates adversarial examples with strong imperceptibility and cross-model defense capability. It achieves a high defense success rate against various proxy models. The average Peak Signal-to-Noise Ratio (PSNR) of forged outputs under adversarial perturbations is reduced to 16.79 dB, representing an improvement of approximately 1.87% over the second-best method. Defense performance against HiSD is improved by approximately 7.5% compared with the second-best method. Defense performance against AttGAN is approximately 12.7% higher than that of the second-best GAN-based defense method. Moreover, the Learned Perceptual Image Patch Similarity (LPIPS) metric shows that the adversarial perturbations have high imperceptibility.  Conclusions  This study proposes a facial attribute-aware attack method for deepfake defense. The method incorporates a frequency-domain identity fusion mechanism to increase the diversity of adversarial feature inputs. Adaptive spatial-frequency attention-based perturbation generators are also designed to extract local facial information and dynamically adjust adversarial features. These designs allow the method to preserve perturbation components that are both imperceptible and attack-effective, leading to strong cross-model generalization. Future work will focus on proactive deepfake defense methods with improved imperceptibility and generalization, especially in cross-model transfer attack scenarios.
Physiological Signal-driven QoE Optimization for Wireless Virtual Reality Transmission
WU Chang, PENG Mingyu, CHEN Yuang, CHEN Yiyuan, GUO Fengqian, QIN Xiaowei, LU Hancheng
2026, 48(7): 2943-2954.   doi: 10.11999/JEIT260067
[Abstract](346) [FullText HTML](200) [PDF 3578KB](33)
Abstract:
  Objective  Virtual Reality (VR) has become a transformative medium for immersive digital experiences because it can deliver high-resolution 360° video with ultra-low Motion-To-Photon (MTP) latency. However, its dependence on wireless transmission creates major challenges. Uncompressed data rates above 1 Gbit/(s·Hz) and latency thresholds below 20 ms place stringent demands on network infrastructure. In mobile scenarios, channel fluctuation and user mobility often compromise service continuity and cause abrupt resolution changes. Traditional Quality of Service (QoS) metrics, such as bandwidth, jitter, and packet loss, provide useful network-level information but cannot adequately reflect subjective user satisfaction. Existing Quality of Experience (QoE) models and Adaptive BitRate (ABR) algorithms often use symmetric metrics, such as Mean Opinion Score (MOS), and overlook the fact that users perceive quality deterioration and quality improvement differently. Sudden resolution downgrading has a stronger negative effect on immersion than the positive effect caused by resolution upgrading. This perceptual asymmetry is consistent with behavioral psychology but remains insufficiently addressed in current transmission schemes. In addition, the separation between Radio Access Network (RAN) resource provisioning and application-layer bitrate adaptation often causes mismatched optimization, video-quality oscillation, and resource underuse. To address these issues, this study establishes a quantitative link between physiological responses and resolution changes. It further develops a physiological signal-driven QoE framework integrated with Deep Reinforcement Learning (DRL) to support adaptive transmission, maximize immersion, and reduce the adverse effects of resolution fluctuation in resource-constrained wireless networks.  Methods  A two-stage method is adopted, including physiological signal analysis and joint optimization framework design. A controlled VR experiment is conducted to quantify the perceptual effect of resolution changes. Nineteen healthy subjects participate in a viewing task using an eye-tracking VR headset, a 32-channel wireless ElectroEncephaloGraphy (EEG) system, ElectroCardioGraphy (ECG) recording, and Galvanic Skin Response (GSR) sensors. The subjects view natural-scene videos in which the resolution levels, including 8 k, 4 k, 1 080 P, 720 P, and 480 P, switch randomly every 8 s. The collected EEG signals are preprocessed by independent component analysis and band-pass filtering. Event-Related Potential (ERP) components are analyzed, with emphasis on the N200 component in the temporal and occipital regions, which reflects visual processing and attention allocation. A Linear Discriminant Analysis (LDA) classifier is used to distinguish different response types. The analysis focuses on the asymmetry between resolution upgrading and downgrading, and on sensitivity to the magnitude of resolution jumps. Based on these physiological findings, a QoE model is formulated by adding penalty terms for resolution degradation and large-amplitude resolution switching. These penalties are weighted more strongly than upgrade rewards to represent user aversion to quality drops. The model is then integrated into an edge-computing environment through a dual-timescale DRL framework. The framework separates control into two cooperative agents: the Scheduling and Utility (SU) agent and the Resolution Scaling (RS) agent. The SU agent operates at the millisecond timescale and performs real-time wireless resource allocation. It uses a Gated Recurrent Unit (GRU) to extract temporal features from Channel State Information (CSI) and transmission history. It then dynamically allocates bandwidth to improve frame delivery success and maintain fairness under VR frame-deadline constraints. The RS agent operates at the frame timescale and determines the resolution of subsequent video frames. Its decision-making is guided by the physiological signal-driven reward function, which penalizes actions that may trigger negative physiological responses, such as sharp resolution drops, unless channel deterioration makes them necessary. Proximal Policy Optimization (PPO) is selected for both agents because of its stable learning behavior in continuous and discrete action spaces. Simulations are conducted using a 3GPP-based wireless channel module with user mobility, shadow fading, and path loss to create a dynamic network environment.  Results and Discussions  The physiological experiment and network simulations validate the proposed framework. In the physiological analysis, a clear N200 response is observed approximately 200 ms after resolution changes. The N200 amplitude is significantly larger during resolution downgrading than during resolution upgrading (p < 0.001), indicating that users are more sensitive to quality deterioration. Large resolution jumps, such as changes from 8 k to 1 080 P, also induce stronger neural responses and more concentrated occipital energy than minor adjustments. The LDA classifier achieves an average Area Under the Curve (AUC) of 74.12% across 19 subjects, confirming that neural responses contain discriminative information about the direction of resolution change. The GSR results support these findings. A dual-branch GSR feature extraction and classification model reaches an average AUC of 78.10% in distinguishing upward and downward switching events. By contrast, ECG signals do not show a stable effect under the current experimental setting and analysis granularity. Therefore, the subsequent QoE model is mainly constructed from EEG and GSR findings. In the network performance evaluation, the proposed physiological signal-driven DRL framework is compared with several baselines, including Proportional-Fair (PF) scheduling, equal resource allocation, and traditional congestion control represented by SCReAM. The training curves show that the dual-agent system converges and learns to coordinate capacity provisioning with resolution decisions. The SU agent smooths short-term channel fluctuation and provides a stable capacity basis, which enables the RS agent to make more reliable resolution decisions. Quantitative results show that the proposed scheme improves the average video resolution by up to 88.7% compared with the equal-resource baseline. More critically, the resolution switching frequency is reduced by up to 81.0%. This reduction is essential because frequent switching, especially downward switching, causes user discomfort, as demonstrated by the physiological analysis. By prioritizing long-term resolution stability and penalizing abrupt drops through the physiological signal-driven reward function, the proposed system reduces the “ping-pong” effect commonly observed in traditional ABR algorithms. Compared with schemes using different penalty weights, the proposed method achieves a better balance. It avoids overly conservative behavior under large penalties, which lowers the average resolution, and unstable visual quality under small penalties, which increases resolution fluctuation. The joint optimization also allocates resources preferentially to users with urgent frame deadlines or higher risks of perceptible quality degradation, while maintaining a frame delivery success rate above 99%.  Conclusions  This paper addresses the conflict between wireless-channel instability and the human need for visually consistent VR streaming. By adopting a physiological signal-driven approach, the asymmetric effect of resolution changes on user experience is quantified, which challenges the symmetric assumptions used in traditional QoE models. Integrating this physiological evidence into a dual-timescale DRL framework enables the RAN to go beyond throughput-oriented optimization. Wireless resource allocation supports stable application-layer adaptation, while application-layer demands guide resource scheduling. The proposed solution improves immersive experience by increasing average resolution and reducing the physiologically disruptive effects of sudden quality degradation. The reduction in resolution switching frequency by more than 80% shows that the system can shield users from network variability. This study also indicates the value of edge intelligence in making resource-allocation decisions based on human perception rather than network statistics alone. Future work should extend the QoE model by considering multisensory factors, such as MTP latency, cybersickness, spatial distortion, stalling, and audiovisual synchronization. Individual differences in physiological sensitivity should also be addressed through personalized modeling. For real-world deployment, privacy protection is essential. Federated learning and local edge updates may allow biometric data to be processed locally while supporting global policy optimization. This work provides a human-centric basis for immersive networking and shifts the focus from QoS to physiologically validated QoE.
A Point Cloud Slice-based UAV SLAM Method for 3D Reconstruction of Large Container Port Areas
HU Zhaozheng, ZUO Zhihang, XU Cong, TAO Qianwen, LIU Chao, MENG Jie
2026, 48(7): 2955-2968.   doi: 10.11999/JEIT251112
[Abstract](432) [FullText HTML](223) [PDF 12796KB](21)
Abstract:
  Objective  With the continuous development of port intelligence, the demand for digital management in container port areas has increased. In large container yards, Three-Dimensional (3D) reconstruction of the yard environment can be achieved using Unmanned Aerial Vehicle (UAV)-based Simultaneous Localization And Mapping (SLAM). However, container port areas contain many repetitive semantic structures. Traditional semantic matching methods therefore show low efficiency and limited accuracy. In addition, lanes between container yards form large feature-sparse regions during UAV-based 3D reconstruction, which can cause odometry degradation. Repetitive scene features also interfere with loop closure detection. To address these problems, this paper proposes a rapid feature extraction method based on point cloud slicing and further optimizes it according to the structural characteristics of container yards. A UAV point cloud slice-based SLAM method, termed Slice-SLAM, is proposed for high-precision 3D reconstruction of large container port areas.  Methods  To improve point cloud semantic extraction, a rapid point cloud slicing method is proposed. The principal direction is extracted rapidly, and the point cloud is divided into multiple layers to obtain multi-layer semantic point clouds efficiently. The slicing strategy is further optimized for container yard scenarios. Principal plane extraction is simplified using the gravity direction, and the elevation range of each container layer is obtained adaptively from point cloud density gradient changes. Multi-layer slice point clouds are then constructed. A progressive adaptive Light Detection And Ranging (LiDAR) odometry method based on slice point clouds is developed. Elevation slices are used to identify degenerate scenarios adaptively, and a layer-wise incremental slice matching and fusion strategy is used. This improves the accuracy, efficiency, and stability of LiDAR odometry. In addition, a factor graph optimization method that integrates slice point cloud information is designed. Fusion voting is performed on the matching results of multi-layer slice point clouds to remove erroneous matches and reduce the effect of repetitive structures on loop closure detection. Slice factors are then used to construct factor graph edges, which improves global optimization and supports efficient and stable 3D reconstruction.  Results and Discussions  The feasibility and effectiveness of the proposed method are verified in CARLA simulation scenarios and real-world tests at a large container port in Wuhan. First, comparisons with three semantic extraction algorithms, namely RANSAC, Region Growth, and 3DG_SEG, demonstrate the efficiency and accuracy of the proposed semantic extraction method. Second, estimated trajectories are compared with those obtained by two open-source LiDAR algorithms, FAST-LIO2 and Faster-LIO, confirming the advantages of the proposed odometry method. Finally, speed and confidence score are compared with those of six algorithms: ICP, NDT, GICP, Fast-GICP, Scan Context+ICP, and Quatro. The loop closure detection module of LIO-SAM is also integrated into FAST-LIO2, and the Scan Context module is integrated into Faster-LIO. The resulting estimated trajectories are compared with those of the proposed method, verifying the effectiveness of the proposed loop closure detection algorithm. The proposed method achieves high 3D reconstruction accuracy and is suitable for practical port operations.  Conclusions  The proposed method uses an efficient point cloud slicing technique and a multi-layer slice matching mechanism. Points within the same elevation range are defined as a slice point cloud, and the segmentation process is defined as point cloud slicing. This design enables efficient and robust 3D reconstruction in large-scale scenes with repetitive features. First, the LiDAR point cloud is aligned with the positive Z-axis using the gravity direction derived from the Inertial Measurement Unit (IMU). A sliding window records density gradient changes to determine the elevation range of each layer adaptively. This simplifies point cloud slicing and reduces the effects of non-standard containers and ground height variations on semantic extraction. Multi-layer slice information is then integrated into the odometry module to detect degenerate scenarios. Under normal conditions, progressive slice matching is used to initialize pose estimation. In degenerate scenarios, iterative Kalman filtering with increased IMU weighting is used. Finally, the fusion voting mechanism removes outliers from multi-layer slice matching results. The optimal match is used to initialize loop closure for global registration of container-region point clouds, enabling dual-stage loop closure detection and slice factor construction. By integrating slice point cloud information into factor graph optimization, the proposed method unifies point clouds in a common coordinate system and achieves efficient and robust 3D reconstruction.
Drug Response Prediction Based on Graph Topology Attention Network
XU Peng, XU Hao, BAO Zhenshen, ZHOU Chi, LIU Wenbin
2026, 48(7): 2969-2978.   doi: 10.11999/JEIT251099
[Abstract](562) [FullText HTML](375) [PDF 1295KB](54)
Abstract:
  Objective  A central goal in modern cancer research is to determine why patients respond differently to the same therapy. This requires computational tools that combine genetic information with drug properties to predict treatment results, which is essential for advancing personalized oncology. Although existing methods have improved cancer drug response prediction, effective drug feature extraction and integration of multi-omics data from cell lines remain challenging. To address these issues, Graph Neural Networks (GNNs) have been increasingly used to process drug molecular graphs. In this study, a model based on a graph topology attention network is proposed to extract features from drug molecular graphs, and an attention mechanism is used to integrate multi-omics data.  Methods  In this study, a drug response prediction method based on Graph Topology Attention Network (GTAT) is proposed. The model integrates topological graph information to predict drug responses in cell lines. Drug SMILES strings are used to generate two different drug representations, and multi-omics data are incorporated to characterize cell lines (Fig. 1). For drug feature extraction, SMILES strings are first parsed to construct molecular graphs, which are then processed by GTAT. This network captures both topological information at the molecular graph level and atom-level features, thereby generating structured molecular representations. At the same time, Extended Connectivity Fingerprints are computed from the same SMILES strings and transformed into continuous feature vectors through a Multi-Layer Perceptron (MLP). The graph-based drug representation and the fingerprint-based representation are then concatenated to form a comprehensive drug feature vector. For cell line representation, multi-omics data are processed through omics-specific neural networks. The resulting features are fused through multi-head self-attention mechanisms, which enable the model to capture contextual interactions across omics modalities and generate an integrated cell line representation. Finally, the drug and cell line features are combined and fed into an MLP classifier to predict drug response results. The proposed model effectively integrates heterogeneous biological data sources and significantly improves prediction accuracy through multimodal learning and attention-based feature fusion.  Results and Discussions  The proposed method achieves competitive performance on both the GDSC and CCLE benchmark datasets (Table 2). Specifically, on the GDSC dataset, the proposed approach outperforms all competing methods across all four metrics, including AUC, AUPR, F1-score, and Accuracy. In particular, the AUPR is improved by approximately 1.92% compared with that of the second-best method, MOFGCN, which demonstrates an advantage in handling class imbalance. On the CCLE dataset, the proposed method still achieves the best performance in terms of AUC and Accuracy. Although its AUPR and F1-score are slightly lower than those of GADRP, the differences are minimal, and the method shows stronger overall discriminative ability, as reflected by AUC. These results validate the effectiveness and strong generalizability of the proposed method in drug sensitivity prediction tasks. The variation in AUPR and F1-score across datasets may be attributed to inherent differences in sample size and class distribution. The limited size of the CCLE dataset, combined with its specific class imbalance, with an approximately 4:1 ratio of resistant to sensitive samples, may restrict the model’s ability to fully learn the underlying data distribution, particularly for minority classes. In contrast, the GDSC dataset shows greater heterogeneity and a more pronounced class imbalance, approximately 8:1, which increases prediction difficulty and leads to lower performance on certain metrics.  Conclusions  Accurate prediction of drug response in cell lines remains a central challenge in precision medicine and has important implications for accelerating drug development and advancing personalized treatment. However, construction of a highly accurate predictive model that effectively integrates multi-source biological information remains difficult because of the complexity of drug molecular structures and the inherent heterogeneity of cell lines. To address this issue, a cell line drug response prediction model based on GTAT is proposed. In this model, GTAT is used to extract molecular graph features of drugs, which are then fused with molecular fingerprint features. Meanwhile, multi-omics features of cell lines are integrated through an attention mechanism. Experimental results demonstrate that the proposed model achieves superior performance compared with existing state-of-the-art benchmark methods on the employed datasets. This study provides a new perspective for cell line drug response prediction. Certain limitations remain, including the use of only three types of omics features for cell line representation and the effect of sample size on predictive performance. Future work will focus on integrating more diverse omics features, applying pre-trained large-scale models, and promoting clinical translation for personalized medicine.
Household Appliance Plastics Identification by Fusing Multi-Level Feature Enhancement and Hierarchical Classification
CHONG Penghao, ZHENG Yunlong, YANG Aosong, GUO Mengci, LI Shifeng
2026, 48(7): 2979-2989.   doi: 10.11999/JEIT260084
[Abstract](510) [FullText HTML](321) [PDF 3437KB](57)
Abstract:
  Objective  Accurate plastic identification remains challenging in waste household appliance recycling under low-resolution spectral conditions. In practical recycling environments, plastics often have complex compositions, surface contamination, and aging effects, which increase classification difficulty. Black plastics are especially difficult to identify because their strong light absorption and spectral overlap in the Visible-Near Infrared (Vis-NIR) range reduce feature separability and degrade classification performance. Under these conditions, conventional single-stage classification models often fail to maintain stable accuracy. To address this problem, an automated identification method is proposed for low-dimensional multispectral feature spaces. The method aims to improve the discriminative capability of limited spectral information and enhance classification accuracy for complex plastic categories.  Methods  A compact Vis-NIR multispectral acquisition system based on the AS7265x sensor is used to collect 18-channel reflectance data in the 410~940 nm range. A handheld acquisition device with a controlled optical structure is designed to reduce environmental interference and ensure measurement consistency (Fig. 3). A total of 576 samples are collected from five typical household appliance plastics, including Acrylonitrile Butadiene Styrene (ABS), High-Impact PolyStyrene (HIPS), PolyPropylene (PP), Acrylonitrile Styrene copolymer (AS), and Polycarbonate/Acrylonitrile Butadiene Styrene (PC+ABS) blends. These samples are obtained from waste household appliances and are subjected to preliminary surface cleaning before spectral acquisition. To improve feature representation, a multi-level feature engineering strategy is adopted. This strategy integrates original spectral intensity features, nonlinear polynomial expansion features, and adjacent-channel ratio features to characterize both global and local spectral information. The nonlinear expansion enhances the representation of reflectance variations, whereas the ratio features capture local spectral-shape changes and reduce external disturbances. These features are combined into a 53-dimensional feature vector. Linear Discriminant Analysis (LDA) is then applied to enhance interclass separability. To address spectral overlap and class imbalance, a Hierarchical Joint Classifier (HJC) is constructed. HJC uses a two-stage classification framework. In the first stage, an XGBoost-based primary classifier performs coarse classification to separate easily distinguishable samples and group spectrally similar black plastics. In the second stage, a TabTransformer-based secondary classifier performs fine-grained classification of difficult samples (Fig. 6). This hierarchical design reduces classification complexity and improves discrimination for challenging categories. Model performance is evaluated using five-fold cross-validation and an independent test set. Accuracy, precision, recall, and F1-score are calculated from confusion matrices (Fig. 7). Comparative experiments are conducted with traditional machine learning methods, ensemble learning models, and deep learning approaches under different feature-processing strategies (Fig. 8, Fig. 9).  Results and Discussions  The proposed HJC achieves a classification accuracy of 97.4% in five-fold cross-validation and 93.1% on the independent test set (Table 4). Compared with single-stage classifiers and methods without feature enhancement, the proposed method provides higher performance and greater stability under low-resolution spectral conditions. Comparative results show that the proposed method outperforms baseline approaches, such as PCA combined with CNN, which achieves an accuracy of approximately 71.3% on the same dataset (Fig. 8). This improvement indicates that the proposed feature engineering strategy effectively strengthens the discriminative capability of low-dimensional spectral data. Combining LDA with feature engineering further improves class separability compared with conventional PCA-based methods. Confusion matrix analysis shows that misclassifications mainly occur between spectrally similar black ABS and black HIPS samples, whereas most other categories achieve high classification accuracy (Fig. 9). These results indicate that spectral overlap remains the main challenge under low-resolution conditions. The hierarchical classification strategy reduces this problem by focusing classification resources on difficult samples, thereby improving the overall generalization ability of the model. Overall, the proposed method shows robustness under practical conditions, including spectral noise, limited channel resolution, and material heterogeneity. These results indicate its suitability for real-world recycling applications.  Conclusions  A hierarchical classification method with multi-level spectral feature engineering is developed for plastic identification under low-resolution Vis-NIR conditions. Nonlinear and spectral-shape features are incorporated into a two-stage framework to improve the identification of spectrally similar materials. The results show stable accuracy across different plastic types. The method is suitable for automated sorting in waste household appliance recycling and can be extended to other material identification tasks with limited spectral information.
Aerial Spatio-Temporal Image Generation via Latent Diffusion Models
SHANG Yuying, HOU Yingyan, LIU Zinan, LU Wanxuan, HUANG Yuhong, WANG Yixiao, YU Hongfeng, FU Kun
2026, 48(7): 2990-3001.   doi: 10.11999/JEIT260165
[Abstract](513) [FullText HTML](311) [PDF 6044KB](38)
Abstract:
  Objective  Aerial Earth observation plays a pivotal role in environmental monitoring, disaster warning, and urban planning. However, constraints such as flight-platform endurance and mission-window timeliness often prevent acquired aerial imagery from fully characterizing the long-term evolution of the Earth’s surface. Although pre-trained latent diffusion models have shown strong potential for image generation, their application in aerial scenarios remains challenging because of the scarcity of high-quality temporal annotation data and semantic-visual misalignment caused by variable observation scales. To address these challenges, this paper proposes ASTIG, a training-free framework for Aerial Spatio-Temporal Image Generation. By leveraging the generative priors of pre-trained latent diffusion models and Large Language Models (LLMs), ASTIG provides a new paradigm for semantically controllable aerial spatio-temporal image generation.  Methods  ASTIG consists of three coordinated components. First, a dynamic semantic decomposition process is proposed to parse complex descriptions of aerial scene evolution into frame-level visual prompts, thereby compensating for the lack of temporal semantic annotations in existing aerial image-text datasets. Second, a Linguistic Binding (LB) strategy is proposed to establish explicit associations between key ground objects and their corresponding visual attributes within the cross-attention mechanism of the diffusion model, thereby improving the semantic response precision of the generated images. Third, a Temporal Anchor Attention (TAA) mechanism is incorporated. It uses dual reference frames to maintain subject stability and background consistency across the generated spatio-temporal image sequence, thus suppressing inter-frame temporal drift under training-free conditions.  Results and Discussions  ASTIG and the baseline methods are evaluated on 7 236 high-quality aerial spatio-temporal descriptions using six automated metrics, including subject consistency, background consistency, temporal flickering, motion smoothness, aesthetic quality, and imaging quality. Quantitative results (Tables 1 and 2) show that ASTIG outperforms the baseline methods in spatio-temporal image generation, with improvements of 3.91% in subject consistency and 4.57% in temporal flickering over the frame-prompt baseline. Qualitative comparisons (Fig. 4) further show its strong ability to model long-term surface evolution in aerial imagery. Ablation studies validate the individual effectiveness of the LB strategy and the TAA mechanism (Table 3 and Fig. 5). Sensitivity analyses of the intervention steps (Table 4 and Fig. 6) and binding strength (Table 5 and Fig. 7) further identify suitable parameter settings. Extension experiments from satellite perspectives (Figs. 8 and 9) also show that ASTIG has the potential to generalize beyond aerial platforms to broader Earth observation scenarios.  Conclusions  This paper proposes ASTIG, a training-free framework for aerial spatio-temporal image generation that addresses the scarcity of high-quality long-term temporal data and semantic-visual misalignment. By leveraging the generative priors of pre-trained latent diffusion models and LLMs, ASTIG integrates a dynamic semantic decomposition process, an LB strategy, and a TAA mechanism to improve temporal semantic construction, semantic response precision, and inter-frame consistency. Experimental results show that ASTIG outperforms existing baseline methods across multiple automated evaluation metrics, providing a new paradigm for aerial spatio-temporal image generation. As a training-free method, ASTIG is still limited by the prior knowledge of the backbone model. Future work will examine geometric correction and nadir-view prior constraints to better align the generated results with the physical properties of satellite imagery.
KE-HNS: Knowledge-Enhanced Personalized Recommendation Model with Hierarchical Noise Suppression
XIE Jun, WANG Dantong, ZHANG Bo, CHEN Guijun, LÜ Jiaqi, LUO Xiongyan
2026, 48(7): 3002-3014.   doi: 10.11999/JEIT260051
[Abstract](342) [FullText HTML](217) [PDF 5092KB](24)
Abstract:
  Objective  In the era of Big Data and Artificial Intelligence (AI), rapid information growth has increased the difficulty of filtering valuable content from redundant data. Personalized recommender systems are key tools for accurate information matching and resource allocation. Knowledge Graphs (KGs) can enrich user-item representations. However, current KG-based recommendation models still face weak noise suppression, coarse-grained user-interest modeling, and imbalanced use of heterogeneous information, which reduce recommendation accuracy. This paper proposes Knowledge-Enhancedpersonalized recommender model with HierarchicalNoise Suppression (KE-HNS), which integrates knowledge enhancement with hierarchical noise suppression. By combining graph representation learning and contrastive learning, KE-HNS addresses noise interference, fine-grained preference modeling, and multi-source information balance, thereby improving recommendation performance.  Methods  KE-HNS adopts a hierarchical noise-suppression paradigm. At the input stage, Input Noise Reduction (INR) is used to reduce noise from two sources. For user-item interactions, a learnable binary mask matrix is used to remove noisy edges. For KG denoising enhancement, triples are scored by importance, low-score triples are identified with a Bottom-K strategy, and noisy triples are masked. At the feature-fusion stage, Isolated Noise Suppression (INS) is used to preserve spatial independence by partitioning entity-attribute spaces according to relation type. This design limits high-order noise propagation and semantic contamination. At the representation-optimization stage, Comparative Noise Suppression (CNS) is implemented through contrastive learning to suppress irrelevant entity noise and strengthen robust semantic signals. To capture fine-grained user interests, Graph Convolutional Networks (GCNs) are used to enhance user representations from historical interactions and related entities. Adaptive weight layers further refine item representations by using entity attributes and relations. To balance heterogeneous information, a dual-view contrastive learning mechanism is constructed between the user-item view and the item-entity view. Positive and negative sample pairs are used to adaptively adjust the weights of different information sources. Finally, user and item representations are matched by inner product to generate the Top-K recommendation list.  Results and Discussions  KE-HNS is evaluated on three public datasets, Book-Crossing, MovieLens-1M, and Last.FM, through performance comparison, ablation experiments, denoising evaluation, case analysis, and complexity assessment. For Click-Through Rate (CTR) prediction, KE-HNS outperforms the best baseline models by 0.94%~1.01% in Area Under the Curve (AUC) and 0.43%~0.90% in F1-score (Table 3). For Top-K recommendation, its Recall@K is higher than those of most advanced methods across nearly all K values, with only a slight gap behind CG-KGR on Last.FM (Fig. 7). The ablation results show that all three denoising components contribute to the performance gains (Table 4). The denoising evaluation shows that KE-HNS effectively suppresses noise and maintains high prediction accuracy under noisy conditions (Fig. 8). The complexity analysis further indicates that the model remains feasible for practical deployment (Table 5).  Conclusions  This paper presents KE-HNS, a personalized recommendation model that combines knowledge enhancement with hierarchical noise suppression. By reducing noise interference and balancing collaborative filtering signals with knowledge-aware semantics, KE-HNS improves recommendation accuracy across multiple benchmark datasets. The model still has limitations in computational efficiency and depends on the coverage and completeness of the KG. Future work may focus on computational optimization and dynamic knowledge integration.
Data-driven Sliding-mode Disturbance-rejection Formation Control for Quadrotor UAV Swarms Under Uncertain Disturbances
LI Qianxiong, LU Xiaoqing
2026, 48(7): 3015-3026.   doi: 10.11999/JEIT260050
[Abstract](380) [FullText HTML](254) [PDF 3092KB](29)
Abstract:
  Objective  Quadrotor Unmanned Aerial Vehicle (UAV) cooperative formation can increase payload capacity and extend the operational range. However, quadrotor UAVs are highly nonlinear and underactuated systems. Differences in size and actuator hardware further weaken the effectiveness of model-based formation-control methods. Therefore, disturbance-rejection formation control is needed for quadrotor UAV swarms with unknown internal models and uncertain external disturbances.  Methods  To address the difficulty of precise modeling for quadrotor UAV swarm formation under uncertain disturbances, this paper proposes a data-driven sliding-mode disturbance-rejection formation control method. First, a data-driven formation-control model is established using the input and output states of each UAV and its neighboring UAVs. Then, an extended state observer and an integral sliding-mode formation controller are designed to estimate uncertain disturbances online and achieve robust formation control. Finally, stability analysis is conducted to derive sufficient conditions under which all UAVs achieve sliding-mode disturbance-rejection formation. The proposed method is verified through simulations and experiments under an unknown system model and uncertain disturbances.  Results and Discussions  The simulation results show that multiple quadrotor UAVs can maintain the desired formation geometry in a wind-disturbed environment (Fig. 4). The formation position error converges to within 0.1 m in 15 s and reconverges rapidly after a 7 m/s gust is applied (Fig. 6). The velocity curves also show rapid convergence among the UAVs (Fig. 5). The experimental results indicate that three UAVs can follow the trajectory of the virtual leader while maintaining the desired triangular formation (Fig. 17). The formation error is mostly kept within 0.1 m (Fig. 18). When the observation matrix fluctuates strongly between 10 s and 20 s, the corresponding formation error is relatively large. When the observation matrix curve becomes smoother between 20 s and 30 s, the formation error also decreases (Fig. 20). Compared with traditional model-based formation-control methods and existing data-driven methods, the proposed method reduces the formation error by 41% and shortens the formation response time by 40%.  Conclusions  This paper proposes a data-driven sliding-mode disturbance-rejection formation control method for quadrotor UAV swarms with unknown internal models and uncertain external disturbances. Under an unknown quadrotor UAV model and a 7 m/s wind disturbance, the proposed method keeps the formation error below 0.1 m. It also reduces the formation error by 41% and shortens the formation response time by 40% compared with traditional model-based formation-control methods and existing data-driven methods. Future work will study multilayer data-driven formation control for heterogeneous UAV-UGV swarm systems. It will also optimize computational cost and scalability in large-scale and complex application scenarios.
Wireless Communication and Internet of Things
One-step Reconstruction Diffusion Model-based Poisoning Attack on QoS-aware Cloud API Recommender Systems
TAN Zeyu, WANG Haoyuan, QI Mingyang, SUN Mengmeng, SHEN Limin, CHEN Zhen
2026, 48(7): 3027-3036.   doi: 10.11999/JEIT260115
[Abstract](335) [FullText HTML](139) [PDF 1284KB](18)
Abstract:
  Objective  In cloud computing, Cloud Application Programming Interfaces (cloud APIs) serve as key carriers for data output, capability reuse, and service delivery. They have become core elements in service-oriented software development and operation. With the rapid growth of cloud APIs, users often find it difficult to select suitable services from many functionally similar candidates. Quality of Service (QoS) is therefore used to differentiate cloud APIs by non-functional attributes. QoS-Aware cloud API Recommender System (QARS) plays an increasingly important role in guiding users toward suitable cloud APIs. However, existing studies mainly focus on improving recommendation accuracy and often ignore security risks caused by the economic value of cloud APIs and the openness of network environments. These risks are particularly evident in poisoning attacks. By injecting fake users, attackers can manipulate recommendation results and reduce the fairness and credibility of QARS. To address this threat from an attack-informed defense perspective, this paper analyzes the attack mechanisms of diffusion model-based poisoning methods and supports the design of targeted defense strategies.  Methods  The poisoning attack process and fake user profiles are first formally defined. Attack scale is then defined to flexibly simulate poisoning attacks under different settings. To analyze the attack principle of diffusion model-based methods, a One-step reconstruction Diffusion Model (ODM) is adopted, and a Preference guided one-step reconstruction Diffusion model-based Poisoning Attack framework (PDPA) is proposed. According to the collaborative principle that similar users tend to have similar preferences for cloud APIs, fake users generated by an attack method should have QoS values and cloud API invocation distributions similar to those of real users. This similarity allows fake users to exert collaborative influence and interfere with user preference modeling in QARS. PDPA is therefore designed to generate fake users that closely match real users. First, ODM separately models the QoS data and invocation distributions of real users. Unlike standard diffusion models, ODM avoids error accumulation caused by noise-dependent iterative denoising. It can generate fake-user invocation behavior similar to real-user behavior, which helps fake users exert effective collaborative influence. Then, to improve attack effectiveness, PDPA systematically selects fake users with invocation preferences for the target cloud API and assigns the maximum QoS value to the target item. This strategy strengthens the attack while reducing the disturbance caused by adding the target cloud API to fake-user invocation behavior, thereby improving stealthiness.  Results and Discussions  Experiments are conducted on the real-world WS-DREAM response-time QoS dataset. First, six recommendation methods, namely LR, MLP, DeepFM, AFM, DCN, and XSimGCL, are used as target recommender systems. Six baseline attack methods are used to simulate poisoning attacks. The results in Table 3 reveal the vulnerability of QARS to poisoning attacks. All attack methods reduce recommendation accuracy. PDPA achieves the best attack effectiveness in most experimental settings because it sufficiently models user invocation preferences, enabling fake users to exert stronger collaborative influence on QARS. Second, fake users generated by ODM and those generated by the standard diffusion model are compared in terms of F1 score and latent-space distribution. The results in Figure 2 show that ODM outperforms the standard diffusion model in stealthiness and produces a latent-space distribution closer to that of real users. Third, ablation studies are conducted for each module of PDPA. The results in Tables 4 and 5 verify that each module is necessary for attack effectiveness and fake-user stealthiness. Finally, Mean Absolute Error (MAE) and F1 score are compared under different attack scales to evaluate the effect of attack scale on attack effectiveness and stealthiness. The results in Figure 3 and Table 6 show that increasing the attack scale improves attack effectiveness but also increases the number of detected fake users.  Conclusions  This paper investigates the threat of poisoning attacks against QARS by analyzing the attack process and key attack parameters. The proposed PDPA simulates poisoning attacks on QARS and reveals their vulnerability. The results show the potential of diffusion models for poisoning attacks and verify the necessity of separately modeling QoS data and cloud API invocations. PDPA also clarifies how diffusion models generate fake users, providing a basis for future targeted countermeasures.
Rotatable-Antenna-Aided Near-Field Wideband Integrated Sensing and Communication System: Hybrid Beamforming Design
XU Hongbo, MO Minghui, XIN Wei, WANG Shuli, WANG Ji, LI Xingwang, ZHENG Le
2026, 48(7): 3037-3046.   doi: 10.11999/JEIT260023
[Abstract](588) [FullText HTML](239) [PDF 1972KB](79)
Abstract:
  Objective  Near-field wideband Integrated Sensing and Communication (ISAC) systems face two main challenges: pronounced near-field effects and wideband beam splitting. These effects reduce communication throughput and sensing reliability, particularly when fixed-orientation antenna arrays and phase-shifter-based beamforming architectures are used. Because such architectures provide limited spatial adaptability and frequency-independent phase control, the spatial-frequency degrees of freedom available in near-field wideband channels cannot be fully used. To address this issue, a Rotatable-Antenna-assisted near-field wideband ISAC architecture is investigated to improve the system sum rate under sensing constraints.  Methods  A near-field wideband ISAC architecture assisted by Rotatable Antennas (RAs) is proposed. By allowing the antenna boresight direction to be adjusted mechanically or electronically, additional angular degrees of freedom are provided at the element level, which enables more flexible spatial coverage and more accurate energy focusing. A True Time Delay (TTD)-based hybrid beamforming architecture is further adopted to provide frequency-dependent phase shifts and compensate for the frequency-independent property of conventional phase shifters. Consistent beam focusing across subcarriers is thus maintained, and wideband beam splitting is effectively suppressed. Based on a spherical-wave near-field channel model that incorporates propagation distance, angular information, and the orientation gain of RAs, a joint optimization problem is formulated to maximize the system sum rate under transmit power constraints, sensing power thresholds, and antenna rotation constraints. Because the resulting problem is highly non-convex, a Penalty-Based Fully Digital Approximation (PBFDA) algorithm is developed. In each iteration, the RA orientations are first optimized by Particle Swarm Optimization (PSO) to improve the weighted channel gain. Then, with the antenna orientations fixed, a reduced-dimensional formulation with Successive Convex Approximation (SCA) is used to solve the fully digital beamforming problem. Finally, a manifold-based Block Coordinate Descent (BCD) algorithm is used to jointly optimize the analog beamformer, digital beamformer, and TTD units, so that the hybrid beamforming solution gradually approaches the fully digital solution (Algorithm 1–Algorithm 4).  Results and Discussions  Simulation results verify the effectiveness of the proposed RA-assisted near-field wideband ISAC framework. The proposed PBFDA algorithm converges monotonically within a limited number of iterations, which confirms its numerical stability and efficiency (Fig. 2). Compared with fixed-antenna architectures, the proposed RA-assisted scheme achieves a clear improvement in system sum rate under the same transmit power constraint (Fig. 3). When the system bandwidth increases, the spectral efficiency of TTD-based hybrid beamforming decreases because the limited number of TTD units and the restricted maximum delay weaken frequency-dependent compensation and aggravate beam splitting. By contrast, the optimal fully digital beamforming scheme maintains nearly unchanged spectral efficiency because each subcarrier can be controlled accurately (Fig. 4). When the sensing power threshold increases, the achievable sum rate decreases for all schemes, which reflects the trade-off between communication and sensing. The proposed method, however, consistently outperforms the benchmark schemes (Fig. 5). The effects of antenna number, antenna directivity factor, and maximum rotation angle are also evaluated. Spectral efficiency increases with the number of antennas because of the higher array gain (Fig. 6). As the antenna directivity factor increases, the RA-assisted system attains further gains through adaptive orientation, whereas fixed-orientation and isotropic schemes degrade (Fig. 7). A larger allowable rotation range also provides greater spatial alignment flexibility and further improves system performance (Fig. 8). Overall, the proposed architecture improves near-field energy focusing and achieves performance close to that of fully digital beamforming with lower hardware complexity.  Conclusions  A Rotatable-Antenna-assisted near-field wideband ISAC system with a TTD-based fully connected hybrid beamforming architecture is investigated. By jointly using antenna rotation and true time delay, the proposed framework effectively mitigates near-field effects and wideband beam splitting. The developed PBFDA algorithm solves the resulting highly non-convex optimization problem efficiently. Numerical results show that the proposed scheme significantly improves the system sum rate under sensing constraints and approaches the performance of fully digital beamforming, which supports its use in near-field wideband ISAC systems.
Secure and Covert MIMO Short packet Communication with Location-Uncertain Malicious Nodes
TIAN Bo, YANG Weiwei, YANG Xiaoqin, BAI Mengmeng
2026, 48(7): 3047-3058.   doi: 10.11999/JEIT260059
[Abstract](392) [FullText HTML](157) [PDF 2732KB](34)
Abstract:
  Objective  This paper investigates secure and covert short-packet communication in Multiple-Input Multiple-Output (MIMO) wireless systems with location-uncertain malicious nodes over quasi-static Rician fading channels. In the considered scenario, a legitimate transmitter sends confidential short packets to a legitimate receiver. Meanwhile, multiple monitoring nodes (Willie nodes) attempt to detect whether transmission occurs, and multiple eavesdropping nodes (Eve nodes) attempt to intercept the confidential information. Because malicious nodes may remain silent and their exact locations are unavailable to the legitimate system, their spatial uncertainty poses major challenges to joint covertness and secrecy analysis. To address this problem, a unified analytical and optimization framework is established for secure and covert short-packet transmission. The framework is used to characterize the coupling among covertness, secrecy, and reliability and to improve the Average Effective Secrecy and Covert Rate (AESCR).  Methods  The transmitter adopts Singular Value Decomposition (SVD)-based precoding, and the legitimate receiver applies Maximum Ratio Combining (MRC) to enhance the legitimate link. Monitoring nodes and eavesdropping nodes are modeled as two independent Poisson Point Processes (PPPs) outside a circular protection zone centered at the transmitter. This model captures the spatial randomness of malicious nodes. For covertness analysis, each monitoring node is assumed to perform optimal Likelihood Ratio Test (LRT)-based detection with full knowledge of the system model, noise power, channel state, and codebook information. Using the Chernoff bound and the Bhattacharyya coefficient, a theoretical lower bound on the minimum detection error probability of a single monitoring node is first derived. Stochastic geometry is then combined with the distribution of the strongest monitoring node to obtain a tractable lower bound on the average minimum detection error probability. For secrecy analysis, the finite blocklength normal approximation is used to account for decoding error and information leakage penalties. The legitimate channel is statistically characterized under Rician fading conditions, and the strongest eavesdropping node is analyzed through stochastic geometry. Based on these results, an approximate analytical expression for the average secrecy rate is derived. AESCR is proposed as a comprehensive performance metric that jointly reflects reliability, secrecy, and covertness. Under the average covertness constraint and the short-packet length constraint, a joint optimization problem for transmit power and packet length is formulated. By using the monotonic properties of the objective function and the covertness constraint, the original coupled optimization problem is transformed into a one-dimensional search problem.  Results and Discussions  Simulation results verify the accuracy of the theoretical derivations and reveal the effects of key system parameters. Both the simulated average minimum detection error probability and its theoretical lower bound decrease as the packet length increases. Higher transmit power further reduces the detection error probability, indicating that excessive power makes transmission more exposed to monitoring nodes (Fig. 2). Increasing the number of monitoring-node antennas strengthens spatial reception capability and further degrades covertness (Fig. 2). Enlarging the protection zone improves covertness because malicious nodes are forced to remain farther away from the transmitter. However, increasing the monitoring-node density weakens this benefit by raising the probability that a strong monitoring node appears near the protection-zone boundary (Fig. 3). The average secrecy rate increases with packet length and gradually approaches the asymptotic secrecy-capacity upper bound because the finite blocklength rate penalty decreases as the packet length grows (Fig. 4). AESCR first increases and then decreases with packet length, confirming the existence of an optimal packet length. This behavior results from the tradeoff between the reduced finite blocklength penalty and increased detection exposure (Fig. 5). Higher malicious-node density and more malicious-node antennas degrade system performance because they enhance both monitoring and eavesdropping capabilities (Fig. 5). Relaxing the covertness constraint improves the achievable AESCR because the system can select a higher transmit power or a more favorable packet length (Fig. 6). Results under different Rician factors show that the proposed analytical framework is applicable to both Rician and Rayleigh fading conditions (Fig. 6). Increasing the number of legitimate receive antennas improves AESCR, and a larger transmit antenna array provides additional SVD precoding gain (Fig. 7). Compared with benchmark schemes, the proposed joint optimization of transmit power and packet length consistently outperforms the scheme with fixed packet length and power-only optimization. This result demonstrates the need to jointly balance reliability, secrecy, and covertness in MIMO short-packet transmission (Fig. 8).  Conclusions  This paper develops a stochastic-geometry-based analytical framework for secure and covert MIMO short-packet communication with location-uncertain multi-antenna malicious nodes. By deriving a lower bound on the average minimum detection error probability, obtaining an approximate analytical expression for the average secrecy rate, and proposing AESCR, the framework reveals the fundamental tradeoff among covertness, secrecy, and reliability under finite blocklength transmission. The results show that increasing the number of legitimate transmit and receive antennas improves secure and covert performance, whereas higher malicious-node density and more malicious-node antennas degrade system performance. The existence of an optimal packet length further shows that packet length and transmit power should be jointly designed. The proposed joint optimization method therefore provides an effective solution for secure and covert short-packet transmission in mission-critical and low-latency wireless systems.
Modulation Recognition Method for High-Speed Mobile Communication Based on Attention Dynamic Fusion and Hybrid Pruning Transformer
ZHENG Qinghe, CHEN Bin, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
2026, 48(7): 3059-3070.   doi: 10.11999/JEIT251211
[Abstract](685) [FullText HTML](370) [PDF 3840KB](68)
Abstract:
  Objective  Automatic modulation recognition is a critical preprocessing step in dynamic spectrum access and anti-jamming communication systems. It directly affects the robustness and spectrum efficiency of noncooperative communication. In high-speed mobile communication scenarios, such as satellite communication, high-speed rail communication, and drone swarm communication, signal modulation features experience severe distortion due to Doppler shifts, time-varying channels, and non-stationary interference. These factors challenge traditional modulation recognition methods that rely on static assumptions, leading to feature mismatch and higher misclassification rates. To address the limited robustness and real-time performance of existing deep learning–based modulation recognition models in high-speed mobile environments, a lightweight dynamic fusion Transformer-based method is proposed.  Methods  The proposed method contains three components: a signal representation fusion block, a Transformer architecture, and a model pruning strategy for lightweight inference. First, a RollingQ mechanism dynamically adjusts the direction of the attention query matrix according to the quality of each signal representation. This design prevents attention fixation and enables balanced use of multiple signal representations. Next, a Multi-head Attention Frequency Enhancement Transformer (MAFE-Transformer) is designed. The model integrates local and global spatiotemporal features through lightweight convolutional enhancement, multi-attention feature extraction, and frequency learning and selection modules. Finally, an attention-based dynamic hybrid pruning strategy removes structural redundancy and accelerates inference, which supports real-time modulation recognition.  Results and Discussions  Experiments are conducted on two public datasets, RadioML 2016.10a and RML22, to evaluate the proposed method. The MAFE-Transformer achieves average classification accuracies of 65.34% and 73.42% on the two datasets. Under low Signal-to-Noise Ratio (SNR) conditions of –20~0 dB, the model maintains strong robustness, particularly on the RML22 dataset with the dynamic channel model ETU70 (Fig. 6). The confusion matrix indicates that classification errors are relatively evenly distributed across different modulation schemes, which reflects balanced classification performance (Fig. 7). Ablation experiments show that the RollingQ-based dynamic fusion mechanism improves accuracy by 4.95% on RadioML 2016.10a and 3.86% on RML22 compared with a single signal representation (Fig. 8). The hybrid pruning strategy reduces inference latency to 0.011 ms per signal while maintaining high accuracy (Fig. 9). Comparative experiments indicate that the proposed model outperforms several advanced deep learning models, including Ms-RaT, MobileViT, MobileRaT, and KA-CNN, by 4%~10% in recognition accuracy. These results indicate strong performance in high-speed mobile communication scenarios (Fig. 10).  Conclusions  A lightweight dynamic fusion Transformer-based automatic modulation recognition method is proposed for high-speed mobile communication environments. The RollingQ mechanism and the MAFE-Transformer architecture, combined with a dynamic hybrid pruning strategy, improve the balance between recognition accuracy and inference efficiency. Experimental results on public datasets confirm the effectiveness and robustness of the method under complex channel conditions with Doppler shifts and time-varying interference. However, the method has not been systematically evaluated under more complex interference conditions, such as impulsive noise or frequency-selective fading. Future work will examine adaptability to non-stationary noise, cross-device generalization, and optimization for edge deployment.
Research on UAV-assisted Dynamic-weight Edge Computing Offloading Strategy
WANG Yijun, WANG Yachu, SHAHD Batool, MIAO Ruixin
2026, 48(7): 3071-3083.   doi: 10.11999/JEIT260054
[Abstract](421) [FullText HTML](159) [PDF 4229KB](46)
Abstract:
  Objective  The increasing demands of the Internet of Things (IoT) for computational resources and real-time processing have highlighted the significance of Mobile Edge Computing (MEC). Traditional MEC relies on terrestrial base stations, resulting in coverage blind spots in remote or specialized environments. Unmanned Aerial Vehicle (UAV)-assisted MEC architectures exploit UAVs’ flexible deployment to expand service coverage. However, existing approaches for multi-terminal, multi-UAV scenarios often fail to optimize task offloading latency, system energy consumption, and adaptability to dynamic environments simultaneously. They also overlook optimal UAV selection when terminal devices are covered by multiple UAVs and lack adaptive mechanisms to adjust optimization objectives during task execution. This study addresses these challenges by integrating cooperative caching, offloading decision-making, and resource allocation strategies.  Methods  A three-tier microcloud-edge-terminal architecture is constructed, comprising a central cloud, multiple UAV edge servers with caching capabilities, and numerous mobile terminal devices. A cooperative caching mechanism reduces transmission delay during task execution. Task offloading adopts a fine-grained partial offloading mode, dividing complex tasks into dependent subtasks modeled through a Directed Acyclic Graph (DAG). The Cooperative Caching-Adaptive Hierarchical MultiVerse Optimizer (CCAH-MVO) algorithm is proposed. A hybrid coding scheme encodes offloading decisions, caching decisions, and resource allocation uniformly. A dynamic weight mechanism adaptively balances delay and energy consumption according to the system’s real-time energy state. Additionally, a UAV selection strategy is implemented for scenarios where terminals are covered by multiple UAVs. By simulating inter-universe material exchange and local refined search, the algorithm efficiently determines the optimal offloading strategy. MATLAB simulations validate the method under various experimental settings.  Results and Discussions  The simulation scenario involves 50 randomly distributed terminal devices and 5 UAVs in a 400 m × 400 m area. UAVs are deployed above terminal cluster centers, while terminals at cluster edges are simultaneously within the coverage of multiple UAVs (Fig. 5). The optimal UAV for each terminal is selected using the UAV selection function (Fig. 6), preventing resource bottlenecks and achieving balanced load distribution. In terms of delay performance, the CCAH-MVO algorithm maintains the lowest task delay across all task volumes, with a gradual increase as the number of tasks grows (Fig. 7). Delay under CCAH-MVO is consistently lower than that under fixed-weight strategies across the full task range, demonstrating the effectiveness of the dynamic adaptive mechanism in preserving low latency (Fig. 10). For energy consumption, differences among the algorithms are minor when task quantities are low. Under high task loads, the activation of the dynamic weight mechanism flattens the energy consumption curve (Fig. 8). When the number of tasks reaches 100, total energy consumption under CCAH-MVO is the lowest among all strategies and remains lower than the fixed-weight approach, reflecting effective control under critical energy conditions (Fig. 9). Regarding total system overhead, the CCAH-MVO algorithm consistently achieves the best performance. The gap with fixed-weight strategies widens when task numbers exceed 80, illustrating the dynamic weight mechanism’s collaborative optimization of delay and energy consumption (Fig. 11). Overall, by integrating the dynamic weight mechanism and balancing load through UAV selection, the CCAH-MVO algorithm effectively mitigates resource constraints and high task processing overhead in complex, dynamic UAV-assisted MEC environments. It ensures precise coordination between task delay and energy consumption across different load stages.  Conclusions  The proposed CCAH-MVO framework, incorporating a microcloud-edge-terminal architecture, cooperative caching mechanism, fine-grained partial offloading, dynamic weight adjustment, and UAV selection strategy, effectively addresses resource scheduling in complex multi-UAV MEC environments. Simulations show adaptive optimization of objectives, intelligent energy management, low latency, and reduced total system overhead, improving service stability and user experience. This research provides a practical solution for efficient UAV edge computing in dynamic environments. Future work will explore dynamic energy efficiency optimization and multi-node collaboration while maintaining low-latency performance.
Joint Channel Estimation and Diagnosis for Blocked RIS-Assisted Multi-User Multipath Millimeter-Wave Systems
LI Shuangzhi, LIU Cong, WANG Ning, HAN Gangtao, GUO Xin
2026, 48(7): 3084-3093.   doi: 10.11999/JEIT260093
[Abstract](467) [FullText HTML](183) [PDF 2820KB](22)
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS) can effectively modulate Millimeter-Wave (mmWave) signals and reshape the wireless propagation environment. In practical deployments, however, RIS elements are vulnerable to adverse weather and physical obstructions, which cause unpredictable distortion and motivate joint channel estimation and blockage diagnosis. Most existing studies focus on single-user systems, whereas multi-user scenarios remain insufficiently studied. This gap creates an opportunity to exploit the common RIS blockage vector and the shared RIS-Base Station (BS) channel across users. This paper therefore proposes a low-complexity framework for joint channel estimation and blockage diagnosis by exploiting the sparsity and correlation of multi-user cascaded channels.  Methods  Under the assumption that all User Equipment (UE) shares the same RIS-BS channel and is affected by a common RIS blockage vector, the problem is divided into two stages. First, a target UE is selected. The sparsity of the mmWave channel and blockage vector, together with the linear dependence among RIS-BS paths, is used to formulate a sparse recovery problem. A hierarchical Bayesian model is then adopted, and an efficient Sparse Bayesian Learning (SBL) algorithm is used for joint recovery. Second, partial Channel State Information (CSI) obtained from the target UE is used to construct a common channel matrix that combines the RIS-BS channel and blockage information. Channel estimation for the remaining UEs is then reformulated as another sparse recovery problem.  Results and Discussions  A low-complexity strategy for cascaded channel estimation and blockage diagnosis is developed by exploiting the sparsity and correlation of multi-user cascaded channels and the commonality of the RIS blockage vector. Ideal estimation results are used as a theoretical lower bound, and the proposed algorithm is compared with two benchmark schemes. Simulation results show that the proposed algorithm consistently outperforms the benchmark schemes (Fig. 1). Specifically, a higher target-user Signal-to-Noise Ratio (SNR) improves the Normalized Mean Square Error (NMSE), which confirms the importance of target-user selection (Fig. 2). The algorithm also shows good convergence as the number of iterations increases (Fig. 3), and its performance approaches the ideal case more closely as the number of time frames increases (Fig. 4). In addition, the method remains robust as the number of blocked elements increases (Fig. 5). More BS antennas further improve performance by enhancing array orthogonality (Fig. 6). By exploiting path correlation, the proposed method achieves better estimation accuracy with slightly lower runtime (Table 1). However, estimation accuracy decreases as the number of paths increases because the model becomes more complex (Figs. 7 and 8).  Conclusions  This paper proposes a joint channel estimation and blockage diagnosis framework for blocked RIS-assisted multi-user multipath mmWave systems. Simulation results show that the method approaches the theoretical performance bound in complex multipath environments. It also maintains clear performance advantages under high blockage rates while reducing computational complexity through the use of common channel structures. This study provides a practical solution to performance degradation in RIS deployment, clarifies the effects of key parameters, and offers guidance for system design. Because practical blockages often exhibit block-sparse or structured-sparse characteristics, future work may incorporate structured priors, such as group sparsity and Markov random fields, into the SBL framework to capture spatial correlation and improve diagnostic accuracy and robustness.
Joint Power Allocation and AP On-Off Control for Long-Term Energy Efficient Cell-Free Massive MIMO Systems
WEI Siqi, GUO Fengqian, CHONG Baolin, CHENG Guo, LU Hancheng
2026, 48(7): 3094-3104.   doi: 10.11999/JEIT260014
[Abstract](568) [FullText HTML](298) [PDF 2818KB](28)
Abstract:
  Objective   With the rapid development of wireless communication technologies, Cell-Free Massive Multiple-Input Multiple-Output (CF-mMIMO) has emerged as an effective paradigm to overcome the limitations of traditional cell-centric networks, such as limited performance for edge users. By deploying a large number of distributed Access Points (APs) connected to a Central Processing Unit (CPU) to cooperatively serve users, CF-mMIMO improves spectral efficiency and macro-diversity gain. However, dense AP deployment also introduces a critical challenge: high energy consumption. In practical systems, if all APs remain continuously active, especially during periods of low traffic load, substantial and unnecessary energy consumption occurs. This behavior reduces network sustainability and conflicts with global “dual-carbon” goals. Existing studies on energy efficiency in CF-mMIMO systems mainly focus on short-term performance optimization. These short-term approaches often ignore long-term traffic dynamics and the requirement of queue stability. Therefore, they lack robustness under time-varying traffic conditions and may cause queue congestion and significant performance fluctuations, which are unacceptable for next-generation wireless networks with strict reliability requirements. Although several recent studies examine long-term energy efficiency optimization, most assume that all APs remain active at all times. Therefore, the energy-saving potential of adaptive AP on-off control is not fully utilized.  Methods   To address these issues, a joint power allocation and AP on-off control strategy is proposed for downlink CF-mMIMO systems. The optimization problem aims to maximize long-term energy efficiency subject to user queue stability and AP power constraints. Because the problem has stochastic and long-term characteristics, the Lyapunov optimization framework is applied to transform the original long-term fractional programming problem into a sequence of deterministic drift-plus-penalty minimization problems solved in each time slot. The resulting per-slot problems remain nonconvex. Therefore, each problem is decomposed into two subproblems: power allocation and AP on-off control. The Successive Convex Approximation (SCA) method is used to convert the nonconvex formulations into solvable convex problems. An alternating optimization algorithm is then developed to jointly solve the two subproblems, which enables adaptive resource configuration under dynamic network conditions and stochastic traffic arrivals.  Results and Discussions   The proposed algorithm is evaluated through extensive simulations. First, the convergence behavior is examined. Numerical results (Fig. 2) show that per-slot energy efficiency increases rapidly and stabilizes after several iterations, which verifies the convergence of the alternating optimization procedure. Second, the effect of the control parameter is analyzed. As the parameter increases, the algorithm places greater emphasis on energy efficiency. Average power consumption decreases and then stabilizes (Fig. 3), whereas long-term energy efficiency increases and eventually stabilizes (Fig. 4). These results confirm the trade-off between energy efficiency and queue stability. Third, the proposed scheme is compared with three baseline methods. The results (Fig. 5) show that the proposed joint optimization approach consistently achieves higher long-term energy efficiency than the baseline methods. Fourth, the necessity of long-term optimization is demonstrated by comparing queue lengths with a short-term baseline (Fig. 6). Under the same traffic arrival rate, the short-term method shows cumulative queue growth, whereas the Lyapunov-based approach maintains queue lengths within a stable range and ensures network stability. Finally, robustness under imperfect Channel State Information (CSI) is evaluated (Fig. 7). Although energy efficiency decreases as channel uncertainty increases, the proposed method consistently outperforms the baseline approaches, which demonstrates strong robustness to channel estimation errors.  Conclusions   A long-term energy efficiency optimization framework is proposed for CF-mMIMO systems with stochastic traffic arrivals. By applying Lyapunov optimization theory, the stochastic long-term problem is transformed into slot-level drift-plus-penalty problems based on queue states. This transformation enables per-slot resource scheduling decisions while maintaining queue stability. On this basis, an efficient joint resource scheduling algorithm that integrates power allocation and AP on-off control is developed. The original problem is decomposed into power allocation and AP on-off control subproblems and solved through alternating optimization. Simulation results show that the proposed method adapts to dynamic traffic conditions. By placing underutilized APs into sleep mode, the algorithm improves long-term system energy efficiency and maintains queue stability. These results provide guidance for the design of green and sustainable wireless networks.
SG-DDPG-based Low-intercept Point Beam Design for FDA-MIMO Short-range Detectors
JIA Jinwei, GAO Min, HAN Zhuangzhi, LIU Limin, YIN Yuanwei
2026, 48(7): 3105-3122.   doi: 10.11999/JEIT260010
[Abstract](433) [FullText HTML](162) [PDF 10678KB](37)
Abstract:
  Objective  Radio short-range detectors are widely used in many detection systems. However, in modern battlefields, the electromagnetic environment is increasingly complex, and radio short-range detectors must withstand various forms of electromagnetic interference. In particular, fourth-generation jammers based on Digital Radio Frequency Memory (DRFM) can implement repeater deception jamming. Such jamming may cause failures such as premature detonation in radio short-range detectors and reduce their damage effectiveness. Anti-repeater deception jamming has therefore become a key issue for short-range detectors. Improving the Low Probability of Intercept (LPI) performance of radio short-range detectors is an effective means of resisting repeater deception jamming. According to the Chinese manuscript, this study focuses on the effect of FDA-MIMO array-element frequency-offset settings on beam synthesis and proposes an SG-DDPG-based method for LPI point beam design.  Methods  Frequency Diverse Array-Multiple-Input Multiple-Output (FDA-MIMO) technology is used in this study, and the key factors affecting beam convergence are analyzed. For the spatial LPI beam design of radio short-range detectors, a performance evaluation model for spatial LPI beams is constructed. An FDA-MIMO LPI point beam design method based on the Stage Guidance-Deep Deterministic Policy Gradient (SG-DDPG) algorithm is then proposed. In the SG-DDPG algorithm, a multidimensional staged guidance reward function is designed. An Actor-Critic model is used to maximize the reward value through gradient ascent. The array-element frequency offsets that provide better beam convergence in the current environment are then obtained. The SG-DDPG algorithm is suitable for LPI point beam design under different fall angles of radio short-range detectors. It overcomes the technical limitation of formula-based frequency-offset calculation, which is only applicable when the detector fall angle is close to vertical.  Results and Discussions  The simulations show that, after the array-element frequency offsets are optimized by the SG-DDPG algorithm, the FDA-MIMO beam achieves a half-power beam width of 1 m in the range dimension and 9.9° in the angular dimension. The proposed method provides better beam convergence and LPI performance than classical frequency-offset design methods. These results indicate that the proposed algorithm offers an effective approach for array-element frequency-offset optimization and LPI point beam design, thereby improving the LPI performance of radio short-range detectors.  Conclusions  This paper presents an FDA-MIMO LPI point beam design method based on the SG-DDPG algorithm, with the array-element frequency offset used as the optimization objective. The simulation results support two main conclusions. First, the proposed method removes the restriction that the fall angle of the radio short-range detector must be close to vertical when the array-element frequency offset is calculated by a formula-based method. The algorithm can be applied to LPI beam design under different fall angles and improves the LPI performance of radio short-range detectors. Second, the proposed method achieves a half-power beam width of only 1 m in the range dimension and 9.9° in the angular dimension, which is better than that of traditional methods. Under different fall angles, the beam formed by the proposed method has the smallest intercept area, indicating the best LPI performance.
Index Modulation Design with Sparse Spatial Constellation and Dynamic Multi-RIS-Block Selection for RIS-MIMO Systems
HUANG Fuchun, ZHU Han, TANG Xiaoqing, YANG Fan, HUANG Jie
2026, 48(7): 3123-3134.   doi: 10.11999/JEIT251289
[Abstract](386) [FullText HTML](167) [PDF 2740KB](19)
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS)-assisted Multiple-Input Multiple-Output (MIMO) Index Modulation (IM) systems face two main challenges: the difficult deployment of a single large-scale RIS panel and the high design complexity of efficient transmit spatial signal vectors. To address these issues, a joint design that combines sparse spatial constellation and dynamic multi-RIS-block selection is proposed. The design improves spectral efficiency, Bit Error Rate (BER) performance, and deployment flexibility.  Methods  Inspired by the Extended Space Index Modulation (ESIM) paradigm, a sparse Spatial Constellation with Two Active Antennas (SCTA) is proposed, forming the SCTA-RIS-SM system. In this design, Pulse Amplitude Modulation (PAM) and Secondary PAM (SPAM) constellations are combined to construct the spatial constellation vector [x1,x2]T, which is modulated onto two active antennas. This design maximizes the Minimum Euclidean Distance (MED) between transmit vectors and improves the anti-interference capability of the system. To address the deployment difficulty of a single large-scale RIS panel, an enhanced SCTA-MBRIS-SM system is further proposed. The system uses a distributed array of small RIS blocks and dynamically selects a subset of blocks for cooperative reflection. Different RIS block selection combinations are used as a new IM dimension. Spectral efficiency and average BER are then analyzed theoretically. Monte Carlo simulations are conducted to compare the proposed systems with several existing schemes.  Results and Discussions  The simulation results show that the proposed SCTA-RIS-SM system achieves clear Signal-to-Noise Ratio (SNR) gains over RIS-SIM, RIS-SM, and DHRIS-SM systems at the same spectral efficiency, such as 10–12 bits/(s·Hz). For instance, when BER = 10–3, SCTA-RIS-SM outperforms RIS-SIM by approximately 1.5–2.5 dB and DHRIS-SM by more than 6 dB. By using additional IM from RIS block selection, SCTA-MBRIS-SM further improves BER performance and spectral efficiency compared with SCTA-RIS-SM, without increasing the number of Radio Frequency (RF) chains. With the same total number of reflecting elements, the proposed multi-RIS-block scheme achieves an SNR gain of up to 5 dB over RIS-SIM when BER = 10–3. The theoretical BER curves agree well with the simulation results in the high-SNR region, confirming the validity of the analytical derivations. The results also indicate that the performance advantage is maintained as the number of transmit antennas increases. In addition, the proposed design is compatible with channel coding.  Conclusions  This paper addresses the challenges of large-scale RIS deployment and high-complexity spatial signal design in RIS-assisted MIMO systems. The proposed SCTA design improves system reliability by optimizing the Euclidean distance distribution in the signal space. Dynamic multi-RIS-block selection transforms hardware deployment constraints into a new dimension for improving spectral efficiency, providing a feasible path for practical large-scale RIS applications. Simulation results confirm that joint optimization of transmit spatial vectors and RIS reflection degrees of freedom is an effective strategy for improving system performance. Future work will focus on robust design under imperfect channel state information, construction of higher-dimensional sparse constellations, extension to extremely large-scale MIMO scenarios, and multi-user communications.
Radar, Sonar,Navigation and Array Signal Processing
Research on Inverse QR Decomposition Optimization for Sparse Adaptive System Identification Algorithms
PENG Yi, ZHANG Pengfei, WANG Xiaoyong, GAO Junqi, LI Changlong, ZHANG Zhiyuan, SUN Tianxiang
2026, 48(7): 3135-3145.   doi: 10.11999/JEIT250562
[Abstract](453) [FullText HTML](160) [PDF 4926KB](35)
Abstract:
  Objective  Traditional sparse-regularized Recursive Least Squares (RLS) algorithms, namely L1/L0-norm Recursive Least Squares (L1/L0-RLS), have theoretical advantages in sparse parameter-space estimation and are widely used in system identification and channel equalization. However, under limited numerical precision, iterative covariance matrix computation may cause rounding errors to accumulate. This can lead to divergence and instability in the least-squares solution.  Methods  To address this problem, an improved algorithm based on the Inverse QR Decomposition (IQRD) framework is proposed. The framework suppresses rounding-error accumulation in traditional regularized RLS algorithms. It also removes the back-substitution step for weight coefficients required in conventional QR decomposition. These features improve numerical robustness and system identification efficiency in finite-precision environments. Specifically, L1-IQRD-RLS and L0-IQRD-RLS algorithms are constructed under an L1/L0-constrained IQRD architecture. A general recursive expression for the weight coefficients is derived. An automatic parameter selection mechanism is also incorporated into the algorithm framework to solve the dynamic optimization problem of the sparse regularization parameter.  Results and Discussions  Monte Carlo simulations are conducted to evaluate the sparse constraints and robustness of the proposed algorithms. The results show that L1-IQRD-RLS and L0-IQRD-RLS maintain long-term numerical stability in an 11-decimal-place fixed-point computing environment. Compared with traditional algorithms, the proposed algorithms show clear advantages in system sparsity representation, parameter estimation variance, and covariance matrix condition number. Measured-data verification further confirms that the improved algorithms maintain numerical stability under limited-precision conditions and are more robust than traditional methods. The measured-data results also show that the regularized RLS algorithms optimized by the IQRD framework have advantages in system sparsity representation, parameter estimation, and numerical stability. Their iterative convergence success rate is higher than that of traditional methods.  Conclusions  This paper addresses sparse system identification in adaptive filtering. Traditional sparse-regularized RLS algorithms still face numerical stability problems under limited numerical precision. To solve this problem, an IQRD framework is constructed to reduce the numerical ill-conditioning caused by accumulated rounding errors in sparse-regularized RLS algorithms. The proposed method improves numerical robustness in low-precision environments. In addition, an automatic parameter selection mechanism is incorporated into the algorithm framework. This reduces repeated parameter tuning and supports stable performance optimization under sparse constraints. In practical electromagnetic signal processing, system identification and beamforming are limited by the finite precision of hardware implementation and often exhibit inherent system sparsity. The proposed algorithm provides a targeted solution. Its finite-word-length robustness suppresses numerical divergence during adaptive weight updates and supports stable implementation on fixed-point processors. The sparse constraints also match the physical characteristics of sparse systems and improve estimation accuracy. This study provides a practical algorithm for high-performance and high-stability sparse-constrained systems on precision-limited hardware platforms.
Optimal Weighted Subspace Fitting-based Direct Position Determination with HF/VHF Collaboration
YANG Gaoyuan, YIN Jiexin, WANG Ding, YANG Bin
2026, 48(7): 3146-3161.   doi: 10.11999/JEIT260001
[Abstract](398) [FullText HTML](216) [PDF 5354KB](23)
Abstract:
  Objective   Passive localization is essential for target detection, navigation, and track tracking, particularly in military applications involving maritime and aerial targets. These targets often transmit across multiple frequency bands, including shortwave High Frequency(HF) and Very High Frequency (VHF). Existing localization methods largely rely on single-band approaches or two-step positioning techniques. Single-band methods underutilize the positional information available across different bands, while two-step methods lose information during intermediate parameter estimation (e.g., Direction-Of-Arrival (DOA); Time-Difference-Of-Arrival (TDOA)), reducing localization accuracy. Collaborative fusion of HF signals (via ionospheric reflection) and VHF signals (via Doppler effects from moving arrays) has been rarely addressed. To overcome low positioning accuracy and limited spatial resolution in over-the-horizon multi-target scenarios, this study proposes a novel collaborative Direct Position Determination (DPD) method designed to integrate the complementary strengths of HF and VHF signals, enhancing localization precision and robustness in complex electromagnetic environments.  Methods  An Optimal Weighted Subspace Fitting (OWSF) DPD algorithm is proposed. Comprehensive signal propagation models are established for heterogeneous observation platforms (Fig. 1). HF signal propagation is modeled using a two-dimensional DOA framework based on ionospheric reflection, incorporating azimuth and elevation angles to handle nonlinear over-the-horizon propagation. VHF signals are modeled using a space-time extended signal framework for a moving Unmanned Aerial Vehicle (UAV), exploiting Doppler effects to create a virtual large-aperture array that captures both one-dimensional angle and Frequency-Of-Arrival (FOA) information. Unlike traditional methods that process each band separately, the OWSF algorithm constructs a unified cost function that fuses the signal and noise subspaces of both HF and VHF data using optimal weighting matrices, balancing the contributions of different signal qualities. Target positions are then estimated by minimizing this cost function via grid search or Newton iteration. The Cramér-Rao Bound (CRB) under Earth-ellipsoid constraints is derived to provide the theoretical performance limit.  Results and Discussions   Simulations are conducted in a centralized processing scenario, where HF stations and UAV VHF signals are transmitted to a central station for joint processing (Fig. 2). The simulation involves three stationary targets and a collaborative system comprising HF stations and a UAV (Fig. 3, Table 2, Table 3). Performance comparisons demonstrate that the OWSF method consistently outperforms traditional two-step positioning methods and single-system DPD methods (DOA-only or FOA-only) in Root Mean Square Error (RMSE) (Fig. 4). When HF SNR is 5 dB lower than VHF SNR, OWSF exhibits superior robustness compared to Subspace Data Fusion (SDF) and Minimum Variance Distortionless Response (MVDR) methods, approaching the CRB at high SNR (Fig. 5). The impact of system parameters is further analyzed, showing that increasing the number of sampling points (Fig. 6) and array elements (Fig. 7) improves accuracy, particularly in low SNR regimes. Regarding spatial resolution, the OWSF algorithm generates sharper spectral peaks for distant targets and successfully resolves closely spaced targets that the SDF-DPD algorithm fails to distinguish (Fig. 7, Fig. 8).  Conclusions   The HF/VHF collaborative DPD method effectively integrates multidimensional observational information from ionospheric reflection and Doppler-based propagation. Simulation results demonstrate substantial improvements in localization accuracy, spatial resolution, and robustness, especially under low-SNR conditions or heterogeneous signal quality between bands. The derived CRB provides a solid theoretical benchmark, confirming that the method overcomes the limitations of single-band and two-step approaches. This approach offers a highly effective solution for over-the-horizon passive localization of multiple stationary targets.
An Ultra-Wideband Low-Profile Dipole Patch Antenna for VHF-Band Probing Radars
TIAN Yuxiao, ZHANG Feng, MA Zhangjun, WANG Jiacheng, JI Yicai
2026, 48(7): 3162-3169.   doi: 10.11999/JEIT260105
[Abstract](486) [FullText HTML](343) [PDF 3135KB](51)
Abstract:
  Objective  In radar systems, the limitations of traditional narrowband antennas in data transmission rate and resolution have become increasingly evident. Ultra-WideBand (UWB) antennas therefore receive broad attention because they provide high range resolution and strong interference suppression capability. However, at low frequencies, existing UWB antennas usually suffer from excessively large physical size, which makes installation on airborne or vehicle-mounted platforms difficult. By contrast, compact antennas that are easier to deploy often exhibit insufficient gain and cannot satisfy the penetration-depth requirement of deep subsurface detection. Thus, achieving a proper balance among antenna size, bandwidth, and gain over an ultra-wideband range remains a major challenge for VHF-band probing radars. To address this issue, a planar dipole antenna loaded with an Artificial Magnetic Conductor (AMC) structure and metallic shorting walls is proposed. The antenna maintains stable radiation performance over a wide frequency range while preserving a low-profile and structurally simple configuration.  Methods  The reflection-phase characteristics of AMC unit cells with different geometries are compared, and square unit cells are selected to construct a 9 × 7 AMC reflective layer. Owing to its in-phase reflection property, the AMC structure removes the conventional requirement for a quarter-wavelength spacing between the antenna and a metallic ground plane, thereby reducing the profile height. The dipole patch adopts an optimized meandered current-bending structure to reduce the lateral size. Metallic shorting walls are further loaded at both ends of the antenna. According to image theory, equivalent currents are generated on the outer surfaces of these metal walls during operation, which effectively extends the electrical length and improves low-frequency performance without increasing the physical size. In addition, two vertical metallic walls are connected to the ground plane on both sides of the antenna to form a reflective back cavity, which strengthens unidirectional radiation and improves antenna gain. As part of the overall co-design, four 125 Ω resistors are inserted between the feed region and the metallic sidewalls. This resistive loading suppresses strong low-frequency resonances and broadens the impedance bandwidth at the cost of acceptable Ohmic loss.  Results and Discussions  A prototype with favorable simulated performance is fabricated and measured in a microwave anechoic chamber. The measured impedance bandwidth for VSWR<2 is 50~400 MHz, which agrees well with the simulated range of 84~366 MHz. The measured impedance matching is slightly better than the simulated result, mainly because cable loss and power-divider loss in the feeding network reduce the reflected power. The measured gain follows the same trend as the simulated gain, with deviations within 1 dBi. Radiation-pattern measurements show that at 100, 200, and 300 MHz, the measured copolarization patterns agree well with the simulated results, and the maximum radiation direction remains normal to the antenna plane, which confirms the effectiveness of the proposed design. As shown in Fig. 5, the current on the radiating patch layer mainly flows along the +x direction and generates a radiated electric field along the +z direction. The current on the AMC unit can be represented by an equivalent current loop oriented along the +z direction. At this frequency, the x-direction current and the parasitic current loop on the AMC jointly enhance the antenna gain. This result explains the gain-improvement mechanism of the AMC structure. When the operating frequency increases to 400 MHz, the electrical size of the antenna reaches approximately \begin{document}$ 1.6\lambda $\end{document}, which causes main-lobe splitting and shifts the maximum radiation direction toward 90°. Although this high-frequency beam splitting introduces spatial clutter, it is an acceptable physical trade-off for achieving the ultra-low profile of 0.07 λL, while the overall UWB characteristic still supports high time-domain resolution in probing radar systems. At 400 MHz, the measured H-plane co-polarization level is slightly higher than the simulated value, possibly because of coupling between the feeding cable and the vertically mounted antenna.  Conclusions  A low-profile UWB planar dipole antenna is proposed for VHF-band probing radar applications. By combining the AMC layer, metallic shorting walls, and resistive loading, the proposed design improves impedance matching while preserving a compact size. The reflective back cavity further improves the realized gain. The fabricated prototype shows good agreement between measurement and simulation. The antenna operates over 100~366 MHz and exhibits a measured VSWR<2 bandwidth of 50~400 MHz. It maintains a compact electrical size of 0.38λL × 0.18λL × 0.07λL, and the maximum measured gain within the operating band reaches 6 dBi. The proposed co-design provides a practical solution for low-frequency probing radar antennas that require wide bandwidth, low profile, and relatively high gain.
Circuit and System Design
Optimizing Satisfiability-Based Automatic Test Pattern Generation Systems: Unified Fault Set Construction,Modeling, and Solving
YAN Dapeng, HE Qirun, GUO Jing, WANG Boning, CAI Zhikuang
2026, 48(7): 3170-3180.   doi: 10.11999/JEIT260025
[Abstract](480) [FullText HTML](316) [PDF 2362KB](32)
Abstract:
  Objective  Boolean SATisfiability-Based Automatic Test Pattern Generation (SAT-Based ATPG) is widely used to generate tests for hard-to-detect single stuck-at faults and to prove fault untestability in combinational logic. When SAT-Based ATPG is applied to large netlists with dense fanout and reconvergence, its runtime and memory consumption are often dominated by three interacting issues. Representative fault lists produced by conventional dominance- or equivalence-based fault collapsing can remain large, increasing the number of SAT calls and enlarging the incremental context that must be maintained across faults. Meanwhile, SAT modeling may introduce redundant Conjunctive Normal Form (CNF) overhead, especially when an explicit faulty-circuit copy is constructed or when propagation constraints are encoded globally without locality control. In addition, fanout-reconvergence structures amplify assignment correlations along sensitized paths, and such correlations are often exposed only after repeated decisions and backtracking when only standard unit propagation is used. The unified optimization objective is therefore to reduce overall CNF size and solving cost while preserving completeness, so that a practical SAT-Based ATPG system remains efficient and stable across circuits of different scales.  Methods  A three-part framework is developed and implemented in an incremental SAT-Based ATPG flow, and the overall workflow is illustrated (Fig. 1). First, a checkpoint-driven dynamic fault-set construction method is proposed. Checkpoints are collected during netlist-to-directed-acyclic-graph conversion, including all primary inputs and all fanout branches, and XOR/XNOR outputs are additionally recorded as supplementary checkpoints to avoid over-collapsing XOR-related fault behavior. Representative faults are initialized on checkpoints by compact rules that combine dominance-oriented fault collapsing with equivalence-aware refinement, and solver-guided repair is performed when an untestable representative fault indicates potential masking under structural constraints. The procedure is summarized in Algorithm 1. Second, an SAT modeling method based on fault sensitization constraints is adopted to avoid explicit faulty-circuit duplication. Fault activation, propagation, and observability are represented by additional fault sensitization constraints over the original circuit variables, and auxiliary variables are introduced only when local bookkeeping is required. Constraint localization is restricted to the fault fanout cone, and cone-boundary and internal vertices are identified through a graph-traversal procedure (Fig. 2). Third, a dynamic implication learning mechanism oriented to fanout-reconvergence pairs is integrated into the incremental solving loop. Reconvergence pairs within the fault fanout cone are monitored under partial assignments, and structure-induced implications are injected either as implied assignments when a reconvergent output becomes functionally determined or as short conflict clauses when a branch-value combination becomes inconsistent with the fault sensitization constraints. The dynamic implication learning procedure is summarized in Algorithm 2.  Results and Discussions  The unified system is evaluated on ISCAS’85 and ISCAS’89 benchmark circuits, with TG-PRO used as the baseline implementation under the same SAT solver and termination settings. The checkpoint-driven dynamic fault-set construction method substantially reduces the representative fault space entering ATPG. Relative to the uncollapsed fault space, the average representative-fault ratio decreases from 51.38% to 42.41%, corresponding to an average fault-space reduction of 57.59%. The best-case ratio reaches 33.19% on large circuits with heavy reconvergence, which indicates that checkpoint-centered representative-fault allocation effectively suppresses redundancy without enlarging the untestable fault set (Table 1). The reduced fault-set size is reflected in preprocessing efficiency, and the total runtime for fault-set construction is consistently reduced, with an average reduction of 8.37% across the evaluated circuits (Fig. 3). For SAT model construction, the fault-sensitization-constraint encoding reduces CNF overhead relative to the baseline model construction. Across the benchmark set, the numbers of CNF clauses and CNF variables are reduced by 11.44% and 3.50%, respectively, which shows that avoiding explicit faulty-circuit duplication and localizing auxiliary constraints to the fault fanout cone effectively lowers memory demand (Table 2). The reduced CNF size and strengthened locality of constraints are further reflected in end-to-end runtime, and the total runtime of SAT modeling and solving is reduced across the evaluated benchmarks (Fig. 4). Dynamic implication learning further improves solving efficiency in reconvergence-heavy structures. Compared with static implication learning, CNF construction time increases by 3.0% on average because of the additional monitoring and injection operations, yet the overall runtime decreases by 4.42% on average, which indicates a favorable cost-benefit trade-off. The overhead attributed to dynamic implication learning accounts for 2.51% of the total runtime aggregated across circuits, which confirms that the injected implications and pruning clauses provide measurable solving benefits at limited extra cost (Table 3).  Conclusions  A unified optimization framework for SAT-Based ATPG is developed by combining checkpoint-driven dynamic fault-set construction, localized fault sensitization constraints for CNF modeling, and fanout-reconvergence-oriented dynamic implication learning. Representative faults are compressed through solver-guided repair of dominance and equivalence relations to avoid masking, CNF growth is controlled through duplication-free modeling localized to the fault fanout cone, and reconvergence correlations are exploited through incremental implication injection to strengthen propagation and enable early conflict pruning. Experimental results on standard benchmark circuits show consistent reductions in representative fault scale, CNF size, and total runtime, providing a practical approach for scaling SAT-Based ATPG to larger designs with complex fanout and reconvergence.
Design of a Timing-Controlled Nonvolatile Flip-Flop for Low-ON/OFF-Current-Ratio FeFETs
DU Shimin, YANG Chang, WANG Lunyao, ZHANG Zhe
2026, 48(7): 3181-3192.   doi: 10.11999/JEIT251059
[Abstract](362) [FullText HTML](252) [PDF 9154KB](24)
Abstract:
  Objective  Nonvolatile Processors (NVPs) are a key technology for Internet of Things (IoT) and energy-harvesting systems, in which computational states must be preserved during unexpected power loss. Conventional volatile processors rely on external Nonvolatile Memory (NVM) for state retention. However, this approach causes high latency and energy overhead. Integrated Nonvolatile Flip-Flops (NVFFs) based on Ferroelectric Field-Effect Transistors (FeFETs) provide a promising alternative by enabling on-chip state backup and recovery. However, existing single-ended FeFET-based flip-flops are prone to contention-induced recovery failures, especially when the FeFET ON/OFF current ratio degrades. This failure arises from contention among internal metal-oxide-semiconductor transistors, which makes internal node settling uncertain and causes unreliable state recovery. To address this issue, this paper proposes a timing-controlled NVFF architecture that replaces contention-based recovery with a two-stage recovery mechanism. The proposed design aims to achieve reliable recovery under degraded FeFET ON/OFF current ratios as low as 102, improve timing metrics such as hold time and clock-to-Q delay, and maintain low energy consumption for IoT applications.  Methods  The proposed design extends the Static Contention-Free Single-Phase-Clocked Flip-Flop (SSCFF), whose fully static structure suppresses internal node contention. On this basis, one FeFET and five additional Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) are integrated to construct a single-ended NVFF. Two control signals, RES and MOD, are used to manage the recovery process. In normal operation, MOD = 0, and the circuit functions as a conventional SSCFF while supporting runtime state backup. In recovery mode, MOD = 1, and the recovery process is divided into two stages. In the precharge stage, when RES = 0, the internal nodes are precharged to VDD. In the selective-discharge stage, RES switches from low to high, and the FeFET resistance state determines whether discharge occurs. If the FeFET is in the Low-Resistance State (LRS), a discharge path is formed, and the node voltage is pulled down to ground. If the FeFET is in the High-Resistance State (HRS), the node retains its charge until the next clock edge. This precharge-selective-discharge sequence removes recovery contention and enables deterministic internal node settling. The design is implemented using a 130 nm Complementary Metal-Oxide-Semiconductor (CMOS) process and an integrated FeFET model. Simulations are performed in Cadence Virtuoso across a supply voltage range of 0.6~0.9 V and FeFET ON/OFF current ratios from 102 to 104. Key metrics, including setup time, hold time, clock-to-Q delay, recovery energy, and recovery success rate, are evaluated and compared with those of a conventional Transmission-Gate Flip-Flop (TGFF).  Results and Discussions  Simulation results show that timing-controlled recovery improves reliability under severe FeFET degradation. At an FeFET ON/OFF current ratio of 102, the proposed flip-flop achieves a 100% recovery success rate in 2000 Monte Carlo simulations. This improvement is attributed to the removal of contention among internal recovery paths. Timing metrics are also improved. The 3σ worst-case hold time is reduced by 64.6%, and the clock-to-Q delay is reduced by 33.9%. Although setup time increases slightly, this increase can be mitigated through device sizing. Recovery energy remains at the fJ level, with values of approximately 10 fJ under the tested conditions. This energy is only slightly higher than that of the TGFF because of the added precharge stage.  Conclusions  An FeFET-based NVFF with timing-controlled two-stage recovery is presented to address the contention-induced failure modes that limit low-voltage recovery reliability. By integrating a single FeFET into an enhanced SSCFF structure and using the RES signal to control precharge and selective discharge, the proposed design maintains a high recovery success rate even under severely degraded FeFET ON/OFF current ratios. It also improves hold time and clock-to-Q delay compared with conventional transmission-gate NVFFs. The proposed architecture provides an effective solution for energy-constrained IoT processors that require fast and reliable state preservation under unpredictable power conditions.
A Frequency Domain Self-Attention Guided MultiscaleInverse Lithography Technology
LUO Binling, WANG Ying, CAI Shuting
2026, 48(7): 3193-3202.   doi: 10.11999/JEIT251382
[Abstract](613) [FullText HTML](258) [PDF 3536KB](42)
Abstract:
  Objective  Optical Proximity Effect (OPE) in lithographic processes causes printed wafer patterns to deviate from target layouts. Therefore, Optical Proximity Correction (OPC) is required for mask optimization before exposure. Traditional rule-based OPC methods show reduced accuracy for complex layouts, whereas model-based OPC methods require high computational cost. Deep learning-based methods have recently been used to accelerate mask generation. However, their limited receptive fields make it difficult to model long-range optical interference, which restricts optimization accuracy. To address these limitations, this work proposes Frequency-Domain Self-Attention-Guided Multiscale Inverse Lithography Technology (FMS-ILT). The method jointly models local geometric details and global optical interference to improve printed image fidelity, edge placement accuracy, and process robustness.  Methods  FMS-ILT uses a residual convolution-based multiscale encoder-decoder architecture. Shallow layers extract fine geometric features, such as edges and corners, whereas deeper layers capture large-scale layout context. Residual blocks and multilevel skip connections are used to preserve high-frequency information and stabilize training. To overcome the limited receptive field of spatial convolutions, a Frequency-domain Self-Attention Mechanism (FSAM) is introduced at the encoder output. Global feature interactions are modeled using the Fourier transform. The resulting attention responses are then mapped back to the spatial domain through the inverse Fourier transform to adaptively reweight feature representations. A two-stage training strategy is adopted. During pretraining, a dual-branch structure jointly learns mask geometry and imaging consistency, providing physically meaningful initialization. During main training, lithography simulation is applied under nominal, maximum, and minimum process corners to refine mask optimization under physical constraints.  Results and Discussions  The comparison results with baseline models are summarized in Tables 2 and 3. FMS-ILT is used as the reference method (Ratio = 1), and all experiments are conducted on the LithoBench dataset. For the overall imaging \begin{document}$ \mathcal{L}2 $\end{document} error, FMS-ILT achieves the lowest value of 19,998, outperforming the baseline models by 2%~107%. For Process Variation Band (PVB), GAN-OPC obtains the best value of 19 156, which is 31% lower than that of FMS-ILT. However, its \begin{document}$ \mathcal{L}2 $\end{document} error and Edge Placement Error (EPE) are 107% and 1 115% higher, respectively, indicating an imbalance between imaging fidelity and edge accuracy. The remaining baseline models show PVB performance comparable to that of FMS-ILT. For EPE, FMS-ILT also shows a clear advantage, achieving an average value of 1.95, which is 47%~1 115% lower than those of the baseline models. These improvements are mainly attributed to the multiscale encoder-decoder fusion mechanism, which integrates local and global features; the combination of attention mechanisms and frequency-domain operations, which guides the model toward critical regions; and the dual-branch pretraining strategy, which introduces physical priors into the network. These modules enable FMS-ILT to achieve balanced performance in imaging fidelity, process stability, and edge accuracy.  Conclusions  This work proposes FMS-ILT for mask optimization in computational lithography. The model uses a residual convolution-based multiscale encoder-decoder architecture to extract rich spatial features. It also incorporates FSAM to jointly model local geometric details and global optical interference. A two-stage training strategy is used. In the pretraining stage, mask generation and target image reconstruction are used as dual-branch tasks to improve the physical consistency between the mask and the printed image. In the main training stage, lithography simulation is introduced to further improve imaging accuracy and process robustness. Experimental results on the public LithoBench dataset show that FMS-ILT achieves strong performance in terms of L2, PVB, and EPE. The method improves printed image quality and provides a feasible and efficient solution for computational lithography.
A Lightweight and High-Reliability Challenge Generation Strategy for APUF
LAN Guohao, ZHANG Hui, DUO Bin, WANG Zibin, ZHOU Rang, LI Dongfen
2026, 48(7): 3203-3212.   doi: 10.11999/JEIT251073
[Abstract](392) [FullText HTML](247) [PDF 910KB](38)
Abstract:
  Objective  The Arbiter Physical Unclonable Function (APUF) is a lightweight security primitive widely used for identity authentication and key generation in resource-constrained devices. However, its response consistency is highly sensitive to environmental perturbations. The same challenge may therefore produce inconsistent responses under different conditions, which reduces the reliability of APUF-based security systems. Existing reliability improvement schemes mainly rely on hardware modification or challenge screening. These schemes often require high resource overhead and have low efficiency. To address these limitations, a Delay-constrained Challenge Generation Strategy (DCGS) is proposed to improve APUF reliability without additional hardware overhead or inefficient candidate screening.  Methods  DCGS models APUF path-delay characteristics and constructs challenges with constrained delay differences to ensure response stability. First, a Logistic Regression (LR) model is established to characterize the relationship between challenge bits and path delays. A delay-weight vector is then derived from the trained LR model to quantify the contribution of each challenge bit to the overall path delay. Second, a two-stage challenge generation mechanism is designed for delay-constraint control. In the first stage, prefix-bit initialization generates different prefix sequences to establish a delay baseline for subsequent bitwise extension. In the second stage, bitwise extension dynamically determines each remaining challenge bit according to the delay-weight vector. During this process, the cumulative delay difference of each challenge is monitored in real time and maintained within a preset delay-difference threshold range. Unlike conventional screening methods that post-process candidate challenges, DCGS directly generates stable challenges by design. This design removes the need for candidate challenge pools and improves generation efficiency.  Results and Discussions  DCGS is evaluated under different noise intensities. At a noise intensity of 0.3, which represents the maximum practical noise level, the reliability of DCGS-generated challenges remains 100% (Fig. 2). For generation efficiency, DCGS requires only 0.017 s to generate 10 000 challenges (Table 4). The response uniformity reaches 50.02% (Table 4), and the uniqueness reaches 50.46% (Table 4). Both metrics are close to the ideal theoretical value of 50%. The security analysis shows that the average bit entropy of DCGS-generated challenges is 0.980 7 (Fig. 3). The conditional entropy is 0.987 8, only 0.002 3 lower than that of random challenges (0.990 1).  Conclusions  This paper proposes DCGS for APUF to address inconsistent responses, low generation efficiency, and high hardware resource consumption in traditional schemes under high-noise conditions. By modeling path-delay characteristics with LR and combining prefix-bit initialization with bitwise extension, the proposed strategy ensures that the generated challenges satisfy the preset delay-difference threshold range. DCGS achieves high reliability, high efficiency, and good response uniformity without increasing hardware overhead. Experimental results show that DCGS improves APUF reliability in complex environments and supports secure applications in resource-constrained devices.
News
more >
Conference
more >
Author Center

Wechat Community