Advanced Search

Current Issue

2026 Vol. 48, No. 7

2026, 48(7): 1-1.
Abstract:
2026, 48(7): 1-4.
Abstract:
Excellence Action Plan Leading Column
A Survey of Processor Security
CHEN Congcong, GU Zhiyang, ZHANG Jiliang
2026, 48(7): 2765-2780. doi: 10.11999/JEIT260026
Abstract:
  Significance   Processor security is a cornerstone of modern information security. Cryptographic algorithms, operating systems, and applications have long relied on processors as trusted computing bases. However, as Moore’s Law slows, modern processors increasingly adopt aggressive microarchitectural optimization techniques to improve performance and energy efficiency, often without sufficient security consideration. This trend has led to frequent security vulnerabilities in recent years. In particular, microarchitectural timing channels, exemplified by Meltdown and Spectre, exploit timing differences caused by microarchitectural state changes to break fundamental hardware and software isolation, affecting billions of devices worldwide. At the same time, the boundary between architectural and microarchitectural behavior has become less clear, giving rise to new attack paradigms and turning timing channels from isolated hardware flaws into cross-layer system security problems.  Progress   Although substantial progress has been made in the study of timing channels, existing surveys still have several limitations. First, the mechanisms of timing channels are highly diverse, and the set of exploitable components continues to grow. Hardware-centric classification schemes are therefore insufficient to capture emerging and previously unknown attacks, and they often obscure the common features shared across different techniques. Second, as traditional microarchitectural channels become better understood and partially mitigated, leakage increasingly shifts to higher-level shared resources, including operating system policies and software-managed shared resources. However, previous studies have often treated software mainly as an execution context rather than a direct source of timing leakage. In addition, current discussions of defenses tend to emphasize individual techniques, with limited analysis of their scope and failure modes.  Contributions   This survey systematically reviews timing channels from a cross-layer perspective and unifies hardware- and software-based timing channels under a common abstraction. Four necessary conditions for timing channel exploitation are identified, and a unified classification framework is established based on the nature of shared mutable state and the mechanisms that make timing differences observable. Within this framework, representative attacks from the past decade are comprehensively reviewed, their attack procedures are systematically analyzed, and their common features are clarified. In addition, existing defense mechanisms are classified according to the leakage conditions they are intended to disrupt, and their scope and possible failure modes are examined. This survey also reviews current automated vulnerability detection methods.  Prospects   Future research on timing channels faces several emerging challenges. New microarchitectural optimization techniques continue to create new attack surfaces, while resource sharing at the software level may produce additional forms of timing leakage. Moreover, emerging platforms, including chiplet-based architectures, cloud computing environments, hardware accelerators, and heterogeneous systems, are likely to expose new types of timing channels that require systematic study.
Review of Non-invasive Brain-Computer Interfaces for Continuous Motor Control
XU Minpeng, JIA Leyi, ZHOU Xiaoyu, CHEN Enze, WANG Junyang, XIAO Xiaolin, MING Dong
2026, 48(7): 2781-2791. doi: 10.11999/JEIT260011
Abstract:
  Significance   Continuous motor control is a core capability of Brain-Computer Interface (BCI) systems for natural and efficient interaction with external devices. Compared with discrete command-based control, continuous control supports real-time and smooth regulation of motion parameters such as position, velocity, and trajectory. This capability is required for applications in assistive mobility, neurorehabilitation, robotic manipulation, and immersive human-machine interaction. Although invasive BCIs have achieved high-performance continuous control through high-quality neural recordings, their dependence on surgical implantation limits long-term use and large-scale deployment. A systematic review of non-invasive continuous motor control BCI technologies is therefore needed to clarify research progress, methodological features, and remaining challenges.  Progress   Advances in non-invasive continuous motor control BCIs are reviewed from four closely related aspects: control paradigms, decoding algorithms, applications, and performance evaluation. At the paradigm level, motor imagery, steady-state visual evoked potentials, P300, and hybrid paradigms have been studied to support continuous control through sustained intention modulation, dynamic stimulus encoding, and hierarchical or shared-control strategies. For decoding algorithms, two main frameworks are identified: motion parameter mapping and motion parameter regression. Motion parameter mapping generates continuous output by temporally integrating discrete classification results or mapping them to velocity or state variables, whereas motion parameter regression directly establishes relationships between Electroencephalogram (EEG) features and continuous kinematic parameters. Recent studies increasingly incorporate nonlinear models and deep learning methods to improve robustness under the non-stationary nature of EEG signals. At the application level, non-invasive continuous control has progressed from two-dimensional cursor tasks to more practical scenarios, including wheelchair navigation, robotic arm manipulation, unmanned systems, and virtual or augmented reality environments. Existing studies also assess continuous control performance using both objective and subjective indicators, including trajectory error, task success rate, information transfer rate, workload, and user experience, reflecting varied experimental designs and control aims.  Conclusions  Existing studies show that non-invasive BCIs can support continuous motor control. However, current research remains at a stage in which multiple methods coexist without a unified framework. At the paradigm level, available approaches differ in their ability to elicit and sustain continuous motor intention reliably. For decoding algorithms, both motion parameter mapping and motion parameter regression are limited by the non-stationary nature of EEG signals, which affects robustness, generalization, and long-term stability. At the application level, many studies remain restricted to specific tasks and controlled environments, and the transfer of continuous control strategies to complex real-world scenarios still requires further validation. Moreover, the lack of standardized evaluation protocols hinders direct comparison and systematic optimization across studies.  Prospects   Future research should improve the stability and reliability of continuous control paradigms, enhance decoding robustness under realistic EEG conditions, and strengthen the match between control strategies and application requirements. Unified evaluation frameworks that integrate objective and subjective indicators should also be established to support methodological convergence and fair comparison. With continued progress, non-invasive continuous motor control BCIs are expected to play a growing role in assistive technologies, rehabilitation systems, and advanced human-machine interaction.
Datasets
TTSPD: A Multimodal Traffic Scene Perception Dataset Integrating Tire Data
YING Zongchen, GUI Lin, YANG Jiahan, ZHANG Fangwei, WANG Junfan, DONG Zhekang
2026, 48(7): 2792-2804. doi: 10.11999/JEIT260022
Abstract:
  Objective  With the rapid development of Intelligent Transportation Systems (ITS) and autonomous driving technologies, accurate traffic environment perception is a fundamental prerequisite for vehicle safety and decision making. Current perception frameworks primarily rely on high-resolution cameras and LiDAR sensors. Although these sensors provide rich information, they create severe challenges across the Perception-Storage-Calculation pipeline. High acquisition costs limit large-scale deployment. In addition, the massive data volume produced by high-dimensional sensors places heavy pressure on onboard storage and computational resources, often exceeding the power and thermal budgets of vehicle-grade edge platforms. These constraints motivate the exploration of alternative sensing paradigms that are cost-effective, compact, and computationally efficient while maintaining reliable perception accuracy. In response, the present study shifts the perception perspective from conventional external sensors to the tire-road contact interface, where abundant physical interaction information naturally exists. The objective is to construct a novel multimodal dataset, termed the Tire-integrated Traffic Scene Perception Dataset (TTSPD), which combines internal tire dynamics with external visual observations. This dataset is used to examine whether low-dimensional tire sensing data can complement or partially substitute high-dimensional visual data for accurate road surface classification. The study also aims to establish a new data morphology that balances perception performance and system efficiency for future intelligent vehicles.  Methods  To construct a high-quality and practically usable multimodal dataset, an integrated hardware-software acquisition framework is developed. From a hardware perspective, a specialized sensing system is designed by coupling tire-mounted multi-parameter sensors with a vehicle-mounted camera. To ensure reliable operation under the harsh mechanical conditions of a rotating tire, sensing nodes are encapsulated using a rubber-based composite material that provides mechanical protection and long-term stability. Wireless transmission is implemented using Bluetooth Low Energy (BLE) 5.0 with an adaptive frequency-hopping mechanism, enabling low-power and reliable communication during high-speed rotation. During data acquisition, the system synchronously collects six types of internal tire signals, including radial acceleration, tire temperature, and tire pressure, producing approximately 1.8 million sampling points. In parallel, a dashboard-mounted camera records high-resolution traffic scene images totaling 309 GB across four representative road surface conditions. To address the heterogeneity between high-frequency one-dimensional tire signals and two-dimensional visual data, a timestamp-based association strategy is adopted to achieve scene-level temporal alignment rather than strict frame-by-frame correspondence. Sensor sequences and image segments are grouped according to shared temporal windows and driving scenarios. This approach ensures semantic and temporal consistency at the scene level. The alignment strategy reflects practical deployment conditions and forms the basis of the final TTSPD dataset for multimodal fusion research.  Results and Discussions  The effectiveness of the proposed TTSPD is evaluated through comprehensive road surface classification experiments using mainstream deep learning models. Initial experiments based solely on visual data demonstrate strong baseline performance, with classification accuracies ranging from 88.50% to 93.75% (Table 7). These results confirm the quality and diversity of the visual modality in the dataset. The primary contribution of this study is the quantification of efficiency gains enabled by tire-based sensing. Comparative experiments progressively reduce the amount of visual data while integrating low-dimensional tire signals, particularly radial acceleration (Table 9). The results show that the multimodal model achieves approximately 95% of the full-data baseline accuracy while using only about 38.75% of the original data volume. This reduction in data dependency produces significant system-level benefits. Storage requirements decrease by approximately 61.25%, and overall model training time decreases by about 54.10% (Fig. 8). These findings indicate that tire dynamics encode high-value physical features related to road texture and surface conditions that complement visual cues. The proposed dataset therefore supports the development of lighter perception pipelines without reducing recognition performance.  Conclusions  This study addresses the long-standing Perception-Storage-Calculation bottleneck in vision-dominated autonomous driving systems by proposing the TTSPD. Multi-parameter sensors are embedded within tires using rubber-based encapsulation, and stable wireless communication is achieved through BLE 5.0. A robust tire-camera data acquisition system is therefore established. The resulting dataset covers four common and safety-critical road surface types: cement, asphalt, damaged, and water-covered roads. It provides a comprehensive foundation for multimodal perception research. Experimental results show that combining low-dimensional tire sensing data with visual information significantly improves perception efficiency. Approximately 95% of peak classification accuracy is achieved using only about 38.75% of the original data volume. This result effectively reduces storage pressure and computational cost, reflected in a 61.25% reduction in data storage and a 54.10% reduction in training time. The TTSPD dataset therefore proposes a practical data morphology that supports efficient and high-performance perception under vehicle-grade computational constraints. It also provides valuable resources for the future development of ITS.
Image and Intelligent Information Processing
A Processing-In-Memory Neural Network Inference System Design for Infrared Gesture Recognition
SHI Xiangyang, LIU Jinchang, LU Qiulin, WANG Ziang, HAN Yongkang, GAO Gen, SUN Haoran, JIANG Xiaoyong, SHI Tuo, LI Qing, MIAO Jinshui
2026, 48(7): 2805-2814. doi: 10.11999/JEIT260122
Abstract:
  Objective  Infrared (IR) image recognition is widely used in security monitoring, autonomous driving, and human-machine interaction, where stable perception under low illumination and noisy backgrounds is required. However, IR images often exhibit low contrast, blurred edges, and weak texture representation, which reduce recognition accuracy. Conventional Deep Neural Networks (DNNs) further aggravate this issue in edge environments because the von Neumann architecture requires intensive computation, frequent memory-processor data transfer, and high power consumption. Memristor devices provide an alternative because they support in-memory computing. However, most existing implementations still depend on CPUs or GPUs to execute nonlinear operations such as activation layers, which reintroduces data transfer overhead. To address this issue, a nearly all-memristor neural network framework for IR image recognition is proposed. In this framework, computationally dominant linear operations are executed entirely on memristor arrays, whereas CPU participation is limited to the final activation step.  Methods  The proposed system maps two fully connected layers of a neural network directly onto memristor crossbar arrays. This mapping enables large-scale matrix-vector multiplications to be executed in memory with high parallelism and low energy consumption. Network weights are encoded as memristor conductance states, and inference is performed by applying voltage inputs and summing output currents. Intermediate nonlinear activations and final classification are computed on the host computer. Because these operations require minimal computation, the CPU overhead in both latency and energy remains negligible. Based on this design, an IR gesture recognition system is constructed to distinguish two hand gestures and a no-gesture state. System evaluation considers recognition accuracy, inference latency, energy consumption, and the amount of data transferred between the memristor arrays and the CPU. A CPU-only neural network is used as the baseline. The framework provides a practical approach for near all-hardware neural network computing while maintaining recognition accuracy and reducing energy consumption and data transfer for edge-scale IR applications.  Results and Discussions  The system is evaluated on a three-class IR gesture recognition task that includes two gestures and a no-gesture state. The memristor-based network achieves approximately 95% accuracy (Table 2), which is close to the 97% obtained with a GPU implementation (Table 2). This result indicates that conductance variation has limited effect on recognition performance. The average inference latency per frame is substantially lower than that of the GPU baseline, and the achieved frame rate satisfies real-time requirements. Power measurements indicate that the memristor array consumes only milliwatts, whereas the GPU requires approximately 25 W (Table 2). Although peripheral circuit consumption is not fully included, the results demonstrate the inherent energy efficiency of in-memory computing. Intermediate and final activation outputs are transferred to the CPU, which removes most memory-processor interactions. This reduction in data movement, combined with array-level parallelism, accounts for the observed improvements in latency and energy consumption. Overall, the framework maintains high recognition accuracy while significantly improving computational efficiency, which indicates strong potential for edge-scale IR recognition.  Conclusions  This study presents a nearly all-memristor infrared neural network framework in which two fully connected layers are executed on memristor arrays and intermediate nonlinear activations are processed on the host computer. When applied to a three-class IR gesture recognition task, the system achieves recognition accuracy comparable to that of a GPU platform while significantly reducing inference latency and energy consumption. By minimizing memory-processor data transfer and exploiting in-memory computing, the framework provides clear advantages for edge applications. The results confirm the feasibility of deploying memristor-based neural networks in practical infrared recognition systems. Future research will focus on integrating memristor-based activation functions, scaling the system to larger circuits, and extending the approach to more complex network architectures and datasets.
Explicit Discrimination-driven Automatic Unknown Class Clustering for Open-World Semi-Supervised Learning
SONG Jialun, DU Lan, CHEN Jian
2026, 48(7): 2815-2828. doi: 10.11999/JEIT251291
Abstract:
  Objective   Traditional target recognition is generally developed under the closed-set assumption, in which all test classes are assumed to be included in the training set. In real-world scenarios, however, unlabeled unknown classes are commonly encountered. Therefore, recognition systems should simultaneously recognize known classes and discover and cluster unknown classes. To address this challenge, a transductive Open-World Semi-Supervised Learning (OWSSL) method driven by explicit discrimination and automatic unknown-class clustering is proposed.   Methods   The proposed method is trained using a small set of labeled known-class samples and a large set of unlabeled test samples containing both known and unknown classes. It consists of two complementary modules. The Dynamic Known-Unknown Class Discrimination (DKUCD) module models the known-class boundary distribution using Extreme Value Theory (EVT) and progressively refines the distribution with high-confidence known-class samples during semi-supervised learning to improve known-unknown discrimination. The Neighbor Intersection-Over-Union Cluster Merging (NIOUCM) module automatically clusters high-confidence unknown-class samples by merging neighboring clusters according to their intersection-over-union relationships. The DKUCD and NIOUCM modules are optimized iteratively to improve discrimination and unknown-class clustering jointly.   Results and Discussions   Experiments conducted on the optical CIFAR-10 dataset and measured radar datasets demonstrate that the proposed method achieves accurate known-class recognition while effectively clustering unknown classes.   Conclusions   By explicitly discriminating between known and unknown classes and automatically estimating unknown-class clusters, the proposed method improves both known-class recognition and unknown-class clustering under open-world conditions.
Spatial Information-guided Diffusion for Domain Adaptation Semantic Segmentation of Remote Sensing Images
LIANG Yan, LI Junfan, SHAO Kai, HU Lin
2026, 48(7): 2829-2842. doi: 10.11999/JEIT260031
Abstract:
  Objective  Domain Adaptation Semantic Segmentation (DASS) is critical for remote sensing applications, including land-cover mapping, urban planning, and environmental monitoring. However, deep learning models often show severe performance degradation under domain shifts caused by imaging variation, geographic differences, and label-semantic heterogeneity. Conventional feature-alignment and generative adversarial network-based methods often fail to preserve semantic consistency. They are also sensitive to noisy supervision, especially when cross-domain gaps are large. This work aims to construct a robust DASS framework for semantically consistent image translation and reliable knowledge transfer.  Methods  A two-stage framework, termed Co-training Spatial-Guided DASS (CoSG-DASS), is proposed by integrating image translation and co-training. In the image-translation stage, a spatial information-guided latent diffusion model enhanced by ControlNet is designed. Semantic pseudo-labels and depth estimates are used as horizontal semantic and vertical spatial conditions to guide target-style image generation. To reduce the effect of noisy pseudo-labels, an Entropy-based Adaptive Guidance Intensity Module (EAGIM) is introduced. EAGIM estimates pixel-level confidence using information entropy and suppresses unreliable features. In the co-training stage, translated target-style images and unlabeled real target-domain images are used to train a segmentation model with a depth-guided segmentation head. Cross-entropy loss and adversarial loss are jointly used for optimization.  Results and Discussions  Extensive experiments are conducted on three cross-domain tasks. CoSG-DASS generates images that better match target-domain distributions. Quantitative results based on Fréchet Inception Distance (FID) show that the proposed method outperforms CycleGAN, UNI-Diff, and CRS-Diff in most settings (Table 1). Visual comparisons (Fig. 6) show that the method reduces edge blurring and category confusion. It also improves the separation of roads and vegetation and preserves small objects, such as vehicles. In the semantic segmentation stage, CoSG-DASS outperforms state-of-the-art domain adaptation methods. It improves mean Intersection over Union (mIoU) by 1.14%, 3.78%, and 2.49% on the cross-geographic task (Vaihingen IRRG→Potsdam IRRG), cross-imaging-mode task (Vaihingen IRRG→Potsdam RGB), and bidirectional label-semantic-heterogeneity tasks between DFC25 and LoveDA, respectively (Tables 24). Visual segmentation results (Fig. 7) confirm its strong boundary preservation and high accuracy in complex scenes. Ablation studies (Table 5) verify the contribution of the core components, including depth control, pseudo-label guidance, EAGIM, and the co-training strategy. Feature-distribution visualization based on Uniform Manifold Approximation and Projection (UMAP) further shows that CoSG-DASS reduces intra-class variation and increases inter-class separation after adaptation (Fig. 8).  Conclusions  CoSG-DASS alleviates domain shifts in remote sensing images through semantic-preserving diffusion-based translation and depth-guided co-training. It improves both image-translation quality and segmentation accuracy over existing methods. The proposed framework provides an effective solution for multi-source remote sensing interpretation. Future work will focus on extreme label-semantic heterogeneity and lightweight diffusion architectures.
A Social-Aware Ant Colony Optimization Algorithm with Reproductive Division of Labor for MCS Task Allocation
SHEN Xiaoning, SHE Juan, WANG Zhilong, LI Jiayuan
2026, 48(7): 2843-2853. doi: 10.11999/JEIT260018
Abstract:
  Objective  With the rapid development of handheld and wearable smart devices, Mobile Crowd Sensing (MCS) has become an efficient data collection paradigm. Effective task allocation can improve system efficiency, requester and participant satisfaction, and platform sustainability. Existing models often neglect task skill requirements, do not use participants’ social networks as auxiliary execution resources in emergencies, and overlook the effect of collaboration efficiency on team-task quality. To address these issues, this paper proposes a Social-Aware MCS Task Allocation model (SAMCSTA) with two objectives: maximizing total platform revenue and total task sensing quality. Social networks are used to build a two-layer collaboration framework of platform participants and social-network friends, which expands available execution resources and improves allocation flexibility. For complex tasks, participant sensing capability is quantified, and collaboration efficiency is introduced to optimize team composition.  Methods  This paper proposes a Multi-objective Ant Colony Optimization based on Reproductive Division of Labor (MACORDL) algorithm. The main innovations are as follows. First, the ant colony is divided into four collaborative subpopulations: queen ants, male ants, scout ants, and worker ants. Local enhancement, memetic crossover, knowledge transfer, and other search strategies are designed for these subpopulations to form a hierarchical collaborative search framework. Second, a statistical-learning-based mating selection strategy is designed to support intelligent transfer of elite genes. Third, the short-term contribution of each subpopulation is predicted from historical performance, which enables dynamic and adaptive allocation of computational resources. Fourth, a cooperative update mechanism for node pheromones and participant pheromones is designed to establish a dual-layer search guidance system.  Results and Discussions  The evaluation uses 8 synthetic instances and 4 real-world instances. Performance is measured by HyperVolume Ratio (HVR) and Inverted Generational Distance (IGD). The Wilcoxon rank-sum test at a significance level of 0.05 is used for statistical comparison. The results show that MACORDL achieves the best HVR and IGD on most instances (Table 2, Table 3). On average, MACORDL improves HVR and IGD by 16.41% and 18.04%, respectively, compared with the second-best algorithm. Visual comparisons further show that the Pareto front obtained by MACORDL has better convergence, distribution uniformity, and breadth (Fig. 4). Although its fine-grained local search can still be improved for a few large-scale instances, MACORDL shows stable performance and good scalability across different problem scales. It helps the platform obtain task allocation schemes with higher revenue and better sensing quality.  Conclusions  This paper studies the task allocation problem in MCS systems by considering interactions among platform participants and between participants and their social-network friends. A social-aware MCS task allocation model is established, and MACORDL is proposed to solve it. Comparative experiments on 8 synthetic instances and 4 real-world instances with different scales show that MACORDL outperforms six representative algorithms on most instances. It obtains allocation schemes and paths that yield higher total platform revenue and better task sensing quality, indicating good scalability. MACORDL uses multiple strategies to balance local exploitation and global exploration. However, the current model assumes that all tasks are released at the initial stage and that complete information is available. Participant privacy protection is also not considered. Future work will focus on MCS task allocation models in dynamic and uncertain environments and on privacy-preserving distributed optimization.
Multi-dimensional Spatio-temporal Feature Enhancement for Lip Reading
MA Jinlin, ZHONG Yaowei, MA Ruishi
2026, 48(7): 2854-2864. doi: 10.11999/JEIT251111
Abstract:
  Objective  Lip reading is a challenging yet important task in computer vision that aims to decode spoken language solely from visual lip movements. The task is difficult mainly because of inherent ambiguity in visual speech signals. On one hand, articulatory movements for different visemes can be highly subtle. For instance, lip displacement differences for confusable pairs such as /p/-/b/ and /m/-/n/ may be as small as 0.3~0.7 mm. These fine-grained spatial variations often fall below the effective resolution limits of conventional 3D convolutional neural networks. On the other hand, natural coarticulation in speech introduces temporal ambiguity, as mouth shapes may transiently blend multiple phonemes and make distinct visual units difficult to separate. These challenges are further aggravated by real-world factors such as uneven lighting and substantial inter-speaker differences in articulation. Current lip-reading models therefore often show limited ability to capture discriminative spatiotemporal features, which leads to suboptimal performance, particularly for phonemes with minimal visual differences. To address these issues, a robust lip-reading framework is developed to capture and exploit fine-grained spatiotemporal dependencies and improve recognition accuracy under diverse and realistic conditions.  Methods  To address the above limitations, a novel lip-reading framework, termed the Multi-dimensional Spatio-Temporal Enhancement Network (MSTEN), is proposed. The framework is designed to strengthen spatial and temporal representations through integrated attention mechanisms and advanced residual learning. It contains three core components that collaboratively model dependencies between spatial and temporal features, which are often insufficiently used in conventional architectures. The first component, the Self-adjusting Spatio-temporal Attention (SaSTA) module, adopts a self-adjusting mechanism that operates simultaneously across the height, width, and temporal dimensions. Query, key, and value tensors are generated through 1×1×1 3D convolutions, flattened across spatial and temporal dimensions, and used to compute attention weights by multiplying the query tensor with the transposed key tensor, followed by softmax normalization. The resulting attention map is multiplied by the value tensor and then combined with the original input through learnable parameters and a residual connection, thereby preserving contextual information and producing globally enhanced features. The second component, the Three-dimensional Enhanced Residual Block (TE-ResBlock), improves spatiotemporal feature extraction through temporal shift, multi-scale convolution, and channel shuffle. The temporal shift operation moves one quarter of the feature channels along the time axis to fuse information from adjacent frames without adding parameters. Multi-scale convolution adopts parallel branches with kernel sizes of 3×3, 3×1, 1×3, and 1×1 to capture features at different receptive fields. The outputs are concatenated and processed by channel shuffle to improve information exchange across groups, and four TE-ResBlocks are stacked to achieve progressive feature refinement. The third component, the Multi-dimensional Adaptive Fusion (MDAF) module, integrates spatial, temporal, and channel information through three submodules. These are a Channel Enhancement Module (CEM), which recalibrates features through max pooling, temporal convolution, and sigmoid activation, a Spatial Enhancement Module (SEM), which expands the receptive field through identity mapping and both standard and dilated convolution, and an Adaptive Temporal Capture Module (ATCM), which emphasizes dynamic movements through frame-difference features and temporal weight maps. The MDAF modules are inserted between TE-ResBlock stacks for iterative refinement. Finally, features extracted by the MSTEN front end are fed into a Densely Connected Temporal Convolutional Network (DC-TCN) back end, which consists of four blocks, each containing three temporal convolutional layers with dense connections, to model long-range phonological dependencies effectively.  Results and Discussions  The proposed framework is evaluated comprehensively on the widely used LRW and GRID datasets. The LRW dataset contains more than 500 000 video clips from over 1 000 speakers. The GRID dataset contains video clips from 34 speakers, each providing 1 000 utterances, with a total duration of 28 h. The proposed model achieves an accuracy of 91.18%, which is an absolute improvement of 2.82 percentage points over a strong ResNet18 baseline, demonstrating its substantial effectiveness. Ablation studies are further conducted to analyze the contribution of each key component. The results show clearly that each proposed module yields a meaningful performance gain. Specifically, the SaSTA module alone improves accuracy by 2.09%, which confirms the key role of global spatiotemporal attention. The TE-ResBlock increases accuracy by 1.73%, which verifies its effectiveness in multi-scale local feature extraction and inter-frame information fusion. The MDAF module provides a further 1.74% improvement, which highlights the benefit of adaptive multi-dimensional feature fusion, as shown in Table 2.  Conclusions  This study advances lip reading through the proposed MSTEN front-end network. The framework is built on three main contributions. First, the SaSTA module proposes an effective mechanism for global context aggregation and performs multi-dimensional feature weighting across height, width, and temporal sequences. Second, the TE-ResBlock addresses central challenges in spatiotemporal modeling through the combined use of temporal displacement, multi-scale convolution, and enhanced channel interaction. Third, the MDAF module enables deep and coordinated integration of spatial, temporal, and channel information. Together, these components improve model performance substantially and achieve accuracies of 91.18% on the challenging LRW dataset and 97.82% on the GRID dataset. Ablation studies further confirm the individual and combined effectiveness of the proposed components. Future work will examine the extension of this framework to audio-visual speech recognition under noisy conditions and the development of domain adaptation strategies to improve robustness in low-resolution or resource-constrained scenarios.
Network Metric System and Scenario-Differentiated Analysis Driven by LLM Literature Mining
XU Qikun, LIU Yaxi, HAN Shuxian, ZHANG Huifeng, HUANGFU Wei
2026, 48(7): 2865-2875. doi: 10.11999/JEIT251120
Abstract:
  Objective  Network metrics provide the foundation for network design, operation, and optimization. Existing studies primarily focus on individual scenarios or representative metrics and lack unified extraction rules and reproducible workflows for large-scale, cross-scenario metric analysis. To address terminology ambiguity, scenario heterogeneity, and the quantification of complex metric relationships, this study proposes a reproducible domain-specific literature mining framework based on a Large Language Model (LLM). The framework automatically extracts and standardizes network metrics, annotates application scenarios, quantifies inter-metric relationships, and establishes a Service-Multiplexing-Versatility (SMV) analytical framework. Rather than providing a complete set of metric calculation methods, the SMV framework serves as a conceptual model for guiding multi-objective tradeoffs in network architecture design and lifecycle management.  Methods  An automated literature mining framework based on a multi-agent LLM architecture is developed (Fig. 1). A dataset comprising 583 articles published in IEEE/ACM Transactions on Networking during 2023~2024 is analyzed. The framework consists of three specialized agents. A terminology normalization agent maps aliases and synonymous expressions to standardized metric names. A scenario annotation agent assigns primary application scenario labels using high-information-density sections of each article. A correlation mining agent identifies the semantic direction and strength of relationships between metric pairs and quantifies these relationships as signed correlation coefficients ranging from –1 to +1. The reliability of the mining results is evaluated through dual-LLM cross-validation and manual sampling review (Fig. 2).  Results and Discussions  The proposed framework extracts 3 978 independent network metrics, of which 138 appear in more than 1% of the analyzed articles (Fig. 3). The metric frequency distribution exhibits a pronounced heavy-tailed distribution, with throughput (79.1%), end-to-end delay (74.6%), and packet error rate (59.5%) representing the most frequently studied metrics (Table 1). The core metric sets show strong scenario dependence (Fig. 4). For example, data center networks primarily emphasize throughput, end-to-end delay, and flow completion time, whereas Internet of Things (IoT) applications additionally prioritize energy consumption and network lifetime (Table 2). Furthermore, scenario-specific correlation matrices reveal markedly different coupling patterns among metrics (Figs. 5 and 6). In data center networks, throughput is strongly negatively correlated with flow completion time and queueing delay, reflecting the fundamental tradeoff associated with congestion control. In edge computing networks, end-to-end delay is negatively correlated with resource utilization, indicating the balance between real-time task offloading and resource utilization.  Conclusions  The strong coupling between network metrics and application scenarios indicates that future network architectures should be evaluated from a multidimensional perspective. Based on the extracted scenario-specific metric relationships, this study proposes the SMV analytical framework (Fig. 7). By jointly considering differentiated service quality requirements (Service), physical infrastructure cost and resource reuse (Multiplexing), and adaptive reconfiguration capability for emerging services and application scenarios (Versatility), the framework provides a theoretical basis for adaptive resource orchestration in AI-native networks. Future work will extend the current static literature mining pipeline into a continuously updated network metric knowledge base and further validate the engineering applicability of the SMV framework in programmable and multimodal networks.
UWF-YOLO: A Lightweight Framework for Underwater Object Detection via Redundant Information Optimization
HOU Guojia, MA Jiaqi, WANG Yuechuan, HUANG Baoxiang, LI Kunqian
2026, 48(7): 2876-2886. doi: 10.11999/JEIT251129
Abstract:
  Objective  The rapid development of underwater imaging technology has increased the significance of underwater object detection for resource exploration and environmental monitoring. Complex underwater environments often degrade image quality through color casts, haze-like effects, and non-uniform illumination. These factors reduce the performance of existing vision-based object detection algorithms, particularly for small objects, and often lead to missed detections and false positives. In addition, current deep learning-based underwater detection models face difficulty balancing detection accuracy and lightweight design under limited computational resources. Therefore, efficient underwater object detection methods are required for water-related vision tasks. Such methods support marine resource exploration, ecological monitoring, underwater robotics, and perception systems for autonomous underwater vehicles.  Methods  A lightweight framework based on redundant information optimization is proposed for underwater object detection. Specifically, a lightweight underwater object detection network, termed UWF-YOLO, is designed based on redundant information optimization. First, the C2f module is reconstructed using the FasterNet Block to optimize both the backbone and neck networks. A feature channel selection mechanism is integrated to reduce redundant feature representations. Furthermore, redundant convolutional features in the conventional YOLO neck limit adaptation to underwater environments. Therefore, Ghost Convolution is introduced to generate Ghost feature maps and improve the multi-scale feature fusion capability of the neck network. Next, parameter sharing is achieved by replacing the original detection head with a redundant optimization group detection head (RRG-Head) based on group convolution, which reduces computational cost. Finally, a structured channel pruning strategy is applied to identify inter-layer dependencies in the computational graph and bind pruning units. Combined with LAMP weight magnitude score normalization to evaluate channel importance, low-contributing groups are pruned and subsequently fine-tuned to compress the network size. In addition, existing underwater detection datasets usually contain monotonous scenes, and the objects are typically small and densely distributed. To address this limitation, an underwater object detection dataset with complex scenes, termed CSUOD, is constructed by collecting real-world underwater images from various websites and platforms. Manual annotation and resolution normalization are then performed to ensure dataset consistency. CSUOD is designed for challenging underwater environments characterized by color casts, haze-like effects, and non-uniform illumination. A total of 1 135 images containing six object categories are manually selected and annotated.  Results and Discussions  Extensive experiments are conducted on three public underwater object detection datasets, namely DUO, RUOD, and TrashCan, and several widely used detection methods are compared. The proposed model is evaluated against mainstream detectors, including YOLOv5s, YOLOv7-tiny, YOLOv8s, YOLOv9-tiny, and Deformable DETR. In terms of computational complexity, the proposed method reduces FLOPs, model size, and parameters by 60.4%, 77.3%, and 78.4%, respectively, compared with the baseline model. Furthermore, the proposed method outperforms YOLOv9-tiny with comparable parameters by 0.3%, 2.3%, and 3.4% in mAP on the three datasets. Additional comparative experiments on the constructed CSUOD dataset also demonstrate improved performance and stable detection capability in complex underwater environments. Qualitative visualization results further demonstrate the robustness and detection stability of the model under various underwater degradations, including haze-like effects and non-uniform illumination.  Conclusions  Quantitative and qualitative experiments on multiple datasets validate the effectiveness and robustness of the proposed method. The proposed framework achieves superior detection performance in complex underwater environments and reduces missed detections and false positives caused by background interference. Experimental results indicate that the proposed UWF-YOLO achieves significant model lightweighting while maintaining detection accuracy comparable to benchmark models. This balance between detection accuracy and low computational cost makes the framework suitable for underwater devices with limited resources. The proposed method also shows strong potential for practical applications such as marine ecological monitoring, underwater resource exploration, and perception systems for autonomous underwater vehicles. It provides a reliable technical foundation for real-time applications, supports integration into embedded platforms, and enables real-time perception and decision-making under different underwater conditions. In addition, the constructed CSUOD dataset helps address the limitations of existing underwater detection datasets and supports further research in underwater object detection. Future work will extend this framework to multi-modal perception systems and larger-scale datasets, enabling adaptive models for dynamic underwater scenarios and supporting broader applications in intelligent ocean observation and autonomous navigation.
A Multimodal Sentiment Analysis Model with Multi-source Knowledge guided Visual Confidence Perception
PENG Juhong, ZHANG Zhi, LIU Peng, GE Wenhui, LIU Chen, LIAO Lingxin, ZHANG Kai
2026, 48(7): 2887-2897. doi: 10.11999/JEIT260063
Abstract:
  Objective  Multimodal sentiment analysis is often affected by visual noise from complex environments, image-text sentiment inconsistency, and imbalanced modality contributions. When all modalities are treated without distinction, visual noise can degrade model performance. A robust mechanism is therefore needed to evaluate visual confidence and filter redundant visual information.  Methods  A Multimodal Sentiment Analysis Model with Multi-source Knowledge-guided Visual confidence Perception (MKVP) is proposed (Fig. 1). A multi-source knowledge guidance matrix is constructed using syntactic-dependency, sentiment-intensity, and aspect-focused operators (Fig. 2). Guided by this matrix, the Visual Confidence Perception (VCP) module measures semantic affinity and dynamically suppresses irrelevant visual noise (Fig. 3). A dual-stream parallel interaction module is then used to support deep cross-modal alignment, and a global gated fusion mechanism further adjusts the fusion weights of different modalities.  Results and Discussions  Extensive experiments are conducted on the MVSA-Single, MVSA-Multiple, and HFM datasets. The proposed MKVP model achieves accuracy and F1 scores of 77.56% and 76.70%, 72.72% and 70.66%, and 87.26% and 86.78%, respectively. Compared with the baseline models, the accuracy and F1 score are improved by 2.45% and 3.68%, 2.19% and 2.21%, and 1.83% and 1.91%, respectively (Table 3). Ablation studies show that each component contributes to performance, especially the VCP module, which filters visual noise and improves feature quality (Table 5). Feature-space visualization further confirms that the VCP module refines semantic representations by promoting clearer clustering of samples with the same sentiment polarity (Fig. 4). Case studies on mismatched image-text samples also verify the ability of the model to resolve cross-modal semantic conflicts (Table 6). Model-complexity analysis shows that MKVP maintains high computational efficiency and low inference latency (Table 8).  Conclusions  The proposed MKVP framework reduces the effects of visual noise and image-text sentiment inconsistency in multimodal sentiment analysis. By using multi-source knowledge to guide visual confidence perception and combining dual-stream interaction with dynamic gated fusion, the model learns robust sentiment representations from noisy multimodal data. This method provides an efficient and reliable solution for complex social media scenarios.
Multipath Scheduling Algorithm for UAV Video Streaming
CAO Changlong, LI Lingzhi, SHI Lianmin, ZHAO Qingyue
2026, 48(7): 2898-2908. doi: 10.11999/JEIT260002
Abstract:
  Objective   With the rapid growth of the low-altitude economy, Unmanned Aerial Vehicle (UAV) technology has been widely used in emergency rescue, disaster monitoring, urban security, and other applications. In these scenarios, stable, low-latency, and high-fidelity video backhaul is critical for task execution. Multipath transport protocols can improve Quality of Experience (QoE) through bandwidth aggregation, providing an effective basis for UAV video streaming. However, under dynamic and heterogeneous network conditions, the performance of multipath transport protocols depends strongly on the design of multipath scheduling algorithms. Existing heuristic schedulers use predefined rules to reduce head-of-line blocking and inter-path load imbalance, but their adaptability remains limited in highly dynamic environments. Learning-based schedulers can learn the mapping between network states and scheduling rewards from real-time feedback, enabling adaptive performance optimization. However, most existing learning-based schedulers are designed for general network scenarios. They are not optimized for UAV networks, and their ability to guarantee QoE has not been fully validated. A multipath scheduling algorithm tailored to UAV video streaming is therefore needed to better exploit the performance potential of multipath transport protocols.  Methods   To address the dynamic and heterogeneous challenges of UAV video streaming, this paper proposes NeuroFly, a multipath scheduling framework based on the NeuralUCB algorithm. In NeuroFly, multipath traffic scheduling is formulated as a Contextual Multi-Armed Bandit (CMAB) problem. The context space is constructed by integrating path state information, video encoding features, and UAV mobility parameters, which jointly characterize the current transmission environment. In the action space, a frame-priority-driven redundant transmission mechanism is proposed. Video frames are assigned different frame priorities according to decoding dependencies, and differentiated redundancy strategies are used to improve the probability of successful video-frame delivery. A multi-objective reward function is further designed to guide policy learning and support adaptive optimization under dynamic and heterogeneous network conditions. In addition, a context monitoring mechanism is integrated into NeuroFly to handle abrupt environmental changes caused by high UAV mobility. This mechanism detects context distribution shifts and triggers a two-stage restart strategy. A soft restart is activated when gradual context drift is detected, removing outdated historical experience. A hard restart is performed under abrupt context changes by clearing the experience replay buffer and reinitializing model parameters, allowing learning to restart under a new distribution.  Results and Discussions   The proposed NeuroFly framework is evaluated in both simulation and field environments. First, Mininet-WiFi is used to simulate realistic UAV network environments and evaluate overall QoE performance. The results (Fig. 4) show that, compared with state-of-the-art heuristic and learning-based schedulers, NeuroFly achieves broad performance gains by fully using aggregated multipath bandwidth. Specifically, the 99th-percentile latency is reduced by 19.9%~51.0%, the average video frame rate is increased by up to 24.6%, image structural similarity is improved by up to 49.2%, and the buffering time ratio is reduced by 13.4%~77.6%. These results demonstrate the strong ability of NeuroFly to guarantee QoE. Field experiments (Fig. 6) further confirm that NeuroFly provides favorable optimization in real UAV operation scenarios. Compared with mainstream transport solutions widely deployed in production environments, NeuroFly achieves better real-time transmission performance and shows strong practical applicability for future large-scale UAV deployment.  Conclusions   This paper addresses network dynamics, path heterogeneity, and time-varying transmission conditions in UAV video streaming over multipath transport protocols. An intelligent multipath scheduling framework, NeuroFly, is proposed based on the NeuralUCB algorithm. In this framework, multipath traffic scheduling is modeled as a CMAB problem. Through the design of the context space, action space, and multi-objective reward function, online learning and adaptive optimization of traffic allocation policies are achieved. To further improve robustness under severe environmental changes, a lightweight context monitoring mechanism is introduced to detect context distribution drift and restart the learning process when needed. Systematic evaluations are conducted on both simulation platforms and real UAV operation environments. The simulation results show that NeuroFly achieves consistent improvements across QoE metrics compared with state-of-the-art heuristic and learning-based schedulers. The field results further indicate that NeuroFly provides reliable guarantees in actual UAV operation scenarios when compared with mature solutions that have been widely deployed in production environments. These results validate the practicality, robustness, and engineering feasibility of NeuroFly, and suggest its potential for large-scale deployment in UAV applications that are sensitive to real-time video quality, including emergency response, power inspection, agricultural monitoring, and logistics delivery.
Box Particle Filter δ-GLMB Algorithm for Multiple Maneuvering Group Targets Tracking
GAN Linhai, WANG Gang, LI Zhihui, SUN Wen, WANG Baotang
2026, 48(7): 2909-2918. doi: 10.11999/JEIT251273
Abstract:
  Objective  Targets that move in a coordinated manner or show similar motion patterns are commonly referred to as group targets. Dense group targets contain many closely spaced individuals and often suffer from poor measurement resolvability, severe measurement overlap, and frequent target disappearance and reappearance. These factors make it difficult to establish stable tracks for individual targets within the group. Such groups are therefore usually treated as a whole to jointly estimate the kinematic state of the centroid and the extended shape. To improve tracking accuracy and computational efficiency for multiple maneuvering group targets under nonlinear measurements, an Interacting Multiple Model Gamma Box Particle δ-Generalized Labeled Multi-Bernoulli (IMM-GBP-δ-GLMB) algorithm is proposed. Tracking efficiency under nonlinear measurements is improved using the Box Particle Filter (BPF). The likelihood function of the GBP algorithm is improved, and the IMM algorithm is introduced to enhance tracking of the extended shape and centroid kinematic state of group targets. Finally, the method is integrated with the GLMB filter to track an unknown number of multiple maneuvering group targets.  Methods  Existing algorithms mainly describe the area-based overlap between the predicted extended state of group targets and the measurement distribution, but they do not fully capture shape similarity. To address this limitation, the likelihood function of the BPF is modified. Geometric parameters, including the semi-major axis, semi-minor axis, and inclination angle, are incorporated into the likelihood function. This improves the modeling of similarity between the predicted extended state and the measurement distribution. The modification is particularly useful for maneuvering group targets, because the inclination angle of the extended shape changes frequently during maneuvering. Based on IMM modeling of group motion, a model index is added to the centroid kinematic state of each box particle. The model index and centroid kinematic state are jointly estimated in each iteration, allowing mode transitions of individual box particles to be tracked and further improving tracking accuracy. The improved IMM-GBP filter is then embedded into the labeled random finite set framework, and the IMM-GBP-δ-GLMB algorithm is derived for effective tracking of multiple maneuvering group targets.  Results and Discussions  Simulation experiments are conducted to compare the proposed IMM-GBP-δ-GLMB algorithm with the IMM Sequential Monte Carlo δ-GLMB (IMM-SMC-δ-GLMB) filter. The proposed algorithm maintains comparable estimation accuracy for the centroid state, extended state, measurement rate, and number of targets, while improving computational efficiency. In the given simulation scenario, the proposed algorithm achieves a 3.8-fold improvement in timeliness, with an approximately 8.5% reduction in tracking accuracy. In scenarios with two and three group targets, the average tracking time growth rate of the proposed algorithm is 96% of that of the IMM-SMC-δ-GLMB filter. This result indicates good temporal robustness as the number of group targets increases. Therefore, the proposed algorithm has strong practical value.  Conclusions  This paper addresses the tracking of multiple maneuvering group targets under nonlinear measurement conditions by proposing the IMM-GBP-δ-GLMB algorithm. The main contributions are as follows: (1) The likelihood function of the BPF is improved to strengthen the measurement of similarity between the target extended shape and the measurement distribution, improving the tracking accuracy of group target states. (2) A motion model label is assigned to each box particle, and transitions in the target motion state are tracked during filtering. This allows the filter to achieve higher tracking accuracy with fewer box particles and improves computational efficiency. (3) The IMM-GBP method is integrated into the δ-GLMB framework to obtain the final IMM-GBP-δ-GLMB filter, which realizes effective tracking of multiple maneuvering group targets.
A High-Performance Eye Tracking Method Based on Event Camera and Dual-Channel Differential Illumination
SONG Sishun, FENG Junchi, PU Chengyu, GUO Yu, LIU Shijie, HE Xin, CHEN Yuwei
2026, 48(7): 2919-2929. doi: 10.11999/JEIT251162
Abstract:
  Objective  Eye tracking has become an essential technology in human-computer interaction, medical diagnostics, cognitive neuroscience, and augmented and virtual reality applications. Traditional eye tracking systems, however, often suffer from two major limitations: low spatial accuracy and limited temporal resolution, especially during high-speed eye movements. These limitations hinder precise gaze estimation and reduce the reliability of real-time interactive systems. To address these challenges, an event camera is integrated with a dual-channel differential illumination strategy to improve the signal-to-noise ratio of corneal reflection events. The Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is introduced to achieve accurate localization of corneal reflection points. On this basis, the coordinates of the corneal reflection points are combined with Singular Value Decomposition (SVD) and the least-squares method to determine the center of corneal curvature, thereby significantly improving gaze direction estimation accuracy. This study provides a practical technical route for next-generation eye tracking systems and offers theoretical support for their use in complex interactive environments.  Methods  An event-camera-based gaze tracking method is proposed that integrates asynchronous eye-movement event data through a dual-channel differential illumination framework, thereby improving gaze direction estimation accuracy under high-speed and dynamic conditions. First, the event camera asynchronously captures brightness-change events with microsecond-level temporal resolution, which enables precise tracking of rapid eye movements. At the same time, the dual-channel differential illumination mechanism suppresses redundant reflections and improves the contrast of corneal reflection points. Second, the DBSCAN algorithm is used to process the event data, effectively removing noise and improving the spatial localization accuracy of corneal reflection features. Finally, a ray-tracing model is reconstructed using SVD and least-squares fitting to determine the center of corneal curvature, thereby enabling robust and high-precision gaze direction estimation. Experimental results on a biomimetic eye-movement dataset show that the proposed method achieves high temporal resolution, localization accuracy, and robustness in dynamic tracking scenarios.  Results and Discussions  Experiments show that the proposed method achieves a temporal resolution of 25 kHz (Fig. 6), which far exceeds that of conventional cameras. Differential illumination significantly improves the signal-to-noise ratio of corneal reflection events. The DBSCAN algorithm localizes corneal reflection points more efficiently than K-means, agglomerative clustering, mean shift, and OPTICS, and achieves accurate results within 10 ms without requiring predefined clusters (Fig. 8, Table 3). For gaze estimation, the proposed method maintains stable accuracy across sampling frequencies from 2 kHz to 25 kHz. At a 15° cone angle, the Mean Error (ME) and Root Mean Square Error (RMSE) are approximately 0.66° and 0.67°, respectively. At 25°, these values increase slightly to 0.87° and 0.90°, respectively (Table 4). Compared with existing state-of-the-art gaze tracking methods, the proposed approach shows better overall performance in both temporal resolution and accuracy (Table 5). Trajectory results (Fig. 9) show close agreement between the estimated and ground-truth gaze paths, and distribution analyses (Fig. 10) confirm that most errors remain below 1°.  Conclusions  A novel eye tracking method is presented that integrates an event camera with dual-channel differential illumination. The method achieves high temporal resolution (25 kHz), improves event signal quality, and reduces localization errors, yielding gaze estimation errors of less than 1°. The proposed approach provides a reliable technical route for next-generation high-performance eye tracking systems. Future work should address sensor noise modeling and computational optimization to further improve real-world applicability.
Defending against Deepfakes by Attribute-Aware Attack
GAO Fan, YAN Weidan, SHAO Wenze, ZHANG Dengyin
2026, 48(7): 2930-2942. doi: 10.11999/JEIT260043
Abstract:
  Objective  Deepfakes can cause serious personal and property damage when misused. To prevent forged images from spreading, existing methods often use adversarial examples to protect facial images from deepfake manipulation. However, traditional gradient-based attacks show limited generalization and low generation efficiency in black-box attack scenarios. Their performance is also weaker than that of current methods based on Generative Adversarial Networks (GANs), which are used to train cross-model adversarial examples. Although GAN-based methods support fast inference, their lack of perceptual constraints often makes the generated adversarial perturbations visually noticeable. The rapid development of deepfake models also raises higher requirements for the generalization ability of adversarial examples. Therefore, imperceptible and generalizable adversarial attack methods are needed for proactive deepfake defense.  Methods  To further improve the transferability and imperceptibility of adversarial examples generated by existing methods, this paper proposes an attribute-aware adversarial example generation method for deepfake defense. The proposed method generates imperceptible perturbations and improves cross-model generalization through a frequency-domain identity fusion mechanism. Specifically, it focuses on the foreground regions of facial images, uses attribute-aware salient segmentation masks to separate facial and hairstyle regions, and combines these masks with adaptive spatial-frequency attention-based perturbation generators to generate region-specific adversarial perturbations. This strategy improves the imperceptibility of adversarial examples and reduces the additional computational cost caused by global processing. From the perspective of data augmentation, this paper further uses phase swapping in the frequency domain to fuse identity-related features from reference face images. This design reduces perturbation overfitting and improves generalization performance.  Results and Discussions  The proposed method is trained and tested on the CelebA-HQ dataset using proxy models. Compared with existing proactive defense methods, the experimental results show that the proposed method generates adversarial examples with strong imperceptibility and cross-model defense capability. It achieves a high defense success rate against various proxy models. The average Peak Signal-to-Noise Ratio (PSNR) of forged outputs under adversarial perturbations is reduced to 16.79 dB, representing an improvement of approximately 1.87% over the second-best method. Defense performance against HiSD is improved by approximately 7.5% compared with the second-best method. Defense performance against AttGAN is approximately 12.7% higher than that of the second-best GAN-based defense method. Moreover, the Learned Perceptual Image Patch Similarity (LPIPS) metric shows that the adversarial perturbations have high imperceptibility.  Conclusions  This study proposes a facial attribute-aware attack method for deepfake defense. The method incorporates a frequency-domain identity fusion mechanism to increase the diversity of adversarial feature inputs. Adaptive spatial-frequency attention-based perturbation generators are also designed to extract local facial information and dynamically adjust adversarial features. These designs allow the method to preserve perturbation components that are both imperceptible and attack-effective, leading to strong cross-model generalization. Future work will focus on proactive deepfake defense methods with improved imperceptibility and generalization, especially in cross-model transfer attack scenarios.
Physiological Signal-driven QoE Optimization for Wireless Virtual Reality Transmission
WU Chang, PENG Mingyu, CHEN Yuang, CHEN Yiyuan, GUO Fengqian, QIN Xiaowei, LU Hancheng
2026, 48(7): 2943-2954. doi: 10.11999/JEIT260067
Abstract:
  Objective  Virtual Reality (VR) has become a transformative medium for immersive digital experiences because it can deliver high-resolution 360° video with ultra-low Motion-To-Photon (MTP) latency. However, its dependence on wireless transmission creates major challenges. Uncompressed data rates above 1 Gbit/(s·Hz) and latency thresholds below 20 ms place stringent demands on network infrastructure. In mobile scenarios, channel fluctuation and user mobility often compromise service continuity and cause abrupt resolution changes. Traditional Quality of Service (QoS) metrics, such as bandwidth, jitter, and packet loss, provide useful network-level information but cannot adequately reflect subjective user satisfaction. Existing Quality of Experience (QoE) models and Adaptive BitRate (ABR) algorithms often use symmetric metrics, such as Mean Opinion Score (MOS), and overlook the fact that users perceive quality deterioration and quality improvement differently. Sudden resolution downgrading has a stronger negative effect on immersion than the positive effect caused by resolution upgrading. This perceptual asymmetry is consistent with behavioral psychology but remains insufficiently addressed in current transmission schemes. In addition, the separation between Radio Access Network (RAN) resource provisioning and application-layer bitrate adaptation often causes mismatched optimization, video-quality oscillation, and resource underuse. To address these issues, this study establishes a quantitative link between physiological responses and resolution changes. It further develops a physiological signal-driven QoE framework integrated with Deep Reinforcement Learning (DRL) to support adaptive transmission, maximize immersion, and reduce the adverse effects of resolution fluctuation in resource-constrained wireless networks.  Methods  A two-stage method is adopted, including physiological signal analysis and joint optimization framework design. A controlled VR experiment is conducted to quantify the perceptual effect of resolution changes. Nineteen healthy subjects participate in a viewing task using an eye-tracking VR headset, a 32-channel wireless ElectroEncephaloGraphy (EEG) system, ElectroCardioGraphy (ECG) recording, and Galvanic Skin Response (GSR) sensors. The subjects view natural-scene videos in which the resolution levels, including 8 k, 4 k, 1 080 P, 720 P, and 480 P, switch randomly every 8 s. The collected EEG signals are preprocessed by independent component analysis and band-pass filtering. Event-Related Potential (ERP) components are analyzed, with emphasis on the N200 component in the temporal and occipital regions, which reflects visual processing and attention allocation. A Linear Discriminant Analysis (LDA) classifier is used to distinguish different response types. The analysis focuses on the asymmetry between resolution upgrading and downgrading, and on sensitivity to the magnitude of resolution jumps. Based on these physiological findings, a QoE model is formulated by adding penalty terms for resolution degradation and large-amplitude resolution switching. These penalties are weighted more strongly than upgrade rewards to represent user aversion to quality drops. The model is then integrated into an edge-computing environment through a dual-timescale DRL framework. The framework separates control into two cooperative agents: the Scheduling and Utility (SU) agent and the Resolution Scaling (RS) agent. The SU agent operates at the millisecond timescale and performs real-time wireless resource allocation. It uses a Gated Recurrent Unit (GRU) to extract temporal features from Channel State Information (CSI) and transmission history. It then dynamically allocates bandwidth to improve frame delivery success and maintain fairness under VR frame-deadline constraints. The RS agent operates at the frame timescale and determines the resolution of subsequent video frames. Its decision-making is guided by the physiological signal-driven reward function, which penalizes actions that may trigger negative physiological responses, such as sharp resolution drops, unless channel deterioration makes them necessary. Proximal Policy Optimization (PPO) is selected for both agents because of its stable learning behavior in continuous and discrete action spaces. Simulations are conducted using a 3GPP-based wireless channel module with user mobility, shadow fading, and path loss to create a dynamic network environment.  Results and Discussions  The physiological experiment and network simulations validate the proposed framework. In the physiological analysis, a clear N200 response is observed approximately 200 ms after resolution changes. The N200 amplitude is significantly larger during resolution downgrading than during resolution upgrading (p < 0.001), indicating that users are more sensitive to quality deterioration. Large resolution jumps, such as changes from 8 k to 1 080 P, also induce stronger neural responses and more concentrated occipital energy than minor adjustments. The LDA classifier achieves an average Area Under the Curve (AUC) of 74.12% across 19 subjects, confirming that neural responses contain discriminative information about the direction of resolution change. The GSR results support these findings. A dual-branch GSR feature extraction and classification model reaches an average AUC of 78.10% in distinguishing upward and downward switching events. By contrast, ECG signals do not show a stable effect under the current experimental setting and analysis granularity. Therefore, the subsequent QoE model is mainly constructed from EEG and GSR findings. In the network performance evaluation, the proposed physiological signal-driven DRL framework is compared with several baselines, including Proportional-Fair (PF) scheduling, equal resource allocation, and traditional congestion control represented by SCReAM. The training curves show that the dual-agent system converges and learns to coordinate capacity provisioning with resolution decisions. The SU agent smooths short-term channel fluctuation and provides a stable capacity basis, which enables the RS agent to make more reliable resolution decisions. Quantitative results show that the proposed scheme improves the average video resolution by up to 88.7% compared with the equal-resource baseline. More critically, the resolution switching frequency is reduced by up to 81.0%. This reduction is essential because frequent switching, especially downward switching, causes user discomfort, as demonstrated by the physiological analysis. By prioritizing long-term resolution stability and penalizing abrupt drops through the physiological signal-driven reward function, the proposed system reduces the “ping-pong” effect commonly observed in traditional ABR algorithms. Compared with schemes using different penalty weights, the proposed method achieves a better balance. It avoids overly conservative behavior under large penalties, which lowers the average resolution, and unstable visual quality under small penalties, which increases resolution fluctuation. The joint optimization also allocates resources preferentially to users with urgent frame deadlines or higher risks of perceptible quality degradation, while maintaining a frame delivery success rate above 99%.  Conclusions  This paper addresses the conflict between wireless-channel instability and the human need for visually consistent VR streaming. By adopting a physiological signal-driven approach, the asymmetric effect of resolution changes on user experience is quantified, which challenges the symmetric assumptions used in traditional QoE models. Integrating this physiological evidence into a dual-timescale DRL framework enables the RAN to go beyond throughput-oriented optimization. Wireless resource allocation supports stable application-layer adaptation, while application-layer demands guide resource scheduling. The proposed solution improves immersive experience by increasing average resolution and reducing the physiologically disruptive effects of sudden quality degradation. The reduction in resolution switching frequency by more than 80% shows that the system can shield users from network variability. This study also indicates the value of edge intelligence in making resource-allocation decisions based on human perception rather than network statistics alone. Future work should extend the QoE model by considering multisensory factors, such as MTP latency, cybersickness, spatial distortion, stalling, and audiovisual synchronization. Individual differences in physiological sensitivity should also be addressed through personalized modeling. For real-world deployment, privacy protection is essential. Federated learning and local edge updates may allow biometric data to be processed locally while supporting global policy optimization. This work provides a human-centric basis for immersive networking and shifts the focus from QoS to physiologically validated QoE.
A Point Cloud Slice-based UAV SLAM Method for 3D Reconstruction of Large Container Port Areas
HU Zhaozheng, ZUO Zhihang, XU Cong, TAO Qianwen, LIU Chao, MENG Jie
2026, 48(7): 2955-2968. doi: 10.11999/JEIT251112
Abstract:
  Objective  With the continuous development of port intelligence, the demand for digital management in container port areas has increased. In large container yards, Three-Dimensional (3D) reconstruction of the yard environment can be achieved using Unmanned Aerial Vehicle (UAV)-based Simultaneous Localization And Mapping (SLAM). However, container port areas contain many repetitive semantic structures. Traditional semantic matching methods therefore show low efficiency and limited accuracy. In addition, lanes between container yards form large feature-sparse regions during UAV-based 3D reconstruction, which can cause odometry degradation. Repetitive scene features also interfere with loop closure detection. To address these problems, this paper proposes a rapid feature extraction method based on point cloud slicing and further optimizes it according to the structural characteristics of container yards. A UAV point cloud slice-based SLAM method, termed Slice-SLAM, is proposed for high-precision 3D reconstruction of large container port areas.  Methods  To improve point cloud semantic extraction, a rapid point cloud slicing method is proposed. The principal direction is extracted rapidly, and the point cloud is divided into multiple layers to obtain multi-layer semantic point clouds efficiently. The slicing strategy is further optimized for container yard scenarios. Principal plane extraction is simplified using the gravity direction, and the elevation range of each container layer is obtained adaptively from point cloud density gradient changes. Multi-layer slice point clouds are then constructed. A progressive adaptive Light Detection And Ranging (LiDAR) odometry method based on slice point clouds is developed. Elevation slices are used to identify degenerate scenarios adaptively, and a layer-wise incremental slice matching and fusion strategy is used. This improves the accuracy, efficiency, and stability of LiDAR odometry. In addition, a factor graph optimization method that integrates slice point cloud information is designed. Fusion voting is performed on the matching results of multi-layer slice point clouds to remove erroneous matches and reduce the effect of repetitive structures on loop closure detection. Slice factors are then used to construct factor graph edges, which improves global optimization and supports efficient and stable 3D reconstruction.  Results and Discussions  The feasibility and effectiveness of the proposed method are verified in CARLA simulation scenarios and real-world tests at a large container port in Wuhan. First, comparisons with three semantic extraction algorithms, namely RANSAC, Region Growth, and 3DG_SEG, demonstrate the efficiency and accuracy of the proposed semantic extraction method. Second, estimated trajectories are compared with those obtained by two open-source LiDAR algorithms, FAST-LIO2 and Faster-LIO, confirming the advantages of the proposed odometry method. Finally, speed and confidence score are compared with those of six algorithms: ICP, NDT, GICP, Fast-GICP, Scan Context+ICP, and Quatro. The loop closure detection module of LIO-SAM is also integrated into FAST-LIO2, and the Scan Context module is integrated into Faster-LIO. The resulting estimated trajectories are compared with those of the proposed method, verifying the effectiveness of the proposed loop closure detection algorithm. The proposed method achieves high 3D reconstruction accuracy and is suitable for practical port operations.  Conclusions  The proposed method uses an efficient point cloud slicing technique and a multi-layer slice matching mechanism. Points within the same elevation range are defined as a slice point cloud, and the segmentation process is defined as point cloud slicing. This design enables efficient and robust 3D reconstruction in large-scale scenes with repetitive features. First, the LiDAR point cloud is aligned with the positive Z-axis using the gravity direction derived from the Inertial Measurement Unit (IMU). A sliding window records density gradient changes to determine the elevation range of each layer adaptively. This simplifies point cloud slicing and reduces the effects of non-standard containers and ground height variations on semantic extraction. Multi-layer slice information is then integrated into the odometry module to detect degenerate scenarios. Under normal conditions, progressive slice matching is used to initialize pose estimation. In degenerate scenarios, iterative Kalman filtering with increased IMU weighting is used. Finally, the fusion voting mechanism removes outliers from multi-layer slice matching results. The optimal match is used to initialize loop closure for global registration of container-region point clouds, enabling dual-stage loop closure detection and slice factor construction. By integrating slice point cloud information into factor graph optimization, the proposed method unifies point clouds in a common coordinate system and achieves efficient and robust 3D reconstruction.
Drug Response Prediction Based on Graph Topology Attention Network
XU Peng, XU Hao, BAO Zhenshen, ZHOU Chi, LIU Wenbin
2026, 48(7): 2969-2978. doi: 10.11999/JEIT251099
Abstract:
  Objective  A central goal in modern cancer research is to determine why patients respond differently to the same therapy. This requires computational tools that combine genetic information with drug properties to predict treatment results, which is essential for advancing personalized oncology. Although existing methods have improved cancer drug response prediction, effective drug feature extraction and integration of multi-omics data from cell lines remain challenging. To address these issues, Graph Neural Networks (GNNs) have been increasingly used to process drug molecular graphs. In this study, a model based on a graph topology attention network is proposed to extract features from drug molecular graphs, and an attention mechanism is used to integrate multi-omics data.  Methods  In this study, a drug response prediction method based on Graph Topology Attention Network (GTAT) is proposed. The model integrates topological graph information to predict drug responses in cell lines. Drug SMILES strings are used to generate two different drug representations, and multi-omics data are incorporated to characterize cell lines (Fig. 1). For drug feature extraction, SMILES strings are first parsed to construct molecular graphs, which are then processed by GTAT. This network captures both topological information at the molecular graph level and atom-level features, thereby generating structured molecular representations. At the same time, Extended Connectivity Fingerprints are computed from the same SMILES strings and transformed into continuous feature vectors through a Multi-Layer Perceptron (MLP). The graph-based drug representation and the fingerprint-based representation are then concatenated to form a comprehensive drug feature vector. For cell line representation, multi-omics data are processed through omics-specific neural networks. The resulting features are fused through multi-head self-attention mechanisms, which enable the model to capture contextual interactions across omics modalities and generate an integrated cell line representation. Finally, the drug and cell line features are combined and fed into an MLP classifier to predict drug response results. The proposed model effectively integrates heterogeneous biological data sources and significantly improves prediction accuracy through multimodal learning and attention-based feature fusion.  Results and Discussions  The proposed method achieves competitive performance on both the GDSC and CCLE benchmark datasets (Table 2). Specifically, on the GDSC dataset, the proposed approach outperforms all competing methods across all four metrics, including AUC, AUPR, F1-score, and Accuracy. In particular, the AUPR is improved by approximately 1.92% compared with that of the second-best method, MOFGCN, which demonstrates an advantage in handling class imbalance. On the CCLE dataset, the proposed method still achieves the best performance in terms of AUC and Accuracy. Although its AUPR and F1-score are slightly lower than those of GADRP, the differences are minimal, and the method shows stronger overall discriminative ability, as reflected by AUC. These results validate the effectiveness and strong generalizability of the proposed method in drug sensitivity prediction tasks. The variation in AUPR and F1-score across datasets may be attributed to inherent differences in sample size and class distribution. The limited size of the CCLE dataset, combined with its specific class imbalance, with an approximately 4:1 ratio of resistant to sensitive samples, may restrict the model’s ability to fully learn the underlying data distribution, particularly for minority classes. In contrast, the GDSC dataset shows greater heterogeneity and a more pronounced class imbalance, approximately 8:1, which increases prediction difficulty and leads to lower performance on certain metrics.  Conclusions  Accurate prediction of drug response in cell lines remains a central challenge in precision medicine and has important implications for accelerating drug development and advancing personalized treatment. However, construction of a highly accurate predictive model that effectively integrates multi-source biological information remains difficult because of the complexity of drug molecular structures and the inherent heterogeneity of cell lines. To address this issue, a cell line drug response prediction model based on GTAT is proposed. In this model, GTAT is used to extract molecular graph features of drugs, which are then fused with molecular fingerprint features. Meanwhile, multi-omics features of cell lines are integrated through an attention mechanism. Experimental results demonstrate that the proposed model achieves superior performance compared with existing state-of-the-art benchmark methods on the employed datasets. This study provides a new perspective for cell line drug response prediction. Certain limitations remain, including the use of only three types of omics features for cell line representation and the effect of sample size on predictive performance. Future work will focus on integrating more diverse omics features, applying pre-trained large-scale models, and promoting clinical translation for personalized medicine.
Household Appliance Plastics Identification by Fusing Multi-Level Feature Enhancement and Hierarchical Classification
CHONG Penghao, ZHENG Yunlong, YANG Aosong, GUO Mengci, LI Shifeng
2026, 48(7): 2979-2989. doi: 10.11999/JEIT260084
Abstract:
  Objective  Accurate plastic identification remains challenging in waste household appliance recycling under low-resolution spectral conditions. In practical recycling environments, plastics often have complex compositions, surface contamination, and aging effects, which increase classification difficulty. Black plastics are especially difficult to identify because their strong light absorption and spectral overlap in the Visible-Near Infrared (Vis-NIR) range reduce feature separability and degrade classification performance. Under these conditions, conventional single-stage classification models often fail to maintain stable accuracy. To address this problem, an automated identification method is proposed for low-dimensional multispectral feature spaces. The method aims to improve the discriminative capability of limited spectral information and enhance classification accuracy for complex plastic categories.  Methods  A compact Vis-NIR multispectral acquisition system based on the AS7265x sensor is used to collect 18-channel reflectance data in the 410~940 nm range. A handheld acquisition device with a controlled optical structure is designed to reduce environmental interference and ensure measurement consistency (Fig. 3). A total of 576 samples are collected from five typical household appliance plastics, including Acrylonitrile Butadiene Styrene (ABS), High-Impact PolyStyrene (HIPS), PolyPropylene (PP), Acrylonitrile Styrene copolymer (AS), and Polycarbonate/Acrylonitrile Butadiene Styrene (PC+ABS) blends. These samples are obtained from waste household appliances and are subjected to preliminary surface cleaning before spectral acquisition. To improve feature representation, a multi-level feature engineering strategy is adopted. This strategy integrates original spectral intensity features, nonlinear polynomial expansion features, and adjacent-channel ratio features to characterize both global and local spectral information. The nonlinear expansion enhances the representation of reflectance variations, whereas the ratio features capture local spectral-shape changes and reduce external disturbances. These features are combined into a 53-dimensional feature vector. Linear Discriminant Analysis (LDA) is then applied to enhance interclass separability. To address spectral overlap and class imbalance, a Hierarchical Joint Classifier (HJC) is constructed. HJC uses a two-stage classification framework. In the first stage, an XGBoost-based primary classifier performs coarse classification to separate easily distinguishable samples and group spectrally similar black plastics. In the second stage, a TabTransformer-based secondary classifier performs fine-grained classification of difficult samples (Fig. 6). This hierarchical design reduces classification complexity and improves discrimination for challenging categories. Model performance is evaluated using five-fold cross-validation and an independent test set. Accuracy, precision, recall, and F1-score are calculated from confusion matrices (Fig. 7). Comparative experiments are conducted with traditional machine learning methods, ensemble learning models, and deep learning approaches under different feature-processing strategies (Fig. 8, Fig. 9).  Results and Discussions  The proposed HJC achieves a classification accuracy of 97.4% in five-fold cross-validation and 93.1% on the independent test set (Table 4). Compared with single-stage classifiers and methods without feature enhancement, the proposed method provides higher performance and greater stability under low-resolution spectral conditions. Comparative results show that the proposed method outperforms baseline approaches, such as PCA combined with CNN, which achieves an accuracy of approximately 71.3% on the same dataset (Fig. 8). This improvement indicates that the proposed feature engineering strategy effectively strengthens the discriminative capability of low-dimensional spectral data. Combining LDA with feature engineering further improves class separability compared with conventional PCA-based methods. Confusion matrix analysis shows that misclassifications mainly occur between spectrally similar black ABS and black HIPS samples, whereas most other categories achieve high classification accuracy (Fig. 9). These results indicate that spectral overlap remains the main challenge under low-resolution conditions. The hierarchical classification strategy reduces this problem by focusing classification resources on difficult samples, thereby improving the overall generalization ability of the model. Overall, the proposed method shows robustness under practical conditions, including spectral noise, limited channel resolution, and material heterogeneity. These results indicate its suitability for real-world recycling applications.  Conclusions  A hierarchical classification method with multi-level spectral feature engineering is developed for plastic identification under low-resolution Vis-NIR conditions. Nonlinear and spectral-shape features are incorporated into a two-stage framework to improve the identification of spectrally similar materials. The results show stable accuracy across different plastic types. The method is suitable for automated sorting in waste household appliance recycling and can be extended to other material identification tasks with limited spectral information.
Aerial Spatio-Temporal Image Generation via Latent Diffusion Models
SHANG Yuying, HOU Yingyan, LIU Zinan, LU Wanxuan, HUANG Yuhong, WANG Yixiao, YU Hongfeng, FU Kun
2026, 48(7): 2990-3001. doi: 10.11999/JEIT260165
Abstract:
  Objective  Aerial Earth observation plays a pivotal role in environmental monitoring, disaster warning, and urban planning. However, constraints such as flight-platform endurance and mission-window timeliness often prevent acquired aerial imagery from fully characterizing the long-term evolution of the Earth’s surface. Although pre-trained latent diffusion models have shown strong potential for image generation, their application in aerial scenarios remains challenging because of the scarcity of high-quality temporal annotation data and semantic-visual misalignment caused by variable observation scales. To address these challenges, this paper proposes ASTIG, a training-free framework for Aerial Spatio-Temporal Image Generation. By leveraging the generative priors of pre-trained latent diffusion models and Large Language Models (LLMs), ASTIG provides a new paradigm for semantically controllable aerial spatio-temporal image generation.  Methods  ASTIG consists of three coordinated components. First, a dynamic semantic decomposition process is proposed to parse complex descriptions of aerial scene evolution into frame-level visual prompts, thereby compensating for the lack of temporal semantic annotations in existing aerial image-text datasets. Second, a Linguistic Binding (LB) strategy is proposed to establish explicit associations between key ground objects and their corresponding visual attributes within the cross-attention mechanism of the diffusion model, thereby improving the semantic response precision of the generated images. Third, a Temporal Anchor Attention (TAA) mechanism is incorporated. It uses dual reference frames to maintain subject stability and background consistency across the generated spatio-temporal image sequence, thus suppressing inter-frame temporal drift under training-free conditions.  Results and Discussions  ASTIG and the baseline methods are evaluated on 7 236 high-quality aerial spatio-temporal descriptions using six automated metrics, including subject consistency, background consistency, temporal flickering, motion smoothness, aesthetic quality, and imaging quality. Quantitative results (Tables 1 and 2) show that ASTIG outperforms the baseline methods in spatio-temporal image generation, with improvements of 3.91% in subject consistency and 4.57% in temporal flickering over the frame-prompt baseline. Qualitative comparisons (Fig. 4) further show its strong ability to model long-term surface evolution in aerial imagery. Ablation studies validate the individual effectiveness of the LB strategy and the TAA mechanism (Table 3 and Fig. 5). Sensitivity analyses of the intervention steps (Table 4 and Fig. 6) and binding strength (Table 5 and Fig. 7) further identify suitable parameter settings. Extension experiments from satellite perspectives (Figs. 8 and 9) also show that ASTIG has the potential to generalize beyond aerial platforms to broader Earth observation scenarios.  Conclusions  This paper proposes ASTIG, a training-free framework for aerial spatio-temporal image generation that addresses the scarcity of high-quality long-term temporal data and semantic-visual misalignment. By leveraging the generative priors of pre-trained latent diffusion models and LLMs, ASTIG integrates a dynamic semantic decomposition process, an LB strategy, and a TAA mechanism to improve temporal semantic construction, semantic response precision, and inter-frame consistency. Experimental results show that ASTIG outperforms existing baseline methods across multiple automated evaluation metrics, providing a new paradigm for aerial spatio-temporal image generation. As a training-free method, ASTIG is still limited by the prior knowledge of the backbone model. Future work will examine geometric correction and nadir-view prior constraints to better align the generated results with the physical properties of satellite imagery.
KE-HNS: Knowledge-Enhanced Personalized Recommendation Model with Hierarchical Noise Suppression
XIE Jun, WANG Dantong, ZHANG Bo, CHEN Guijun, LÜ Jiaqi, LUO Xiongyan
2026, 48(7): 3002-3014. doi: 10.11999/JEIT260051
Abstract:
  Objective  In the era of Big Data and Artificial Intelligence (AI), rapid information growth has increased the difficulty of filtering valuable content from redundant data. Personalized recommender systems are key tools for accurate information matching and resource allocation. Knowledge Graphs (KGs) can enrich user-item representations. However, current KG-based recommendation models still face weak noise suppression, coarse-grained user-interest modeling, and imbalanced use of heterogeneous information, which reduce recommendation accuracy. This paper proposes Knowledge-Enhancedpersonalized recommender model with HierarchicalNoise Suppression (KE-HNS), which integrates knowledge enhancement with hierarchical noise suppression. By combining graph representation learning and contrastive learning, KE-HNS addresses noise interference, fine-grained preference modeling, and multi-source information balance, thereby improving recommendation performance.  Methods  KE-HNS adopts a hierarchical noise-suppression paradigm. At the input stage, Input Noise Reduction (INR) is used to reduce noise from two sources. For user-item interactions, a learnable binary mask matrix is used to remove noisy edges. For KG denoising enhancement, triples are scored by importance, low-score triples are identified with a Bottom-K strategy, and noisy triples are masked. At the feature-fusion stage, Isolated Noise Suppression (INS) is used to preserve spatial independence by partitioning entity-attribute spaces according to relation type. This design limits high-order noise propagation and semantic contamination. At the representation-optimization stage, Comparative Noise Suppression (CNS) is implemented through contrastive learning to suppress irrelevant entity noise and strengthen robust semantic signals. To capture fine-grained user interests, Graph Convolutional Networks (GCNs) are used to enhance user representations from historical interactions and related entities. Adaptive weight layers further refine item representations by using entity attributes and relations. To balance heterogeneous information, a dual-view contrastive learning mechanism is constructed between the user-item view and the item-entity view. Positive and negative sample pairs are used to adaptively adjust the weights of different information sources. Finally, user and item representations are matched by inner product to generate the Top-K recommendation list.  Results and Discussions  KE-HNS is evaluated on three public datasets, Book-Crossing, MovieLens-1M, and Last.FM, through performance comparison, ablation experiments, denoising evaluation, case analysis, and complexity assessment. For Click-Through Rate (CTR) prediction, KE-HNS outperforms the best baseline models by 0.94%~1.01% in Area Under the Curve (AUC) and 0.43%~0.90% in F1-score (Table 3). For Top-K recommendation, its Recall@K is higher than those of most advanced methods across nearly all K values, with only a slight gap behind CG-KGR on Last.FM (Fig. 7). The ablation results show that all three denoising components contribute to the performance gains (Table 4). The denoising evaluation shows that KE-HNS effectively suppresses noise and maintains high prediction accuracy under noisy conditions (Fig. 8). The complexity analysis further indicates that the model remains feasible for practical deployment (Table 5).  Conclusions  This paper presents KE-HNS, a personalized recommendation model that combines knowledge enhancement with hierarchical noise suppression. By reducing noise interference and balancing collaborative filtering signals with knowledge-aware semantics, KE-HNS improves recommendation accuracy across multiple benchmark datasets. The model still has limitations in computational efficiency and depends on the coverage and completeness of the KG. Future work may focus on computational optimization and dynamic knowledge integration.
Data-driven Sliding-mode Disturbance-rejection Formation Control for Quadrotor UAV Swarms Under Uncertain Disturbances
LI Qianxiong, LU Xiaoqing
2026, 48(7): 3015-3026. doi: 10.11999/JEIT260050
Abstract:
  Objective  Quadrotor Unmanned Aerial Vehicle (UAV) cooperative formation can increase payload capacity and extend the operational range. However, quadrotor UAVs are highly nonlinear and underactuated systems. Differences in size and actuator hardware further weaken the effectiveness of model-based formation-control methods. Therefore, disturbance-rejection formation control is needed for quadrotor UAV swarms with unknown internal models and uncertain external disturbances.  Methods  To address the difficulty of precise modeling for quadrotor UAV swarm formation under uncertain disturbances, this paper proposes a data-driven sliding-mode disturbance-rejection formation control method. First, a data-driven formation-control model is established using the input and output states of each UAV and its neighboring UAVs. Then, an extended state observer and an integral sliding-mode formation controller are designed to estimate uncertain disturbances online and achieve robust formation control. Finally, stability analysis is conducted to derive sufficient conditions under which all UAVs achieve sliding-mode disturbance-rejection formation. The proposed method is verified through simulations and experiments under an unknown system model and uncertain disturbances.  Results and Discussions  The simulation results show that multiple quadrotor UAVs can maintain the desired formation geometry in a wind-disturbed environment (Fig. 4). The formation position error converges to within 0.1 m in 15 s and reconverges rapidly after a 7 m/s gust is applied (Fig. 6). The velocity curves also show rapid convergence among the UAVs (Fig. 5). The experimental results indicate that three UAVs can follow the trajectory of the virtual leader while maintaining the desired triangular formation (Fig. 17). The formation error is mostly kept within 0.1 m (Fig. 18). When the observation matrix fluctuates strongly between 10 s and 20 s, the corresponding formation error is relatively large. When the observation matrix curve becomes smoother between 20 s and 30 s, the formation error also decreases (Fig. 20). Compared with traditional model-based formation-control methods and existing data-driven methods, the proposed method reduces the formation error by 41% and shortens the formation response time by 40%.  Conclusions  This paper proposes a data-driven sliding-mode disturbance-rejection formation control method for quadrotor UAV swarms with unknown internal models and uncertain external disturbances. Under an unknown quadrotor UAV model and a 7 m/s wind disturbance, the proposed method keeps the formation error below 0.1 m. It also reduces the formation error by 41% and shortens the formation response time by 40% compared with traditional model-based formation-control methods and existing data-driven methods. Future work will study multilayer data-driven formation control for heterogeneous UAV-UGV swarm systems. It will also optimize computational cost and scalability in large-scale and complex application scenarios.
Wireless Communication and Internet of Things
One-step Reconstruction Diffusion Model-based Poisoning Attack on QoS-aware Cloud API Recommender Systems
TAN Zeyu, WANG Haoyuan, QI Mingyang, SUN Mengmeng, SHEN Limin, CHEN Zhen
2026, 48(7): 3027-3036. doi: 10.11999/JEIT260115
Abstract:
  Objective  In cloud computing, Cloud Application Programming Interfaces (cloud APIs) serve as key carriers for data output, capability reuse, and service delivery. They have become core elements in service-oriented software development and operation. With the rapid growth of cloud APIs, users often find it difficult to select suitable services from many functionally similar candidates. Quality of Service (QoS) is therefore used to differentiate cloud APIs by non-functional attributes. QoS-Aware cloud API Recommender System (QARS) plays an increasingly important role in guiding users toward suitable cloud APIs. However, existing studies mainly focus on improving recommendation accuracy and often ignore security risks caused by the economic value of cloud APIs and the openness of network environments. These risks are particularly evident in poisoning attacks. By injecting fake users, attackers can manipulate recommendation results and reduce the fairness and credibility of QARS. To address this threat from an attack-informed defense perspective, this paper analyzes the attack mechanisms of diffusion model-based poisoning methods and supports the design of targeted defense strategies.  Methods  The poisoning attack process and fake user profiles are first formally defined. Attack scale is then defined to flexibly simulate poisoning attacks under different settings. To analyze the attack principle of diffusion model-based methods, a One-step reconstruction Diffusion Model (ODM) is adopted, and a Preference guided one-step reconstruction Diffusion model-based Poisoning Attack framework (PDPA) is proposed. According to the collaborative principle that similar users tend to have similar preferences for cloud APIs, fake users generated by an attack method should have QoS values and cloud API invocation distributions similar to those of real users. This similarity allows fake users to exert collaborative influence and interfere with user preference modeling in QARS. PDPA is therefore designed to generate fake users that closely match real users. First, ODM separately models the QoS data and invocation distributions of real users. Unlike standard diffusion models, ODM avoids error accumulation caused by noise-dependent iterative denoising. It can generate fake-user invocation behavior similar to real-user behavior, which helps fake users exert effective collaborative influence. Then, to improve attack effectiveness, PDPA systematically selects fake users with invocation preferences for the target cloud API and assigns the maximum QoS value to the target item. This strategy strengthens the attack while reducing the disturbance caused by adding the target cloud API to fake-user invocation behavior, thereby improving stealthiness.  Results and Discussions  Experiments are conducted on the real-world WS-DREAM response-time QoS dataset. First, six recommendation methods, namely LR, MLP, DeepFM, AFM, DCN, and XSimGCL, are used as target recommender systems. Six baseline attack methods are used to simulate poisoning attacks. The results in Table 3 reveal the vulnerability of QARS to poisoning attacks. All attack methods reduce recommendation accuracy. PDPA achieves the best attack effectiveness in most experimental settings because it sufficiently models user invocation preferences, enabling fake users to exert stronger collaborative influence on QARS. Second, fake users generated by ODM and those generated by the standard diffusion model are compared in terms of F1 score and latent-space distribution. The results in Figure 2 show that ODM outperforms the standard diffusion model in stealthiness and produces a latent-space distribution closer to that of real users. Third, ablation studies are conducted for each module of PDPA. The results in Tables 4 and 5 verify that each module is necessary for attack effectiveness and fake-user stealthiness. Finally, Mean Absolute Error (MAE) and F1 score are compared under different attack scales to evaluate the effect of attack scale on attack effectiveness and stealthiness. The results in Figure 3 and Table 6 show that increasing the attack scale improves attack effectiveness but also increases the number of detected fake users.  Conclusions  This paper investigates the threat of poisoning attacks against QARS by analyzing the attack process and key attack parameters. The proposed PDPA simulates poisoning attacks on QARS and reveals their vulnerability. The results show the potential of diffusion models for poisoning attacks and verify the necessity of separately modeling QoS data and cloud API invocations. PDPA also clarifies how diffusion models generate fake users, providing a basis for future targeted countermeasures.
Rotatable-Antenna-Aided Near-Field Wideband Integrated Sensing and Communication System: Hybrid Beamforming Design
XU Hongbo, MO Minghui, XIN Wei, WANG Shuli, WANG Ji, LI Xingwang, ZHENG Le
2026, 48(7): 3037-3046. doi: 10.11999/JEIT260023
Abstract:
  Objective  Near-field wideband Integrated Sensing and Communication (ISAC) systems face two main challenges: pronounced near-field effects and wideband beam splitting. These effects reduce communication throughput and sensing reliability, particularly when fixed-orientation antenna arrays and phase-shifter-based beamforming architectures are used. Because such architectures provide limited spatial adaptability and frequency-independent phase control, the spatial-frequency degrees of freedom available in near-field wideband channels cannot be fully used. To address this issue, a Rotatable-Antenna-assisted near-field wideband ISAC architecture is investigated to improve the system sum rate under sensing constraints.  Methods  A near-field wideband ISAC architecture assisted by Rotatable Antennas (RAs) is proposed. By allowing the antenna boresight direction to be adjusted mechanically or electronically, additional angular degrees of freedom are provided at the element level, which enables more flexible spatial coverage and more accurate energy focusing. A True Time Delay (TTD)-based hybrid beamforming architecture is further adopted to provide frequency-dependent phase shifts and compensate for the frequency-independent property of conventional phase shifters. Consistent beam focusing across subcarriers is thus maintained, and wideband beam splitting is effectively suppressed. Based on a spherical-wave near-field channel model that incorporates propagation distance, angular information, and the orientation gain of RAs, a joint optimization problem is formulated to maximize the system sum rate under transmit power constraints, sensing power thresholds, and antenna rotation constraints. Because the resulting problem is highly non-convex, a Penalty-Based Fully Digital Approximation (PBFDA) algorithm is developed. In each iteration, the RA orientations are first optimized by Particle Swarm Optimization (PSO) to improve the weighted channel gain. Then, with the antenna orientations fixed, a reduced-dimensional formulation with Successive Convex Approximation (SCA) is used to solve the fully digital beamforming problem. Finally, a manifold-based Block Coordinate Descent (BCD) algorithm is used to jointly optimize the analog beamformer, digital beamformer, and TTD units, so that the hybrid beamforming solution gradually approaches the fully digital solution (Algorithm 1–Algorithm 4).  Results and Discussions  Simulation results verify the effectiveness of the proposed RA-assisted near-field wideband ISAC framework. The proposed PBFDA algorithm converges monotonically within a limited number of iterations, which confirms its numerical stability and efficiency (Fig. 2). Compared with fixed-antenna architectures, the proposed RA-assisted scheme achieves a clear improvement in system sum rate under the same transmit power constraint (Fig. 3). When the system bandwidth increases, the spectral efficiency of TTD-based hybrid beamforming decreases because the limited number of TTD units and the restricted maximum delay weaken frequency-dependent compensation and aggravate beam splitting. By contrast, the optimal fully digital beamforming scheme maintains nearly unchanged spectral efficiency because each subcarrier can be controlled accurately (Fig. 4). When the sensing power threshold increases, the achievable sum rate decreases for all schemes, which reflects the trade-off between communication and sensing. The proposed method, however, consistently outperforms the benchmark schemes (Fig. 5). The effects of antenna number, antenna directivity factor, and maximum rotation angle are also evaluated. Spectral efficiency increases with the number of antennas because of the higher array gain (Fig. 6). As the antenna directivity factor increases, the RA-assisted system attains further gains through adaptive orientation, whereas fixed-orientation and isotropic schemes degrade (Fig. 7). A larger allowable rotation range also provides greater spatial alignment flexibility and further improves system performance (Fig. 8). Overall, the proposed architecture improves near-field energy focusing and achieves performance close to that of fully digital beamforming with lower hardware complexity.  Conclusions  A Rotatable-Antenna-assisted near-field wideband ISAC system with a TTD-based fully connected hybrid beamforming architecture is investigated. By jointly using antenna rotation and true time delay, the proposed framework effectively mitigates near-field effects and wideband beam splitting. The developed PBFDA algorithm solves the resulting highly non-convex optimization problem efficiently. Numerical results show that the proposed scheme significantly improves the system sum rate under sensing constraints and approaches the performance of fully digital beamforming, which supports its use in near-field wideband ISAC systems.
Secure and Covert MIMO Short packet Communication with Location-Uncertain Malicious Nodes
TIAN Bo, YANG Weiwei, YANG Xiaoqin, BAI Mengmeng
2026, 48(7): 3047-3058. doi: 10.11999/JEIT260059
Abstract:
  Objective  This paper investigates secure and covert short-packet communication in Multiple-Input Multiple-Output (MIMO) wireless systems with location-uncertain malicious nodes over quasi-static Rician fading channels. In the considered scenario, a legitimate transmitter sends confidential short packets to a legitimate receiver. Meanwhile, multiple monitoring nodes (Willie nodes) attempt to detect whether transmission occurs, and multiple eavesdropping nodes (Eve nodes) attempt to intercept the confidential information. Because malicious nodes may remain silent and their exact locations are unavailable to the legitimate system, their spatial uncertainty poses major challenges to joint covertness and secrecy analysis. To address this problem, a unified analytical and optimization framework is established for secure and covert short-packet transmission. The framework is used to characterize the coupling among covertness, secrecy, and reliability and to improve the Average Effective Secrecy and Covert Rate (AESCR).  Methods  The transmitter adopts Singular Value Decomposition (SVD)-based precoding, and the legitimate receiver applies Maximum Ratio Combining (MRC) to enhance the legitimate link. Monitoring nodes and eavesdropping nodes are modeled as two independent Poisson Point Processes (PPPs) outside a circular protection zone centered at the transmitter. This model captures the spatial randomness of malicious nodes. For covertness analysis, each monitoring node is assumed to perform optimal Likelihood Ratio Test (LRT)-based detection with full knowledge of the system model, noise power, channel state, and codebook information. Using the Chernoff bound and the Bhattacharyya coefficient, a theoretical lower bound on the minimum detection error probability of a single monitoring node is first derived. Stochastic geometry is then combined with the distribution of the strongest monitoring node to obtain a tractable lower bound on the average minimum detection error probability. For secrecy analysis, the finite blocklength normal approximation is used to account for decoding error and information leakage penalties. The legitimate channel is statistically characterized under Rician fading conditions, and the strongest eavesdropping node is analyzed through stochastic geometry. Based on these results, an approximate analytical expression for the average secrecy rate is derived. AESCR is proposed as a comprehensive performance metric that jointly reflects reliability, secrecy, and covertness. Under the average covertness constraint and the short-packet length constraint, a joint optimization problem for transmit power and packet length is formulated. By using the monotonic properties of the objective function and the covertness constraint, the original coupled optimization problem is transformed into a one-dimensional search problem.  Results and Discussions  Simulation results verify the accuracy of the theoretical derivations and reveal the effects of key system parameters. Both the simulated average minimum detection error probability and its theoretical lower bound decrease as the packet length increases. Higher transmit power further reduces the detection error probability, indicating that excessive power makes transmission more exposed to monitoring nodes (Fig. 2). Increasing the number of monitoring-node antennas strengthens spatial reception capability and further degrades covertness (Fig. 2). Enlarging the protection zone improves covertness because malicious nodes are forced to remain farther away from the transmitter. However, increasing the monitoring-node density weakens this benefit by raising the probability that a strong monitoring node appears near the protection-zone boundary (Fig. 3). The average secrecy rate increases with packet length and gradually approaches the asymptotic secrecy-capacity upper bound because the finite blocklength rate penalty decreases as the packet length grows (Fig. 4). AESCR first increases and then decreases with packet length, confirming the existence of an optimal packet length. This behavior results from the tradeoff between the reduced finite blocklength penalty and increased detection exposure (Fig. 5). Higher malicious-node density and more malicious-node antennas degrade system performance because they enhance both monitoring and eavesdropping capabilities (Fig. 5). Relaxing the covertness constraint improves the achievable AESCR because the system can select a higher transmit power or a more favorable packet length (Fig. 6). Results under different Rician factors show that the proposed analytical framework is applicable to both Rician and Rayleigh fading conditions (Fig. 6). Increasing the number of legitimate receive antennas improves AESCR, and a larger transmit antenna array provides additional SVD precoding gain (Fig. 7). Compared with benchmark schemes, the proposed joint optimization of transmit power and packet length consistently outperforms the scheme with fixed packet length and power-only optimization. This result demonstrates the need to jointly balance reliability, secrecy, and covertness in MIMO short-packet transmission (Fig. 8).  Conclusions  This paper develops a stochastic-geometry-based analytical framework for secure and covert MIMO short-packet communication with location-uncertain multi-antenna malicious nodes. By deriving a lower bound on the average minimum detection error probability, obtaining an approximate analytical expression for the average secrecy rate, and proposing AESCR, the framework reveals the fundamental tradeoff among covertness, secrecy, and reliability under finite blocklength transmission. The results show that increasing the number of legitimate transmit and receive antennas improves secure and covert performance, whereas higher malicious-node density and more malicious-node antennas degrade system performance. The existence of an optimal packet length further shows that packet length and transmit power should be jointly designed. The proposed joint optimization method therefore provides an effective solution for secure and covert short-packet transmission in mission-critical and low-latency wireless systems.
Modulation Recognition Method for High-Speed Mobile Communication Based on Attention Dynamic Fusion and Hybrid Pruning Transformer
ZHENG Qinghe, CHEN Bin, YU Lisu, HUANG Chongwen, JIANG Weiwei, SHU Feng, ZHAO Yizhe
2026, 48(7): 3059-3070. doi: 10.11999/JEIT251211
Abstract:
  Objective  Automatic modulation recognition is a critical preprocessing step in dynamic spectrum access and anti-jamming communication systems. It directly affects the robustness and spectrum efficiency of noncooperative communication. In high-speed mobile communication scenarios, such as satellite communication, high-speed rail communication, and drone swarm communication, signal modulation features experience severe distortion due to Doppler shifts, time-varying channels, and non-stationary interference. These factors challenge traditional modulation recognition methods that rely on static assumptions, leading to feature mismatch and higher misclassification rates. To address the limited robustness and real-time performance of existing deep learning–based modulation recognition models in high-speed mobile environments, a lightweight dynamic fusion Transformer-based method is proposed.  Methods  The proposed method contains three components: a signal representation fusion block, a Transformer architecture, and a model pruning strategy for lightweight inference. First, a RollingQ mechanism dynamically adjusts the direction of the attention query matrix according to the quality of each signal representation. This design prevents attention fixation and enables balanced use of multiple signal representations. Next, a Multi-head Attention Frequency Enhancement Transformer (MAFE-Transformer) is designed. The model integrates local and global spatiotemporal features through lightweight convolutional enhancement, multi-attention feature extraction, and frequency learning and selection modules. Finally, an attention-based dynamic hybrid pruning strategy removes structural redundancy and accelerates inference, which supports real-time modulation recognition.  Results and Discussions  Experiments are conducted on two public datasets, RadioML 2016.10a and RML22, to evaluate the proposed method. The MAFE-Transformer achieves average classification accuracies of 65.34% and 73.42% on the two datasets. Under low Signal-to-Noise Ratio (SNR) conditions of –20~0 dB, the model maintains strong robustness, particularly on the RML22 dataset with the dynamic channel model ETU70 (Fig. 6). The confusion matrix indicates that classification errors are relatively evenly distributed across different modulation schemes, which reflects balanced classification performance (Fig. 7). Ablation experiments show that the RollingQ-based dynamic fusion mechanism improves accuracy by 4.95% on RadioML 2016.10a and 3.86% on RML22 compared with a single signal representation (Fig. 8). The hybrid pruning strategy reduces inference latency to 0.011 ms per signal while maintaining high accuracy (Fig. 9). Comparative experiments indicate that the proposed model outperforms several advanced deep learning models, including Ms-RaT, MobileViT, MobileRaT, and KA-CNN, by 4%~10% in recognition accuracy. These results indicate strong performance in high-speed mobile communication scenarios (Fig. 10).  Conclusions  A lightweight dynamic fusion Transformer-based automatic modulation recognition method is proposed for high-speed mobile communication environments. The RollingQ mechanism and the MAFE-Transformer architecture, combined with a dynamic hybrid pruning strategy, improve the balance between recognition accuracy and inference efficiency. Experimental results on public datasets confirm the effectiveness and robustness of the method under complex channel conditions with Doppler shifts and time-varying interference. However, the method has not been systematically evaluated under more complex interference conditions, such as impulsive noise or frequency-selective fading. Future work will examine adaptability to non-stationary noise, cross-device generalization, and optimization for edge deployment.
Research on UAV-assisted Dynamic-weight Edge Computing Offloading Strategy
WANG Yijun, WANG Yachu, SHAHD Batool, MIAO Ruixin
2026, 48(7): 3071-3083. doi: 10.11999/JEIT260054
Abstract:
  Objective  The increasing demands of the Internet of Things (IoT) for computational resources and real-time processing have highlighted the significance of Mobile Edge Computing (MEC). Traditional MEC relies on terrestrial base stations, resulting in coverage blind spots in remote or specialized environments. Unmanned Aerial Vehicle (UAV)-assisted MEC architectures exploit UAVs’ flexible deployment to expand service coverage. However, existing approaches for multi-terminal, multi-UAV scenarios often fail to optimize task offloading latency, system energy consumption, and adaptability to dynamic environments simultaneously. They also overlook optimal UAV selection when terminal devices are covered by multiple UAVs and lack adaptive mechanisms to adjust optimization objectives during task execution. This study addresses these challenges by integrating cooperative caching, offloading decision-making, and resource allocation strategies.  Methods  A three-tier microcloud-edge-terminal architecture is constructed, comprising a central cloud, multiple UAV edge servers with caching capabilities, and numerous mobile terminal devices. A cooperative caching mechanism reduces transmission delay during task execution. Task offloading adopts a fine-grained partial offloading mode, dividing complex tasks into dependent subtasks modeled through a Directed Acyclic Graph (DAG). The Cooperative Caching-Adaptive Hierarchical MultiVerse Optimizer (CCAH-MVO) algorithm is proposed. A hybrid coding scheme encodes offloading decisions, caching decisions, and resource allocation uniformly. A dynamic weight mechanism adaptively balances delay and energy consumption according to the system’s real-time energy state. Additionally, a UAV selection strategy is implemented for scenarios where terminals are covered by multiple UAVs. By simulating inter-universe material exchange and local refined search, the algorithm efficiently determines the optimal offloading strategy. MATLAB simulations validate the method under various experimental settings.  Results and Discussions  The simulation scenario involves 50 randomly distributed terminal devices and 5 UAVs in a 400 m × 400 m area. UAVs are deployed above terminal cluster centers, while terminals at cluster edges are simultaneously within the coverage of multiple UAVs (Fig. 5). The optimal UAV for each terminal is selected using the UAV selection function (Fig. 6), preventing resource bottlenecks and achieving balanced load distribution. In terms of delay performance, the CCAH-MVO algorithm maintains the lowest task delay across all task volumes, with a gradual increase as the number of tasks grows (Fig. 7). Delay under CCAH-MVO is consistently lower than that under fixed-weight strategies across the full task range, demonstrating the effectiveness of the dynamic adaptive mechanism in preserving low latency (Fig. 10). For energy consumption, differences among the algorithms are minor when task quantities are low. Under high task loads, the activation of the dynamic weight mechanism flattens the energy consumption curve (Fig. 8). When the number of tasks reaches 100, total energy consumption under CCAH-MVO is the lowest among all strategies and remains lower than the fixed-weight approach, reflecting effective control under critical energy conditions (Fig. 9). Regarding total system overhead, the CCAH-MVO algorithm consistently achieves the best performance. The gap with fixed-weight strategies widens when task numbers exceed 80, illustrating the dynamic weight mechanism’s collaborative optimization of delay and energy consumption (Fig. 11). Overall, by integrating the dynamic weight mechanism and balancing load through UAV selection, the CCAH-MVO algorithm effectively mitigates resource constraints and high task processing overhead in complex, dynamic UAV-assisted MEC environments. It ensures precise coordination between task delay and energy consumption across different load stages.  Conclusions  The proposed CCAH-MVO framework, incorporating a microcloud-edge-terminal architecture, cooperative caching mechanism, fine-grained partial offloading, dynamic weight adjustment, and UAV selection strategy, effectively addresses resource scheduling in complex multi-UAV MEC environments. Simulations show adaptive optimization of objectives, intelligent energy management, low latency, and reduced total system overhead, improving service stability and user experience. This research provides a practical solution for efficient UAV edge computing in dynamic environments. Future work will explore dynamic energy efficiency optimization and multi-node collaboration while maintaining low-latency performance.
Joint Channel Estimation and Diagnosis for Blocked RIS-Assisted Multi-User Multipath Millimeter-Wave Systems
LI Shuangzhi, LIU Cong, WANG Ning, HAN Gangtao, GUO Xin
2026, 48(7): 3084-3093. doi: 10.11999/JEIT260093
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS) can effectively modulate Millimeter-Wave (mmWave) signals and reshape the wireless propagation environment. In practical deployments, however, RIS elements are vulnerable to adverse weather and physical obstructions, which cause unpredictable distortion and motivate joint channel estimation and blockage diagnosis. Most existing studies focus on single-user systems, whereas multi-user scenarios remain insufficiently studied. This gap creates an opportunity to exploit the common RIS blockage vector and the shared RIS-Base Station (BS) channel across users. This paper therefore proposes a low-complexity framework for joint channel estimation and blockage diagnosis by exploiting the sparsity and correlation of multi-user cascaded channels.  Methods  Under the assumption that all User Equipment (UE) shares the same RIS-BS channel and is affected by a common RIS blockage vector, the problem is divided into two stages. First, a target UE is selected. The sparsity of the mmWave channel and blockage vector, together with the linear dependence among RIS-BS paths, is used to formulate a sparse recovery problem. A hierarchical Bayesian model is then adopted, and an efficient Sparse Bayesian Learning (SBL) algorithm is used for joint recovery. Second, partial Channel State Information (CSI) obtained from the target UE is used to construct a common channel matrix that combines the RIS-BS channel and blockage information. Channel estimation for the remaining UEs is then reformulated as another sparse recovery problem.  Results and Discussions  A low-complexity strategy for cascaded channel estimation and blockage diagnosis is developed by exploiting the sparsity and correlation of multi-user cascaded channels and the commonality of the RIS blockage vector. Ideal estimation results are used as a theoretical lower bound, and the proposed algorithm is compared with two benchmark schemes. Simulation results show that the proposed algorithm consistently outperforms the benchmark schemes (Fig. 1). Specifically, a higher target-user Signal-to-Noise Ratio (SNR) improves the Normalized Mean Square Error (NMSE), which confirms the importance of target-user selection (Fig. 2). The algorithm also shows good convergence as the number of iterations increases (Fig. 3), and its performance approaches the ideal case more closely as the number of time frames increases (Fig. 4). In addition, the method remains robust as the number of blocked elements increases (Fig. 5). More BS antennas further improve performance by enhancing array orthogonality (Fig. 6). By exploiting path correlation, the proposed method achieves better estimation accuracy with slightly lower runtime (Table 1). However, estimation accuracy decreases as the number of paths increases because the model becomes more complex (Figs. 7 and 8).  Conclusions  This paper proposes a joint channel estimation and blockage diagnosis framework for blocked RIS-assisted multi-user multipath mmWave systems. Simulation results show that the method approaches the theoretical performance bound in complex multipath environments. It also maintains clear performance advantages under high blockage rates while reducing computational complexity through the use of common channel structures. This study provides a practical solution to performance degradation in RIS deployment, clarifies the effects of key parameters, and offers guidance for system design. Because practical blockages often exhibit block-sparse or structured-sparse characteristics, future work may incorporate structured priors, such as group sparsity and Markov random fields, into the SBL framework to capture spatial correlation and improve diagnostic accuracy and robustness.
Joint Power Allocation and AP On-Off Control for Long-Term Energy Efficient Cell-Free Massive MIMO Systems
WEI Siqi, GUO Fengqian, CHONG Baolin, CHENG Guo, LU Hancheng
2026, 48(7): 3094-3104. doi: 10.11999/JEIT260014
Abstract:
  Objective   With the rapid development of wireless communication technologies, Cell-Free Massive Multiple-Input Multiple-Output (CF-mMIMO) has emerged as an effective paradigm to overcome the limitations of traditional cell-centric networks, such as limited performance for edge users. By deploying a large number of distributed Access Points (APs) connected to a Central Processing Unit (CPU) to cooperatively serve users, CF-mMIMO improves spectral efficiency and macro-diversity gain. However, dense AP deployment also introduces a critical challenge: high energy consumption. In practical systems, if all APs remain continuously active, especially during periods of low traffic load, substantial and unnecessary energy consumption occurs. This behavior reduces network sustainability and conflicts with global “dual-carbon” goals. Existing studies on energy efficiency in CF-mMIMO systems mainly focus on short-term performance optimization. These short-term approaches often ignore long-term traffic dynamics and the requirement of queue stability. Therefore, they lack robustness under time-varying traffic conditions and may cause queue congestion and significant performance fluctuations, which are unacceptable for next-generation wireless networks with strict reliability requirements. Although several recent studies examine long-term energy efficiency optimization, most assume that all APs remain active at all times. Therefore, the energy-saving potential of adaptive AP on-off control is not fully utilized.  Methods   To address these issues, a joint power allocation and AP on-off control strategy is proposed for downlink CF-mMIMO systems. The optimization problem aims to maximize long-term energy efficiency subject to user queue stability and AP power constraints. Because the problem has stochastic and long-term characteristics, the Lyapunov optimization framework is applied to transform the original long-term fractional programming problem into a sequence of deterministic drift-plus-penalty minimization problems solved in each time slot. The resulting per-slot problems remain nonconvex. Therefore, each problem is decomposed into two subproblems: power allocation and AP on-off control. The Successive Convex Approximation (SCA) method is used to convert the nonconvex formulations into solvable convex problems. An alternating optimization algorithm is then developed to jointly solve the two subproblems, which enables adaptive resource configuration under dynamic network conditions and stochastic traffic arrivals.  Results and Discussions   The proposed algorithm is evaluated through extensive simulations. First, the convergence behavior is examined. Numerical results (Fig. 2) show that per-slot energy efficiency increases rapidly and stabilizes after several iterations, which verifies the convergence of the alternating optimization procedure. Second, the effect of the control parameter is analyzed. As the parameter increases, the algorithm places greater emphasis on energy efficiency. Average power consumption decreases and then stabilizes (Fig. 3), whereas long-term energy efficiency increases and eventually stabilizes (Fig. 4). These results confirm the trade-off between energy efficiency and queue stability. Third, the proposed scheme is compared with three baseline methods. The results (Fig. 5) show that the proposed joint optimization approach consistently achieves higher long-term energy efficiency than the baseline methods. Fourth, the necessity of long-term optimization is demonstrated by comparing queue lengths with a short-term baseline (Fig. 6). Under the same traffic arrival rate, the short-term method shows cumulative queue growth, whereas the Lyapunov-based approach maintains queue lengths within a stable range and ensures network stability. Finally, robustness under imperfect Channel State Information (CSI) is evaluated (Fig. 7). Although energy efficiency decreases as channel uncertainty increases, the proposed method consistently outperforms the baseline approaches, which demonstrates strong robustness to channel estimation errors.  Conclusions   A long-term energy efficiency optimization framework is proposed for CF-mMIMO systems with stochastic traffic arrivals. By applying Lyapunov optimization theory, the stochastic long-term problem is transformed into slot-level drift-plus-penalty problems based on queue states. This transformation enables per-slot resource scheduling decisions while maintaining queue stability. On this basis, an efficient joint resource scheduling algorithm that integrates power allocation and AP on-off control is developed. The original problem is decomposed into power allocation and AP on-off control subproblems and solved through alternating optimization. Simulation results show that the proposed method adapts to dynamic traffic conditions. By placing underutilized APs into sleep mode, the algorithm improves long-term system energy efficiency and maintains queue stability. These results provide guidance for the design of green and sustainable wireless networks.
SG-DDPG-based Low-intercept Point Beam Design for FDA-MIMO Short-range Detectors
JIA Jinwei, GAO Min, HAN Zhuangzhi, LIU Limin, YIN Yuanwei
2026, 48(7): 3105-3122. doi: 10.11999/JEIT260010
Abstract:
  Objective  Radio short-range detectors are widely used in many detection systems. However, in modern battlefields, the electromagnetic environment is increasingly complex, and radio short-range detectors must withstand various forms of electromagnetic interference. In particular, fourth-generation jammers based on Digital Radio Frequency Memory (DRFM) can implement repeater deception jamming. Such jamming may cause failures such as premature detonation in radio short-range detectors and reduce their damage effectiveness. Anti-repeater deception jamming has therefore become a key issue for short-range detectors. Improving the Low Probability of Intercept (LPI) performance of radio short-range detectors is an effective means of resisting repeater deception jamming. According to the Chinese manuscript, this study focuses on the effect of FDA-MIMO array-element frequency-offset settings on beam synthesis and proposes an SG-DDPG-based method for LPI point beam design.  Methods  Frequency Diverse Array-Multiple-Input Multiple-Output (FDA-MIMO) technology is used in this study, and the key factors affecting beam convergence are analyzed. For the spatial LPI beam design of radio short-range detectors, a performance evaluation model for spatial LPI beams is constructed. An FDA-MIMO LPI point beam design method based on the Stage Guidance-Deep Deterministic Policy Gradient (SG-DDPG) algorithm is then proposed. In the SG-DDPG algorithm, a multidimensional staged guidance reward function is designed. An Actor-Critic model is used to maximize the reward value through gradient ascent. The array-element frequency offsets that provide better beam convergence in the current environment are then obtained. The SG-DDPG algorithm is suitable for LPI point beam design under different fall angles of radio short-range detectors. It overcomes the technical limitation of formula-based frequency-offset calculation, which is only applicable when the detector fall angle is close to vertical.  Results and Discussions  The simulations show that, after the array-element frequency offsets are optimized by the SG-DDPG algorithm, the FDA-MIMO beam achieves a half-power beam width of 1 m in the range dimension and 9.9° in the angular dimension. The proposed method provides better beam convergence and LPI performance than classical frequency-offset design methods. These results indicate that the proposed algorithm offers an effective approach for array-element frequency-offset optimization and LPI point beam design, thereby improving the LPI performance of radio short-range detectors.  Conclusions  This paper presents an FDA-MIMO LPI point beam design method based on the SG-DDPG algorithm, with the array-element frequency offset used as the optimization objective. The simulation results support two main conclusions. First, the proposed method removes the restriction that the fall angle of the radio short-range detector must be close to vertical when the array-element frequency offset is calculated by a formula-based method. The algorithm can be applied to LPI beam design under different fall angles and improves the LPI performance of radio short-range detectors. Second, the proposed method achieves a half-power beam width of only 1 m in the range dimension and 9.9° in the angular dimension, which is better than that of traditional methods. Under different fall angles, the beam formed by the proposed method has the smallest intercept area, indicating the best LPI performance.
Index Modulation Design with Sparse Spatial Constellation and Dynamic Multi-RIS-Block Selection for RIS-MIMO Systems
HUANG Fuchun, ZHU Han, TANG Xiaoqing, YANG Fan, HUANG Jie
2026, 48(7): 3123-3134. doi: 10.11999/JEIT251289
Abstract:
  Objective  Reconfigurable Intelligent Surface (RIS)-assisted Multiple-Input Multiple-Output (MIMO) Index Modulation (IM) systems face two main challenges: the difficult deployment of a single large-scale RIS panel and the high design complexity of efficient transmit spatial signal vectors. To address these issues, a joint design that combines sparse spatial constellation and dynamic multi-RIS-block selection is proposed. The design improves spectral efficiency, Bit Error Rate (BER) performance, and deployment flexibility.  Methods  Inspired by the Extended Space Index Modulation (ESIM) paradigm, a sparse Spatial Constellation with Two Active Antennas (SCTA) is proposed, forming the SCTA-RIS-SM system. In this design, Pulse Amplitude Modulation (PAM) and Secondary PAM (SPAM) constellations are combined to construct the spatial constellation vector [x1,x2]T, which is modulated onto two active antennas. This design maximizes the Minimum Euclidean Distance (MED) between transmit vectors and improves the anti-interference capability of the system. To address the deployment difficulty of a single large-scale RIS panel, an enhanced SCTA-MBRIS-SM system is further proposed. The system uses a distributed array of small RIS blocks and dynamically selects a subset of blocks for cooperative reflection. Different RIS block selection combinations are used as a new IM dimension. Spectral efficiency and average BER are then analyzed theoretically. Monte Carlo simulations are conducted to compare the proposed systems with several existing schemes.  Results and Discussions  The simulation results show that the proposed SCTA-RIS-SM system achieves clear Signal-to-Noise Ratio (SNR) gains over RIS-SIM, RIS-SM, and DHRIS-SM systems at the same spectral efficiency, such as 10–12 bits/(s·Hz). For instance, when BER = 10–3, SCTA-RIS-SM outperforms RIS-SIM by approximately 1.5–2.5 dB and DHRIS-SM by more than 6 dB. By using additional IM from RIS block selection, SCTA-MBRIS-SM further improves BER performance and spectral efficiency compared with SCTA-RIS-SM, without increasing the number of Radio Frequency (RF) chains. With the same total number of reflecting elements, the proposed multi-RIS-block scheme achieves an SNR gain of up to 5 dB over RIS-SIM when BER = 10–3. The theoretical BER curves agree well with the simulation results in the high-SNR region, confirming the validity of the analytical derivations. The results also indicate that the performance advantage is maintained as the number of transmit antennas increases. In addition, the proposed design is compatible with channel coding.  Conclusions  This paper addresses the challenges of large-scale RIS deployment and high-complexity spatial signal design in RIS-assisted MIMO systems. The proposed SCTA design improves system reliability by optimizing the Euclidean distance distribution in the signal space. Dynamic multi-RIS-block selection transforms hardware deployment constraints into a new dimension for improving spectral efficiency, providing a feasible path for practical large-scale RIS applications. Simulation results confirm that joint optimization of transmit spatial vectors and RIS reflection degrees of freedom is an effective strategy for improving system performance. Future work will focus on robust design under imperfect channel state information, construction of higher-dimensional sparse constellations, extension to extremely large-scale MIMO scenarios, and multi-user communications.
Radar, Sonar,Navigation and Array Signal Processing
Research on Inverse QR Decomposition Optimization for Sparse Adaptive System Identification Algorithms
PENG Yi, ZHANG Pengfei, WANG Xiaoyong, GAO Junqi, LI Changlong, ZHANG Zhiyuan, SUN Tianxiang
2026, 48(7): 3135-3145. doi: 10.11999/JEIT250562
Abstract:
  Objective  Traditional sparse-regularized Recursive Least Squares (RLS) algorithms, namely L1/L0-norm Recursive Least Squares (L1/L0-RLS), have theoretical advantages in sparse parameter-space estimation and are widely used in system identification and channel equalization. However, under limited numerical precision, iterative covariance matrix computation may cause rounding errors to accumulate. This can lead to divergence and instability in the least-squares solution.  Methods  To address this problem, an improved algorithm based on the Inverse QR Decomposition (IQRD) framework is proposed. The framework suppresses rounding-error accumulation in traditional regularized RLS algorithms. It also removes the back-substitution step for weight coefficients required in conventional QR decomposition. These features improve numerical robustness and system identification efficiency in finite-precision environments. Specifically, L1-IQRD-RLS and L0-IQRD-RLS algorithms are constructed under an L1/L0-constrained IQRD architecture. A general recursive expression for the weight coefficients is derived. An automatic parameter selection mechanism is also incorporated into the algorithm framework to solve the dynamic optimization problem of the sparse regularization parameter.  Results and Discussions  Monte Carlo simulations are conducted to evaluate the sparse constraints and robustness of the proposed algorithms. The results show that L1-IQRD-RLS and L0-IQRD-RLS maintain long-term numerical stability in an 11-decimal-place fixed-point computing environment. Compared with traditional algorithms, the proposed algorithms show clear advantages in system sparsity representation, parameter estimation variance, and covariance matrix condition number. Measured-data verification further confirms that the improved algorithms maintain numerical stability under limited-precision conditions and are more robust than traditional methods. The measured-data results also show that the regularized RLS algorithms optimized by the IQRD framework have advantages in system sparsity representation, parameter estimation, and numerical stability. Their iterative convergence success rate is higher than that of traditional methods.  Conclusions  This paper addresses sparse system identification in adaptive filtering. Traditional sparse-regularized RLS algorithms still face numerical stability problems under limited numerical precision. To solve this problem, an IQRD framework is constructed to reduce the numerical ill-conditioning caused by accumulated rounding errors in sparse-regularized RLS algorithms. The proposed method improves numerical robustness in low-precision environments. In addition, an automatic parameter selection mechanism is incorporated into the algorithm framework. This reduces repeated parameter tuning and supports stable performance optimization under sparse constraints. In practical electromagnetic signal processing, system identification and beamforming are limited by the finite precision of hardware implementation and often exhibit inherent system sparsity. The proposed algorithm provides a targeted solution. Its finite-word-length robustness suppresses numerical divergence during adaptive weight updates and supports stable implementation on fixed-point processors. The sparse constraints also match the physical characteristics of sparse systems and improve estimation accuracy. This study provides a practical algorithm for high-performance and high-stability sparse-constrained systems on precision-limited hardware platforms.
Optimal Weighted Subspace Fitting-based Direct Position Determination with HF/VHF Collaboration
YANG Gaoyuan, YIN Jiexin, WANG Ding, YANG Bin
2026, 48(7): 3146-3161. doi: 10.11999/JEIT260001
Abstract:
  Objective   Passive localization is essential for target detection, navigation, and track tracking, particularly in military applications involving maritime and aerial targets. These targets often transmit across multiple frequency bands, including shortwave High Frequency(HF) and Very High Frequency (VHF). Existing localization methods largely rely on single-band approaches or two-step positioning techniques. Single-band methods underutilize the positional information available across different bands, while two-step methods lose information during intermediate parameter estimation (e.g., Direction-Of-Arrival (DOA); Time-Difference-Of-Arrival (TDOA)), reducing localization accuracy. Collaborative fusion of HF signals (via ionospheric reflection) and VHF signals (via Doppler effects from moving arrays) has been rarely addressed. To overcome low positioning accuracy and limited spatial resolution in over-the-horizon multi-target scenarios, this study proposes a novel collaborative Direct Position Determination (DPD) method designed to integrate the complementary strengths of HF and VHF signals, enhancing localization precision and robustness in complex electromagnetic environments.  Methods  An Optimal Weighted Subspace Fitting (OWSF) DPD algorithm is proposed. Comprehensive signal propagation models are established for heterogeneous observation platforms (Fig. 1). HF signal propagation is modeled using a two-dimensional DOA framework based on ionospheric reflection, incorporating azimuth and elevation angles to handle nonlinear over-the-horizon propagation. VHF signals are modeled using a space-time extended signal framework for a moving Unmanned Aerial Vehicle (UAV), exploiting Doppler effects to create a virtual large-aperture array that captures both one-dimensional angle and Frequency-Of-Arrival (FOA) information. Unlike traditional methods that process each band separately, the OWSF algorithm constructs a unified cost function that fuses the signal and noise subspaces of both HF and VHF data using optimal weighting matrices, balancing the contributions of different signal qualities. Target positions are then estimated by minimizing this cost function via grid search or Newton iteration. The Cramér-Rao Bound (CRB) under Earth-ellipsoid constraints is derived to provide the theoretical performance limit.  Results and Discussions   Simulations are conducted in a centralized processing scenario, where HF stations and UAV VHF signals are transmitted to a central station for joint processing (Fig. 2). The simulation involves three stationary targets and a collaborative system comprising HF stations and a UAV (Fig. 3, Table 2, Table 3). Performance comparisons demonstrate that the OWSF method consistently outperforms traditional two-step positioning methods and single-system DPD methods (DOA-only or FOA-only) in Root Mean Square Error (RMSE) (Fig. 4). When HF SNR is 5 dB lower than VHF SNR, OWSF exhibits superior robustness compared to Subspace Data Fusion (SDF) and Minimum Variance Distortionless Response (MVDR) methods, approaching the CRB at high SNR (Fig. 5). The impact of system parameters is further analyzed, showing that increasing the number of sampling points (Fig. 6) and array elements (Fig. 7) improves accuracy, particularly in low SNR regimes. Regarding spatial resolution, the OWSF algorithm generates sharper spectral peaks for distant targets and successfully resolves closely spaced targets that the SDF-DPD algorithm fails to distinguish (Fig. 7, Fig. 8).  Conclusions   The HF/VHF collaborative DPD method effectively integrates multidimensional observational information from ionospheric reflection and Doppler-based propagation. Simulation results demonstrate substantial improvements in localization accuracy, spatial resolution, and robustness, especially under low-SNR conditions or heterogeneous signal quality between bands. The derived CRB provides a solid theoretical benchmark, confirming that the method overcomes the limitations of single-band and two-step approaches. This approach offers a highly effective solution for over-the-horizon passive localization of multiple stationary targets.
An Ultra-Wideband Low-Profile Dipole Patch Antenna for VHF-Band Probing Radars
TIAN Yuxiao, ZHANG Feng, MA Zhangjun, WANG Jiacheng, JI Yicai
2026, 48(7): 3162-3169. doi: 10.11999/JEIT260105
Abstract:
  Objective  In radar systems, the limitations of traditional narrowband antennas in data transmission rate and resolution have become increasingly evident. Ultra-WideBand (UWB) antennas therefore receive broad attention because they provide high range resolution and strong interference suppression capability. However, at low frequencies, existing UWB antennas usually suffer from excessively large physical size, which makes installation on airborne or vehicle-mounted platforms difficult. By contrast, compact antennas that are easier to deploy often exhibit insufficient gain and cannot satisfy the penetration-depth requirement of deep subsurface detection. Thus, achieving a proper balance among antenna size, bandwidth, and gain over an ultra-wideband range remains a major challenge for VHF-band probing radars. To address this issue, a planar dipole antenna loaded with an Artificial Magnetic Conductor (AMC) structure and metallic shorting walls is proposed. The antenna maintains stable radiation performance over a wide frequency range while preserving a low-profile and structurally simple configuration.  Methods  The reflection-phase characteristics of AMC unit cells with different geometries are compared, and square unit cells are selected to construct a 9 × 7 AMC reflective layer. Owing to its in-phase reflection property, the AMC structure removes the conventional requirement for a quarter-wavelength spacing between the antenna and a metallic ground plane, thereby reducing the profile height. The dipole patch adopts an optimized meandered current-bending structure to reduce the lateral size. Metallic shorting walls are further loaded at both ends of the antenna. According to image theory, equivalent currents are generated on the outer surfaces of these metal walls during operation, which effectively extends the electrical length and improves low-frequency performance without increasing the physical size. In addition, two vertical metallic walls are connected to the ground plane on both sides of the antenna to form a reflective back cavity, which strengthens unidirectional radiation and improves antenna gain. As part of the overall co-design, four 125 Ω resistors are inserted between the feed region and the metallic sidewalls. This resistive loading suppresses strong low-frequency resonances and broadens the impedance bandwidth at the cost of acceptable Ohmic loss.  Results and Discussions  A prototype with favorable simulated performance is fabricated and measured in a microwave anechoic chamber. The measured impedance bandwidth for VSWR<2 is 50~400 MHz, which agrees well with the simulated range of 84~366 MHz. The measured impedance matching is slightly better than the simulated result, mainly because cable loss and power-divider loss in the feeding network reduce the reflected power. The measured gain follows the same trend as the simulated gain, with deviations within 1 dBi. Radiation-pattern measurements show that at 100, 200, and 300 MHz, the measured copolarization patterns agree well with the simulated results, and the maximum radiation direction remains normal to the antenna plane, which confirms the effectiveness of the proposed design. As shown in Fig. 5, the current on the radiating patch layer mainly flows along the +x direction and generates a radiated electric field along the +z direction. The current on the AMC unit can be represented by an equivalent current loop oriented along the +z direction. At this frequency, the x-direction current and the parasitic current loop on the AMC jointly enhance the antenna gain. This result explains the gain-improvement mechanism of the AMC structure. When the operating frequency increases to 400 MHz, the electrical size of the antenna reaches approximately \begin{document}$ 1.6\lambda $\end{document}, which causes main-lobe splitting and shifts the maximum radiation direction toward 90°. Although this high-frequency beam splitting introduces spatial clutter, it is an acceptable physical trade-off for achieving the ultra-low profile of 0.07 λL, while the overall UWB characteristic still supports high time-domain resolution in probing radar systems. At 400 MHz, the measured H-plane co-polarization level is slightly higher than the simulated value, possibly because of coupling between the feeding cable and the vertically mounted antenna.  Conclusions  A low-profile UWB planar dipole antenna is proposed for VHF-band probing radar applications. By combining the AMC layer, metallic shorting walls, and resistive loading, the proposed design improves impedance matching while preserving a compact size. The reflective back cavity further improves the realized gain. The fabricated prototype shows good agreement between measurement and simulation. The antenna operates over 100~366 MHz and exhibits a measured VSWR<2 bandwidth of 50~400 MHz. It maintains a compact electrical size of 0.38λL × 0.18λL × 0.07λL, and the maximum measured gain within the operating band reaches 6 dBi. The proposed co-design provides a practical solution for low-frequency probing radar antennas that require wide bandwidth, low profile, and relatively high gain.
Circuit and System Design
Optimizing Satisfiability-Based Automatic Test Pattern Generation Systems: Unified Fault Set Construction,Modeling, and Solving
YAN Dapeng, HE Qirun, GUO Jing, WANG Boning, CAI Zhikuang
2026, 48(7): 3170-3180. doi: 10.11999/JEIT260025
Abstract:
  Objective  Boolean SATisfiability-Based Automatic Test Pattern Generation (SAT-Based ATPG) is widely used to generate tests for hard-to-detect single stuck-at faults and to prove fault untestability in combinational logic. When SAT-Based ATPG is applied to large netlists with dense fanout and reconvergence, its runtime and memory consumption are often dominated by three interacting issues. Representative fault lists produced by conventional dominance- or equivalence-based fault collapsing can remain large, increasing the number of SAT calls and enlarging the incremental context that must be maintained across faults. Meanwhile, SAT modeling may introduce redundant Conjunctive Normal Form (CNF) overhead, especially when an explicit faulty-circuit copy is constructed or when propagation constraints are encoded globally without locality control. In addition, fanout-reconvergence structures amplify assignment correlations along sensitized paths, and such correlations are often exposed only after repeated decisions and backtracking when only standard unit propagation is used. The unified optimization objective is therefore to reduce overall CNF size and solving cost while preserving completeness, so that a practical SAT-Based ATPG system remains efficient and stable across circuits of different scales.  Methods  A three-part framework is developed and implemented in an incremental SAT-Based ATPG flow, and the overall workflow is illustrated (Fig. 1). First, a checkpoint-driven dynamic fault-set construction method is proposed. Checkpoints are collected during netlist-to-directed-acyclic-graph conversion, including all primary inputs and all fanout branches, and XOR/XNOR outputs are additionally recorded as supplementary checkpoints to avoid over-collapsing XOR-related fault behavior. Representative faults are initialized on checkpoints by compact rules that combine dominance-oriented fault collapsing with equivalence-aware refinement, and solver-guided repair is performed when an untestable representative fault indicates potential masking under structural constraints. The procedure is summarized in Algorithm 1. Second, an SAT modeling method based on fault sensitization constraints is adopted to avoid explicit faulty-circuit duplication. Fault activation, propagation, and observability are represented by additional fault sensitization constraints over the original circuit variables, and auxiliary variables are introduced only when local bookkeeping is required. Constraint localization is restricted to the fault fanout cone, and cone-boundary and internal vertices are identified through a graph-traversal procedure (Fig. 2). Third, a dynamic implication learning mechanism oriented to fanout-reconvergence pairs is integrated into the incremental solving loop. Reconvergence pairs within the fault fanout cone are monitored under partial assignments, and structure-induced implications are injected either as implied assignments when a reconvergent output becomes functionally determined or as short conflict clauses when a branch-value combination becomes inconsistent with the fault sensitization constraints. The dynamic implication learning procedure is summarized in Algorithm 2.  Results and Discussions  The unified system is evaluated on ISCAS’85 and ISCAS’89 benchmark circuits, with TG-PRO used as the baseline implementation under the same SAT solver and termination settings. The checkpoint-driven dynamic fault-set construction method substantially reduces the representative fault space entering ATPG. Relative to the uncollapsed fault space, the average representative-fault ratio decreases from 51.38% to 42.41%, corresponding to an average fault-space reduction of 57.59%. The best-case ratio reaches 33.19% on large circuits with heavy reconvergence, which indicates that checkpoint-centered representative-fault allocation effectively suppresses redundancy without enlarging the untestable fault set (Table 1). The reduced fault-set size is reflected in preprocessing efficiency, and the total runtime for fault-set construction is consistently reduced, with an average reduction of 8.37% across the evaluated circuits (Fig. 3). For SAT model construction, the fault-sensitization-constraint encoding reduces CNF overhead relative to the baseline model construction. Across the benchmark set, the numbers of CNF clauses and CNF variables are reduced by 11.44% and 3.50%, respectively, which shows that avoiding explicit faulty-circuit duplication and localizing auxiliary constraints to the fault fanout cone effectively lowers memory demand (Table 2). The reduced CNF size and strengthened locality of constraints are further reflected in end-to-end runtime, and the total runtime of SAT modeling and solving is reduced across the evaluated benchmarks (Fig. 4). Dynamic implication learning further improves solving efficiency in reconvergence-heavy structures. Compared with static implication learning, CNF construction time increases by 3.0% on average because of the additional monitoring and injection operations, yet the overall runtime decreases by 4.42% on average, which indicates a favorable cost-benefit trade-off. The overhead attributed to dynamic implication learning accounts for 2.51% of the total runtime aggregated across circuits, which confirms that the injected implications and pruning clauses provide measurable solving benefits at limited extra cost (Table 3).  Conclusions  A unified optimization framework for SAT-Based ATPG is developed by combining checkpoint-driven dynamic fault-set construction, localized fault sensitization constraints for CNF modeling, and fanout-reconvergence-oriented dynamic implication learning. Representative faults are compressed through solver-guided repair of dominance and equivalence relations to avoid masking, CNF growth is controlled through duplication-free modeling localized to the fault fanout cone, and reconvergence correlations are exploited through incremental implication injection to strengthen propagation and enable early conflict pruning. Experimental results on standard benchmark circuits show consistent reductions in representative fault scale, CNF size, and total runtime, providing a practical approach for scaling SAT-Based ATPG to larger designs with complex fanout and reconvergence.
Design of a Timing-Controlled Nonvolatile Flip-Flop for Low-ON/OFF-Current-Ratio FeFETs
DU Shimin, YANG Chang, WANG Lunyao, ZHANG Zhe
2026, 48(7): 3181-3192. doi: 10.11999/JEIT251059
Abstract:
  Objective  Nonvolatile Processors (NVPs) are a key technology for Internet of Things (IoT) and energy-harvesting systems, in which computational states must be preserved during unexpected power loss. Conventional volatile processors rely on external Nonvolatile Memory (NVM) for state retention. However, this approach causes high latency and energy overhead. Integrated Nonvolatile Flip-Flops (NVFFs) based on Ferroelectric Field-Effect Transistors (FeFETs) provide a promising alternative by enabling on-chip state backup and recovery. However, existing single-ended FeFET-based flip-flops are prone to contention-induced recovery failures, especially when the FeFET ON/OFF current ratio degrades. This failure arises from contention among internal metal-oxide-semiconductor transistors, which makes internal node settling uncertain and causes unreliable state recovery. To address this issue, this paper proposes a timing-controlled NVFF architecture that replaces contention-based recovery with a two-stage recovery mechanism. The proposed design aims to achieve reliable recovery under degraded FeFET ON/OFF current ratios as low as 102, improve timing metrics such as hold time and clock-to-Q delay, and maintain low energy consumption for IoT applications.  Methods  The proposed design extends the Static Contention-Free Single-Phase-Clocked Flip-Flop (SSCFF), whose fully static structure suppresses internal node contention. On this basis, one FeFET and five additional Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs) are integrated to construct a single-ended NVFF. Two control signals, RES and MOD, are used to manage the recovery process. In normal operation, MOD = 0, and the circuit functions as a conventional SSCFF while supporting runtime state backup. In recovery mode, MOD = 1, and the recovery process is divided into two stages. In the precharge stage, when RES = 0, the internal nodes are precharged to VDD. In the selective-discharge stage, RES switches from low to high, and the FeFET resistance state determines whether discharge occurs. If the FeFET is in the Low-Resistance State (LRS), a discharge path is formed, and the node voltage is pulled down to ground. If the FeFET is in the High-Resistance State (HRS), the node retains its charge until the next clock edge. This precharge-selective-discharge sequence removes recovery contention and enables deterministic internal node settling. The design is implemented using a 130 nm Complementary Metal-Oxide-Semiconductor (CMOS) process and an integrated FeFET model. Simulations are performed in Cadence Virtuoso across a supply voltage range of 0.6~0.9 V and FeFET ON/OFF current ratios from 102 to 104. Key metrics, including setup time, hold time, clock-to-Q delay, recovery energy, and recovery success rate, are evaluated and compared with those of a conventional Transmission-Gate Flip-Flop (TGFF).  Results and Discussions  Simulation results show that timing-controlled recovery improves reliability under severe FeFET degradation. At an FeFET ON/OFF current ratio of 102, the proposed flip-flop achieves a 100% recovery success rate in 2000 Monte Carlo simulations. This improvement is attributed to the removal of contention among internal recovery paths. Timing metrics are also improved. The 3σ worst-case hold time is reduced by 64.6%, and the clock-to-Q delay is reduced by 33.9%. Although setup time increases slightly, this increase can be mitigated through device sizing. Recovery energy remains at the fJ level, with values of approximately 10 fJ under the tested conditions. This energy is only slightly higher than that of the TGFF because of the added precharge stage.  Conclusions  An FeFET-based NVFF with timing-controlled two-stage recovery is presented to address the contention-induced failure modes that limit low-voltage recovery reliability. By integrating a single FeFET into an enhanced SSCFF structure and using the RES signal to control precharge and selective discharge, the proposed design maintains a high recovery success rate even under severely degraded FeFET ON/OFF current ratios. It also improves hold time and clock-to-Q delay compared with conventional transmission-gate NVFFs. The proposed architecture provides an effective solution for energy-constrained IoT processors that require fast and reliable state preservation under unpredictable power conditions.
A Frequency Domain Self-Attention Guided MultiscaleInverse Lithography Technology
LUO Binling, WANG Ying, CAI Shuting
2026, 48(7): 3193-3202. doi: 10.11999/JEIT251382
Abstract:
  Objective  Optical Proximity Effect (OPE) in lithographic processes causes printed wafer patterns to deviate from target layouts. Therefore, Optical Proximity Correction (OPC) is required for mask optimization before exposure. Traditional rule-based OPC methods show reduced accuracy for complex layouts, whereas model-based OPC methods require high computational cost. Deep learning-based methods have recently been used to accelerate mask generation. However, their limited receptive fields make it difficult to model long-range optical interference, which restricts optimization accuracy. To address these limitations, this work proposes Frequency-Domain Self-Attention-Guided Multiscale Inverse Lithography Technology (FMS-ILT). The method jointly models local geometric details and global optical interference to improve printed image fidelity, edge placement accuracy, and process robustness.  Methods  FMS-ILT uses a residual convolution-based multiscale encoder-decoder architecture. Shallow layers extract fine geometric features, such as edges and corners, whereas deeper layers capture large-scale layout context. Residual blocks and multilevel skip connections are used to preserve high-frequency information and stabilize training. To overcome the limited receptive field of spatial convolutions, a Frequency-domain Self-Attention Mechanism (FSAM) is introduced at the encoder output. Global feature interactions are modeled using the Fourier transform. The resulting attention responses are then mapped back to the spatial domain through the inverse Fourier transform to adaptively reweight feature representations. A two-stage training strategy is adopted. During pretraining, a dual-branch structure jointly learns mask geometry and imaging consistency, providing physically meaningful initialization. During main training, lithography simulation is applied under nominal, maximum, and minimum process corners to refine mask optimization under physical constraints.  Results and Discussions  The comparison results with baseline models are summarized in Tables 2 and 3. FMS-ILT is used as the reference method (Ratio = 1), and all experiments are conducted on the LithoBench dataset. For the overall imaging \begin{document}$ \mathcal{L}2 $\end{document} error, FMS-ILT achieves the lowest value of 19,998, outperforming the baseline models by 2%~107%. For Process Variation Band (PVB), GAN-OPC obtains the best value of 19 156, which is 31% lower than that of FMS-ILT. However, its \begin{document}$ \mathcal{L}2 $\end{document} error and Edge Placement Error (EPE) are 107% and 1 115% higher, respectively, indicating an imbalance between imaging fidelity and edge accuracy. The remaining baseline models show PVB performance comparable to that of FMS-ILT. For EPE, FMS-ILT also shows a clear advantage, achieving an average value of 1.95, which is 47%~1 115% lower than those of the baseline models. These improvements are mainly attributed to the multiscale encoder-decoder fusion mechanism, which integrates local and global features; the combination of attention mechanisms and frequency-domain operations, which guides the model toward critical regions; and the dual-branch pretraining strategy, which introduces physical priors into the network. These modules enable FMS-ILT to achieve balanced performance in imaging fidelity, process stability, and edge accuracy.  Conclusions  This work proposes FMS-ILT for mask optimization in computational lithography. The model uses a residual convolution-based multiscale encoder-decoder architecture to extract rich spatial features. It also incorporates FSAM to jointly model local geometric details and global optical interference. A two-stage training strategy is used. In the pretraining stage, mask generation and target image reconstruction are used as dual-branch tasks to improve the physical consistency between the mask and the printed image. In the main training stage, lithography simulation is introduced to further improve imaging accuracy and process robustness. Experimental results on the public LithoBench dataset show that FMS-ILT achieves strong performance in terms of L2, PVB, and EPE. The method improves printed image quality and provides a feasible and efficient solution for computational lithography.
A Lightweight and High-Reliability Challenge Generation Strategy for APUF
LAN Guohao, ZHANG Hui, DUO Bin, WANG Zibin, ZHOU Rang, LI Dongfen
2026, 48(7): 3203-3212. doi: 10.11999/JEIT251073
Abstract:
  Objective  The Arbiter Physical Unclonable Function (APUF) is a lightweight security primitive widely used for identity authentication and key generation in resource-constrained devices. However, its response consistency is highly sensitive to environmental perturbations. The same challenge may therefore produce inconsistent responses under different conditions, which reduces the reliability of APUF-based security systems. Existing reliability improvement schemes mainly rely on hardware modification or challenge screening. These schemes often require high resource overhead and have low efficiency. To address these limitations, a Delay-constrained Challenge Generation Strategy (DCGS) is proposed to improve APUF reliability without additional hardware overhead or inefficient candidate screening.  Methods  DCGS models APUF path-delay characteristics and constructs challenges with constrained delay differences to ensure response stability. First, a Logistic Regression (LR) model is established to characterize the relationship between challenge bits and path delays. A delay-weight vector is then derived from the trained LR model to quantify the contribution of each challenge bit to the overall path delay. Second, a two-stage challenge generation mechanism is designed for delay-constraint control. In the first stage, prefix-bit initialization generates different prefix sequences to establish a delay baseline for subsequent bitwise extension. In the second stage, bitwise extension dynamically determines each remaining challenge bit according to the delay-weight vector. During this process, the cumulative delay difference of each challenge is monitored in real time and maintained within a preset delay-difference threshold range. Unlike conventional screening methods that post-process candidate challenges, DCGS directly generates stable challenges by design. This design removes the need for candidate challenge pools and improves generation efficiency.  Results and Discussions  DCGS is evaluated under different noise intensities. At a noise intensity of 0.3, which represents the maximum practical noise level, the reliability of DCGS-generated challenges remains 100% (Fig. 2). For generation efficiency, DCGS requires only 0.017 s to generate 10 000 challenges (Table 4). The response uniformity reaches 50.02% (Table 4), and the uniqueness reaches 50.46% (Table 4). Both metrics are close to the ideal theoretical value of 50%. The security analysis shows that the average bit entropy of DCGS-generated challenges is 0.980 7 (Fig. 3). The conditional entropy is 0.987 8, only 0.002 3 lower than that of random challenges (0.990 1).  Conclusions  This paper proposes DCGS for APUF to address inconsistent responses, low generation efficiency, and high hardware resource consumption in traditional schemes under high-noise conditions. By modeling path-delay characteristics with LR and combining prefix-bit initialization with bitwise extension, the proposed strategy ensures that the generated challenges satisfy the preset delay-difference threshold range. DCGS achieves high reliability, high efficiency, and good response uniformity without increasing hardware overhead. Experimental results show that DCGS improves APUF reliability in complex environments and supports secure applications in resource-constrained devices.