SearcharxivSearch

arXiv subjects

Wenwen Zhang

Publications and source records attributed to Wenwen Zhang.

At least 19 recordsLinked to original sources

CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion

Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with precise curvilinear structure objects remains an open challenge. To address this, we propose CSGen, a hierarchical multimodal diffusion model that synthesizes high-fidelity images precisely aligned with multiple control conditions. The CSGen is built upon three key innovations: 1) We construct a multi-domain and multimodal dataset, including over 24K samples from 5 domains and 7 different types of annotations, to train the unified generation model. 2) We propose a novel hierarchical progressive control strategy that decouples topology clues from visual context by a phased signal injection, mitigating semantic drift while ensuring the topological integrity of sparse structures. 3) We design a sparsity-aware loss re-weighting mechanism to address the extreme sparsity of curvilinear structures, significantly enhancing the attention on thin and fragile structures during optimization. Extensive experiments demonstrate that CSGen generates images with superior structure accuracy and visual realism, significantly improving downstream segmentation performance while maintaining robustness across diverse prompts. Our results confirm CSGen as a scalable, data-centric paradigm for the analysis of complex curvilinear structures in diverse multimedia applications. Code and dataset are available at https://github.com/ShanZard/CSGen.

cs.CV

How Infrastructure and Streetscape Shape E-Scooter Route Choice: Evidence from Washington, DC

E-scooters have emerged as an important micromobility mode for short urban trips, yet evidence on route choice behavior remains limited. This study examines e-scooter route choice in Washington, DC using GPS trajectory data and a Path Size Logit model. In addition to roadway and infrastructure characteristics, the model incorporates visual streetscape features extracted from Google Street View imagery using computer vision techniques. The results show that the effectiveness of cycling infrastructure depends strongly on roadway context. On major roads, only protected bicycle facilities significantly increase route attractiveness, whereas on minor roads both protected and designated lanes provide utility gains. Sidewalks constitute the most frequently used riding environment, yet only asphalt-paved sidewalks are associated with positive utility, suggesting that sidewalk riding may reflect the absence of attractive on-street alternatives rather than a preference for pedestrian infrastructure. Tree coverage, particularly during summer, as well as building and wall coverage, are positively associated with route choice. Likelihood ratio tests and value-of-distance analysis indicate that roadway infrastructure exerts a stronger influence on route choice than visual streetscape features, although the latter provide additional explanatory power. These findings support targeted infrastructure investment and the integration of streetscape improvements as a complementary strategy for enhancing micromobility route attractiveness.

stat.AP

Skew-Parameterized Geometric Constants in Banach Spaces

Building on the family of geometric constants introduced by Amini-Harandi and Rahimi, we define a skew-parameterized family on real normed spaces by replacing the classical symmetric pair with a rotated coefficient pair. This modification reveals new extremal behavior, particularly for asymmetric homogeneous weight functions. We establish reduction formulas, comparison inequalities, and parameter-stability estimates. Under an appropriate differential balance condition, we determine the exact values of these constants on Hilbert spaces and obtain a converse characterization in dimensions at least three. We further derive sufficient conditions for uniform non-squareness and normal structure, together with explicit computations in classical normed spaces. These results provide a unified framework for detecting Hilbertian and fixed-point-related geometric properties of Banach spaces.

math.FA

Learning high-dimensional quantum entanglement through physics-guided neural networks

High-gain spontaneous parametric down-conversion (SPDC) produces bright squeezed vacuum with rich high-dimensional entanglement, but its output is inherently multimodal and non-perturbative, making the full modal characterization a major computational bottleneck. We propose a physics-guided deep neural network that reconstructs the source's modal fingerprint: the high-dimensional correlation signature across radial and azimuthal indices. We designed a FiLM-modulated convolutional architecture that predicts the joint (m,l) distribution, and training is driven by a hybrid loss that couples data-driven metrics (JSD, KL, MSE, Wasserstein) with a soft orbital-angular-momentum (OAM) conservation term, providing an essential inductive bias toward physically consistent solutions. Across gain regimes, our method achieves high-fidelity reconstruction with average JSD of 1.96e-3, WEMD of 1.54e-3, and KL divergence of 7.85e-3, delivering an approximate 128-fold speedup over full numerical simulation and more than 30% accuracy gains over U-Net baselines. These results demonstrate that physics-guided learning, via a soft OAM-conservation regularizer and physically generated training targets, enables rapid and data-efficient modal characterization. Compared with traditional numerical simulation, our mesh-free method has demonstrated good generalization with limited or contaminated training data and has enabled fast "online" prediction of the quantum dynamics of a high-dimensional entanglement system for real-world experimental implementation.

quant-ph

Sustainable and Efficient Renewable-Driven Energy Trading via Neural-Enhanced Time-Adaptive Robust Nash Bargaining between Hydrogen-Enriched Gas and Active Distribution Networks

Integrated hydrogen-enriched compressed natural gas (HCNG) and active distribution network (ADN) is providing efficient and sustainable flexibility for consuming renewable energies. Yet, cross-sector privacy and uncertain high-renewable scenarios block stable coordination. They also worsen decision performance and convergence. To conquer the barrier, a neural enhanced time-adaptive robust Nash bargaining strategy is proposed.In the first stage, to clear energy trading between ADN and gas distribution network (GDN) and promote its sustainability, a privacy preserved Nash Bargaining based on the alternating direction method of multipliers (ADMM) is applied. The next robust dispatch stage explores the worst renewable scenarios and derisks ADNs profit collapse from uncertainties. The convergence of the entire energy trading scheme is theoretically proved. As such, sustainable returns from the participation of solid oxide fuel cell (SOFC) and HCNG are facilitated. Finally, a time complexity and social welfare co-driven neural-enhanced time-adaptive strategy is proposed. The strategy assesses the influence of time resolution on social benefits and solving time in multi-energy trading. Based on the assessment, a neural network surrogate model is trained to accelerate the trading process in a close looped manner. Numerical assessment reveals that, the proposed strategy reaps a stable social welfare of nearly 1.6% to total cost, and benefit-steady situations for both ADN and GDN, even in the worst renewable scenarios. Moreover, it reduces runtime to 102.47s, improving computational efficiency by over 69.86% versus the fixed time-scale baseline, almost without sacrifice in economy.

eess.SY

BSoNet: Deep Learning Solution for Optimizing Image Quality of Portable Backscatter Imaging Systems

Portable backscatter imaging systems (PBI) integrate an X-ray source and detector in a single unit, utilizing Compton scattering photons to rapidly acquire superficial or shallow structural information of an inspected object through single-sided imaging. The application of this technology overcomes the limitations of traditional transmission X-ray detection, offering greater flexibility and portability, making it the preferred tool for the rapid and accurate identification of potential threats in scenarios such as borders, ports, and industrial nondestructive security inspections. However, the image quality is significantly compromised due to the limited number of Compton backscattered photons. The insufficient photon counts result primarily from photon absorption in materials, the pencil-beam scanning design, and short signal sampling times. It therefore yields severe image noise and an extremely low signal-to-noise ratio, greatly reducing the accuracy and reliability of PBI systems. To address these challenges, this paper introduces BSoNet, a novel deep learning-based approach specifically designed to optimize the image quality of PBI systems. The approach significantly enhances image clarity, recognition, and contrast while meeting practical application requirements. It transforms PBI systems into more effective and reliable inspection tools, contributing significantly to strengthening security protection.

eess.IV

Real-Time Forecasting of Pathological Gait via IMU Navigation: A Few-Shot and Generative Learning Framework for Wearable Devices

Current gait analysis faces challenges in various aspects, including limited and poorly labeled data within existing wearable electronics databases, difficulties in collecting patient data due to privacy concerns, and the inadequacy of the Zero-Velocity Update Technique (ZUPT) in accurately analyzing pathological gait patterns. To address these limitations, we introduce GaitMotion, a novel machine-learning framework that employs few-shot learning on a multitask dataset collected via wearable IMU sensors for real-time pathological gait analysis. GaitMotion enhances data quality through detailed, ground-truth-labeled sequences and achieves accurate step and stride segmentation and stride length estimation, which are essential for diagnosing neurological disorders. We incorporate a generative augmentation component, which synthesizes rare or underrepresented pathological gait patterns. GaitMotion achieves a 65\% increase in stride length estimation accuracy compared to ZUPT. In addition, its application to real patient datasets via transfer learning confirms its robust predictive capability. By integrating generative AI into wearable gait analysis, GaitMotion not only refines the precision of pathological gait forecasting but also demonstrates a scalable framework for leveraging synthetic data in biomechanical pattern recognition, paving the way for more personalized and data-efficient digital health services.

eess.SP

MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction

Accurate 3D semantic occupancy perception is essential for autonomous driving in complex environments with diverse and irregular objects. While vision-centric methods suffer from geometric inaccuracies, LiDAR-based approaches often lack rich semantic information. To address these limitations, MS-Occ, a novel multi-stage LiDAR-camera fusion framework which includes middle-stage fusion and late-stage fusion, is proposed, integrating LiDAR's geometric fidelity with camera-based semantic richness via hierarchical cross-modal fusion. The framework introduces innovations at two critical stages: (1) In the middle-stage feature fusion, the Gaussian-Geo module leverages Gaussian kernel rendering on sparse LiDAR depth maps to enhance 2D image features with dense geometric priors, and the Semantic-Aware module enriches LiDAR voxels with semantic context via deformable cross-attention; (2) In the late-stage voxel fusion, the Adaptive Fusion (AF) module dynamically balances voxel features across modalities, while the High Classification Confidence Voxel Fusion (HCCVF) module resolves semantic inconsistencies using self-attention-based refinement. Experiments on two large-scale benchmarks demonstrate state-of-the-art performance. On nuScenes-OpenOccupancy, MS-Occ achieves an Intersection over Union (IoU) of 32.1% and a mean IoU (mIoU) of 25.3%, surpassing the state-of-the-art by +0.7% IoU and +2.4% mIoU. Furthermore, on the SemanticKITTI benchmark, our method achieves a new state-of-the-art mIoU of 24.08%, robustly validating its generalization capabilities.Ablation studies further confirm the effectiveness of each individual module, highlighting substantial improvements in the perception of small objects and reinforcing the practical value of MS-Occ for safety-critical autonomous driving scenarios.

cs.CV

Semantic Parsing for Question Answering over Knowledge Graphs

In this paper, we propose a novel method for question answering over knowledge graphs based on graph-to-segment mapping, designed to improve the understanding of natural language questions. Our approach is grounded in semantic parsing, a key technique for interpreting question utterances. The main challenges arise from handling implicit entities and relations, as well as complex constraints such as temporal conditions, ordinality, and aggregation within the context of a knowledge graph. To address these issues, our framework integrates both rule-based and neural methods to parse and construct accurate, comprehensive semantic segment sequences. These sequences are then assembled into semantic query graphs, providing precise representations of question utterances. We formulate question semantic parsing as a sequence generation task, employing an encoder-decoder neural network to map natural language questions into semantic segments. Furthermore, to enhance the identification of implicit entities and relations, we incorporate a graph neural network that leverages knowledge graph context to enrich question representations. Experimental evaluations on two benchmark datasets demonstrate the effectiveness and superior performance of our model in semantic parsing for knowledge graph question answering.

cs.CL

Hybrid Deep Reconstruction for Vignetting-Free Upconversion Imaging through Scattering in ENZ Materials

Optical imaging through turbid or heterogeneous environments (collectively referred to as complex media) is fundamentally challenged by scattering, which scrambles structured spatial and phase information. To address this, we propose a hybrid-supervised deep learning framework to reconstruct high-fidelity images from nonlinear scattering measurements acquired with a time-gated epsilon-near-zero (ENZ) imaging system. The system leverages four-wave mixing (FWM) in subwavelength indium tin oxide (ITO) films to temporally isolate ballistic photons, thus rejecting multiply scattered light and enhancing contrast. To recover structured features from these signals, we introduce DeepTimeGate, a U-Net-based supervised model that performs initial reconstruction, followed by a Deep Image Prior (DIP) refinement stage using self-supervised learning. Our approach demonstrates strong performance across different imaging scenarios, including binary resolution patterns and complex vortex-phase masks, under varied scattering conditions. Compared to raw scattering inputs, it boosts average PSNR by 124%, SSIM by 231%, and achieves a 10 times improvement in intersection-over-union (IoU). Beyond enhancing fidelity, our method removes the vignetting effect and expands the effective field-of-view compared to the ENZ-based optical time gate output. These results suggest broad applicability in biomedical imaging, in-solution diagnostics, and other scenarios where conventional optical imaging fails due to scattering.

physics.optics

CT-Mamba: A Hybrid Convolutional State Space Model for Low-Dose CT Denoising

Low-dose CT (LDCT) significantly reduces the radiation dose received by patients, however, dose reduction introduces additional noise and artifacts. Currently, denoising methods based on convolutional neural networks (CNNs) face limitations in long-range modeling capabilities, while Transformer-based denoising methods, although capable of powerful long-range modeling, suffer from high computational complexity. Furthermore, the denoised images predicted by deep learning-based techniques inevitably exhibit differences in noise distribution compared to normal-dose CT (NDCT) images, which can also impact the final image quality and diagnostic outcomes. This paper proposes CT-Mamba, a hybrid convolutional State Space Model for LDCT image denoising. The model combines the local feature extraction advantages of CNNs with Mamba's strength in capturing long-range dependencies, enabling it to capture both local details and global context. Additionally, we introduce an innovative spatially coherent Z-shaped scanning scheme to ensure spatial continuity between adjacent pixels in the image. We design a Mamba-driven deep noise power spectrum (NPS) loss function to guide model training, ensuring that the noise texture of the denoised LDCT images closely resembles that of NDCT images, thereby enhancing overall image quality and diagnostic value. Experimental results have demonstrated that CT-Mamba performs excellently in reducing noise in LDCT images, enhancing detail preservation, and optimizing noise texture distribution, and exhibits higher statistical similarity with the radiomics features of NDCT images. The proposed CT-Mamba demonstrates outstanding performance in LDCT denoising and holds promise as a representative approach for applying the Mamba framework to LDCT denoising tasks.

eess.IV

Accurate Diagnosis of Respiratory Viruses Using an Explainable Machine Learning with Mid-Infrared Biomolecular Fingerprinting of Nasopharyngeal Secretions

Accurate identification of respiratory viruses (RVs) is critical for outbreak control and public health. This study presents a diagnostic system that combines Attenuated Total Reflectance Fourier Transform Infrared Spectroscopy (ATR-FTIR) from nasopharyngeal secretions with an explainable Rotary Position Embedding-Sparse Attention Transformer (RoPE-SAT) model to accurately identify multiple RVs within 10 minutes. Spectral data (4000-00 cm-1) were collected, and the bio-fingerprint region (1800-900 cm-1) was employed for analysis. Standard normal variate (SNV) normalization and second-order derivation were applied to reduce scattering and baseline drift. Gradient-weighted class activation mapping (Grad-CAM) was employed to generate saliency maps, highlighting spectral regions most relevant to classification and enhancing the interpretability of model outputs. Two independent cohorts from Beijing Youan Hospital, processed with different viral transport media (VTMs) and drying methods, were evaluated, with one including influenza B, SARS-CoV-2, and healthy controls, and the other including mycoplasma, SARS-CoV-2, and healthy controls. The model achieved sensitivity and specificity above 94.40% across both cohorts. By correlating model-selected infrared regions with known biomolecular signatures, we verified that the system effectively recognizes virus-specific spectral fingerprints, including lipids, Amide I, Amide II, Amide III, nucleic acids, and carbohydrates, and leverages their weighted contributions for accurate classification.

cs.LG

Graph-Driven Models for Gas Mixture Identification and Concentration Estimation on Heterogeneous Sensor Array Signals

Accurately identifying gas mixtures and estimating their concentrations are crucial across various industrial applications using gas sensor arrays. However, existing models face challenges in generalizing across heterogeneous datasets, which limits their scalability and practical applicability. To address this problem, this study develops two novel deep-learning models that integrate temporal graph structures for enhanced performance: a Graph-Enhanced Capsule Network (GraphCapsNet) employing dynamic routing for gas mixture classification and a Graph-Enhanced Attention Network (GraphANet) leveraging self-attention for concentration estimation. Both models were validated on datasets from the University of California, Irvine (UCI) Machine Learning Repository and a custom dataset, demonstrating superior performance in gas mixture identification and concentration estimation compared to recent models. In classification tasks, GraphCapsNet achieved over 98.00% accuracy across multiple datasets, while in concentration estimation, GraphANet attained an R2 score exceeding 0.96 across various gas components. Both GraphCapsNet and GraphANet exhibited significantly higher accuracy and stability, positioning them as promising solutions for scalable gas analysis in industrial settings.

cs.LG

On-Site Precise Screening of SARS-CoV-2 Systems Using a Channel-Wise Attention-Based PLS-1D-CNN Model with Limited Infrared Signatures

During the early stages of respiratory virus outbreaks, such as severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the efficient utilize of limited nasopharyngeal swabs for rapid and accurate screening is crucial for public health. In this study, we present a methodology that integrates attenuated total reflection-Fourier transform infrared spectroscopy (ATR-FTIR) with the adaptive iteratively reweighted penalized least squares (airPLS) preprocessing algorithm and a channel-wise attention-based partial least squares one-dimensional convolutional neural network (PLS-1D-CNN) model, enabling accurate screening of infected individuals within 10 minutes. Two cohorts of nasopharyngeal swab samples, comprising 126 and 112 samples from suspected SARS-CoV-2 Omicron variant cases, were collected at Beijing You'an Hospital for verification. Given that ATR-FTIR spectra are highly sensitive to variations in experimental conditions, which can affect their quality, we propose a biomolecular importance (BMI) evaluation method to assess signal quality across different conditions, validated by comparing BMI with PLS-GBM and PLS-RF results. For the ATR-FTIR signals in cohort 2, which exhibited a higher BMI, airPLS was utilized for signal preprocessing, followed by the application of the channel-wise attention-based PLS-1D-CNN model for screening. The experimental results demonstrate that our model outperforms recently reported methods in the field of respiratory virus spectrum detection, achieving a recognition screening accuracy of 96.48%, a sensitivity of 96.24%, a specificity of 97.14%, an F1-score of 96.12%, and an AUC of 0.99. It meets the World Health Organization (WHO) recommended criteria for an acceptable product: sensitivity of 95.00% or greater and specificity of 97.00% or greater for testing prior SARS-CoV-2 infection in moderate to high volume scenarios.

eess.SP

Unsupervised Attention-Based Multi-Source Domain Adaptation Framework for Drift Compensation in Electronic Nose Systems

Continuous, long-term monitoring of hazardous, noxious, explosive, and flammable gases in industrial environments using electronic nose (E-nose) systems faces the significant challenge of reduced gas identification accuracy due to time-varying drift in gas sensors. To address this issue, we propose a novel unsupervised attention-based multi-source domain shared-private feature fusion adaptation (AMDS-PFFA) framework for gas identification with drift compensation in E-nose systems. The AMDS-PFFA model effectively leverages labeled data from multiple source domains collected during the initial stage to accurately identify gases in unlabeled gas sensor array drift signals from the target domain. To validate the model's effectiveness, extensive experimental evaluations were conducted using both the University of California, Irvine (UCI) standard drift gas dataset, collected over 36 months, and drift signal data from our self-developed E-nose system, spanning 30 months. Compared to recent drift compensation methods, the AMDS-PFFA model achieves the highest average gas recognition accuracy with strong convergence, attaining 83.20% on the UCI dataset and 93.96% on data from our self-developed E-nose system across all target domain batches. These results demonstrate the superior performance of the AMDS-PFFA model in gas identification with drift compensation, significantly outperforming existing methods.

eess.SP

Intelligent Knee Sleeves: A Real-time Multimodal Dataset for 3D Lower Body Motion Estimation Using Smart Textile

The kinematics of human movements and locomotion are closely linked to the activation and contractions of muscles. To investigate this, we present a multimodal dataset with benchmarks collected using a novel pair of Intelligent Knee Sleeves (Texavie MarsWear Knee Sleeves) for human pose estimation. Our system utilizes synchronized datasets that comprise time-series data from the Knee Sleeves and the corresponding ground truth labels from the visualized motion capture camera system. We employ these to generate 3D human models solely based on the wearable data of individuals performing different activities. We demonstrate the effectiveness of this camera-free system and machine learning algorithms in the assessment of various movements and exercises, including extension to unseen exercises and individuals. The results show an average error of 7.21 degrees across all eight lower body joints when compared to the ground truth, indicating the effectiveness and reliability of the Knee Sleeve system for the prediction of different lower body joints beyond the knees. The results enable human pose estimation in a seamless manner without being limited by visual occlusion or the field of view of cameras. Our results show the potential of multimodal wearable sensing in a variety of applications from home fitness to sports, healthcare, and physical rehabilitation focusing on pose and movement estimation.

cs.CV

HSD-PAM: High Speed Super Resolution Deep Penetration Photoacoustic Microscopy Imaging Boosted by Dual Branch Fusion Network

Photoacoustic microscopy (PAM) is a novel implementation of photoacoustic imaging (PAI) for visualizing the 3D bio-structure, which is realized by raster scanning of the tissue. However, as three involved critical imaging parameters, imaging speed, lateral resolution, and penetration depth have mutual effect to one the other. The improvement of one parameter results in the degradation of other two parameters, which constrains the overall performance of the PAM system. Here, we propose to break these limitations by hardware and software co-design. Starting with low lateral resolution, low sampling rate AR-PAM imaging which possesses the deep penetration capability, we aim to enhance the lateral resolution and up sampling the images, so that high speed, super resolution, and deep penetration for the PAM system (HSD-PAM) can be achieved. Data-driven based algorithm is a promising approach to solve this issue, thereby a dedicated novel dual branch fusion network is proposed, which includes a high resolution branch and a high speed branch. Since the availability of switchable AR-OR-PAM imaging system, the corresponding low resolution, undersample AR-PAM and high resolution, full sampled OR-PAM image pairs are utilized for training the network. Extensive simulation and in vivo experiments have been conducted to validate the trained model, enhancement results have proved the proposed algorithm achieved the best perceptual and quantitative image quality. As a result, the imaging speed is increased 16 times and the imaging lateral resolution is improved 5 times, while the deep penetration merit of AR-PAM modality is still reserved.

eess.IV

Weight Bank Addition Photonic Accelerator for Artificial Intelligence

Neural networks powered by artificial intelligence play a pivotal role in current estimation and classification applications due to the escalating computational demands of evolving deep learning systems. The hindrances posed by existing computational limitations threaten to impede the further progression of these neural networks. In response to these issues, we propose neuromorphic networks founded on photonics that offer superior processing speed than electronic counterparts, thereby enhancing support for real time, three dimensional, and virtual reality applications. The weight bank, an integral component of these networks has a direct bearing on their overall performance. Our study demonstrates the implementation of a weight bank utilizing parallelly cascaded micro ring resonators. We present our observations on neuromorphic networks based on silicon on insulators, where cascaded MRRs play a crucial role in mitigating interchannel and intrachannel cross talk, a persistent issue in wavelength division multiplexing systems. Additionally, we design a standard silicon photonic accelerator to perform weight addition. Optimized to offer increased speed and reduced energy consumption, this photonic accelerator ensures comparable processing power to electronic devices.

physics.optics