SearcharxivSearch

arXiv subjects

Wenjun Xia

Publications and source records attributed to Wenjun Xia.

At least 19 recordsLinked to original sources

Front-end and Back-end Computational Modeling of 40-Hz Auditory Steady-State Response Abnormalities in Schizophrenia

40-Hz ASSR is reduced in schizophrenia, but it is unclear if this reflects altered auditory input or cortical E/I dynamics. We hypothesized that similar group differences could arise via distinct model mechanisms. EEG gamma% and ITPC from 21 HC and 21 SCZ constrained an auditory front-end coupled to a Wilson-Cowan E/I model. We compared front-end-restricted, back-end-restricted, and full-joint parameter searches, plus perturbation and fixed-point analyses. HC means exceeded SCZ for both metrics (not individually significant). All three models reproduced HC>SCZ but located group differences differently: front-end via input transformation, back-end via cortical dynamics, full-joint via both. The full-joint solution was most robust to perturbation. Fixed-point analysis revealed similar outputs with distinct local dynamics. All models reproduced the HC>SCZ pattern, suggesting schizophrenia pathophysiology may involve altered sensory encoding, altered cortical E/I, or both. This framework enables future patient-level mechanistic comparison and, after validation, may support individualized stratification.

q-bio.NC

Shared-Structure 4D Spectral Gaussian Representation for Sparse-View Spectral CT Reconstruction

Sparse-view spectral computed tomography (CT) reconstructs energy-resolved attenuation volumes from limited projection views, requiring simultaneous handling of angular undersampling and spectral coupling. We propose a SharedStructure 4D Spectral Gaussian Representation (4D-SG) that learns shared Gaussian geometry from full spectrum structural projections and uses a Gaussian-wise Spectral Density Curve Network (GSC-Net) to predict Gaussian raw density transformations. This factorization separates shared spatial structure from spectral attenuation variation, avoids independent channel geometry optimization, and establishes a continuous 4D-SG representation from discrete spectral measurements for unobserved spectral channel queries. Experiments on six synthesized, simulated projection, and real projection datasets with 50 views demonstrate the best average performance. Compared with the strongest Gaussian baseline, 4D-SG improves PSNR from 35.56 dB to 36.61 dB, increases SSIM from 0.909 to 0.914, and reduces LPIPS from 0.208 to 0.194, demonstrating its effectiveness for sparse-view spectral CT reconstruction.

cs.CV

FORCE-Interior: Measurement-Consistent Adaptation of a Poisson-Flow Generative Prior for Interior CT

Interior tomography reconstructs a region of interest (ROI) from truncated projections, an ill-posed problem with non-unique solutions and truncation-induced bias. Existing deep-learning methods can be sensitive to changes in ROI geometry and noise, while expressive generative priors may produce measurement-inconsistent content without measurement constraints. We propose FORCE-Interior, a training-free adaptation of a pretrained Poisson-flow generative prior to interior CT. A full-field-of-view (FOV) OS-SART warm start avoids forcing all measured attenuation into the ROI, and truncation-mask-aware OS-SART updates enforce data consistency throughout sampling. In our experiment, FORCE-Interior achieves the best PSNR, SSIM, and LPIPS at the two more severely truncated synthetic ROI sizes and competitive performance at the largest ROI, while maintaining low projection-domain residuals. These findings support the measurement-consistent adaptation of a reusable generative CT prior, while further clinical and patient-level validation remains necessary.

eess.IV

Real-time fall detection based on vision for low-power edge platforms

Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss event in a coupled dynamical system. We introduce a novel dual-LTC architecture comprising a Center-of-Mass (CoM) subsystem and a Base-of-Support (BoS) subsystem, both instantiated as Liquid Time-Constant (LTC) neural networks to continuously model inertial trajectory evolution and ground-contact adjustment through adaptive time constants, Physical interpretability of falling motion. A learnable coupling module emulates physical interaction between the two subsystems, while a Stability Manifold classifier operates in the joint latent space to detect boundary crossing via Lyapunov-inspired stability metrics. Complementary counterfactual trajectory projection and Time-to-Collision (TTC) estimation further enable irreversibility assessment and early warning. The architecture is designed to support a three-state prediction paradigm (Normal, Falling, Fallen); in this preliminary study, we validate the core stability discrimination capability on a two-class dataset (Normal vs. Falling), leaving the full three-state temporal transition to future work. Unlike conventional CNN--RNN pipelines, the proposed formulation encodes continuous-time mechanical inertia, yielding a sub-50K-parameter network capable of real-time inference on resource-constrained edge devices. Extensive experiments demonstrate competitive accuracy with superior physical interpretability, validating its efficacy for low-compute visual fall detection.

q-bio.NC

From Data Completeness to Data Sufficiency: A Task-Driven Imaging Framework for Intraoperative CBCT under Quality-Time-Dose Trade-offs

Mobile C-arm cone-beam computed tomography (CBCT) has been widely used for real-time intraoperative 3D imaging. However, current practice often mechanically applies the fan-beam CT criterion of "180{\deg} plus fan angle" in pursuit of "data completeness" in reconstruction. This review argues that, under the single circular trajectory of three-dimensional cone-beam geometry, complete data are mathematically unattainable; moreover, blindly increasing sampling may exacerbate the trade-off among intraoperative image quality (Q), imaging time (T), and radiation dose (D). Against this background, this review reframes the evaluation of intraoperative CBCT around "data sufficiency" rather than "data completeness." This perspective moves beyond the excessive pursuit of absolute mathematical and analytic accuracy, and instead emphasizes task-specific minimum image-quality thresholds required for clinical decision-making. By synthesizing evidence from multiple clinical scenarios, this review suggests that approximation errors can be acceptable when clinical decision-making requirements are satisfied, thereby achieving a Q-T-D balance.

eess.IV

Classification Accuracy of Minimal Spiking Neural Networks Follows a Log-Reciprocal Function

We investigate classification accuracy in minimal LIF-based spiking neural networks, examining its dependence on neuron count, stimulus nodes, and category number. Using an LLM to guide functional-form discovery, we compare power-law, exponential decay, and log-reciprocal candidates. The log-reciprocal model offers the strongest explanatory power: accuracy decays as 1/log(C), with neuron and stimulus effects marginal. This LLM-assisted approach efficiently identifies concise, interpretable descriptions, outperforming fixed-template methods. Our findings highlight AI's utility in computational neuroscience for uncovering interpretable relationships under resource constraints.

q-bio.NC

Single-Node Wilson--Cowan Model Accounts for Speech-Evoked $\gamma$-Band Deficits in Schizophrenia

Cortical gamma ($\gamma$)-band activity reflects local excitation-inhibition (E/I) balance. In schizophrenia (SCZ), reduced task-evoked gamma suggests altered E/I dynamics, but it is unclear whether differences stem from input properties or systematic shifts in E/I operating point and gain. We coupled a cochlear-inspired speech front end to a Wilson-Cowan E/I model to simulate gamma responses across three conditions: Healthy, SCZ-speech, and SCZ-semantics. Metrics included event-related spectral perturbation (ERSP$_\gamma$) and threshold-time fraction ($\gamma%$). A stable hierarchy emerged: Healthy(speech/semantics) $>$ SCZ(speech) $>$ SCZ(semantics), robust under equal-energy control and gain perturbations. Network dynamics coincided with single-node solutions, supporting interpretability. Pharmacological analogs showed bidirectional effects: reduced inhibition lowered $\gamma$, while reduced excitation increased $\gamma$, with no self-sustained oscillations. Findings indicate SCZ gamma deficits align more with shifts in E/I operating point and gain than input differences. This pipeline provides a testable, reusable mechanistic framework for speech-evoked gamma and a baseline for cross-population studies.

q-bio.NC

Audio Outperforms Text for Visual Decoding

Decoding visual semantic representations from human brain activity is a significant challenge. While recent zero-shot decoding approaches have improved performance by leveraging aligned image-text datasets, they overlook a fundamental aspect of human cognition: semantic understanding is inherently anchored in the auditory modality of speech, not text. To address this, our study introduces the first comparative framework for evaluating auditory versus textual semantic modalities in zero-shot visual neural decoding. We propose a novel brain-visual-auditory multimodal alignment model that directly utilizes auditory representations to encapsulate semantics, serving as a substitute for traditional textual descriptors. Our experimental results demonstrate that the auditory modality not only surpasses the textual modality in decoding accuracy but also achieves higher computational efficiency. These findings indicate that auditory semantic representations are more closely aligned with neural activity patterns during visual processing. This work reveals the critical and previously underestimated role of auditory semantics in decoding visual cognition and provides new insights for developing brain-computer interfaces that are more congruent with natural human cognitive mechanisms.

q-bio.NC

High-Resolution Magnetic Particle Imaging System Matrix Recovery Using a Vision Transformer with Residual Feature Network

This study presents a hybrid deep learning framework, the Vision Transformer with Residual Feature Network (VRF-Net), for recovering high-resolution system matrices in Magnetic Particle Imaging (MPI). MPI resolution often suffers from downsampling and coil sensitivity variations. VRF-Net addresses these challenges by combining transformer-based global attention with residual convolutional refinement, enabling recovery of both large-scale structures and fine details. To reflect realistic MPI conditions, the system matrix is degraded using a dual-stage downsampling strategy. Training employed paired-image super-resolution on the public Open MPI dataset and a simulated dataset incorporating variable coil sensitivity profiles. For system matrix recovery on the Open MPI dataset, VRF-Net achieved nRMSE = 0.403, pSNR = 39.08 dB, and SSIM = 0.835 at 2x scaling, and maintained strong performance even at challenging scale 8x (pSNR = 31.06 dB, SSIM = 0.717). For the simulated dataset, VRF-Net achieved nRMSE = 4.44, pSNR = 28.52 dB, and SSIM = 0.771 at 2x scaling, with stable performance at higher scales. On average, it reduced nRMSE by 88.2%, increased pSNR by 44.7%, and improved SSIM by 34.3% over interpolation and CNN-based methods. In image reconstruction of Open MPI phantoms, VRF-Net further reduced reconstruction error to nRMSE = 1.79 at 2x scaling, while preserving structural fidelity (pSNR = 41.58 dB, SSIM = 0.960), outperforming existing methods. These findings demonstrate that VRF-Net enables sharper, artifact-free system matrix recovery and robust image reconstruction across multiple scales, offering a promising direction for future in vivo applications.

physics.med-ph

Poisson Flow Consistency Training

The Poisson Flow Consistency Model (PFCM) is a consistency-style model based on the robust Poisson Flow Generative Model++ (PFGM++) which has achieved success in unconditional image generation and CT image denoising. Yet the PFCM can only be trained in distillation which limits the potential of the PFCM in many data modalities. The objective of this research was to create a method to train the PFCM in isolation called Poisson Flow Consistency Training (PFCT). The perturbation kernel was leveraged to remove the pretrained PFGM++, and the sinusoidal discretization schedule and Beta noise distribution were introduced in order to facilitate adaptability and improve sample quality. The model was applied to the task of low dose computed tomography image denoising and improved the low dose image in terms of LPIPS and SSIM. It also displayed similar denoising effectiveness as models like the Consistency Model. PFCT is established as a valid method of training the PFCM from its effectiveness in denoising CT images, showing potential with competitive results to other generative models. Further study is needed in the precise optimization of PFCT and in its applicability to other generative modeling tasks. The framework of PFCT creates more flexibility for the ways in which a PFCM can be created and can be applied to the field of generative modeling.

cs.CV

Tomographic Foundation Model -- FORCE: Flow-Oriented Reconstruction Conditioning Engine

Computed tomography (CT) is a major medical imaging modality. Clinical CT scenarios, such as low-dose screening, sparse-view scanning, and metal implants, often lead to severe noise and artifacts in reconstructed images, requiring improved reconstruction techniques. The introduction of deep learning has significantly advanced CT image reconstruction. However, obtaining paired training data remains rather challenging due to patient motion and other constraints. Although deep learning methods can still perform well with approximately paired data, they inherently carry the risk of hallucination due to data inconsistencies and model instability. In this paper, we integrate the data fidelity with the state-of-the-art generative AI model, referred to as the Poisson flow generative model (PFGM) with a generalized version PFGM++, and propose a novel CT framework: Flow-Oriented Reconstruction Conditioning Engine (FORCE). In our experiments, the proposed method shows superior performance in various CT imaging tasks, outperforming existing unsupervised reconstruction approaches.

eess.IV

Valley Polarization and Anomalous Valley Hall Effect in Altermagnet Ti2Se2S with Multipiezo Properties

Recently, altermagnets demonstrate numerous newfangle physical phenomena due to their inherent antiferromagnetic coupling and spontaneous spin splitting, that are anticipated to enable innovative spintronic devices. However, the rare two-dimensional altermagnets have been reported, making it difficult to meet the requirements for high-performance spintronic devices on account of the growth big data. Here, we predict a stable monolayer Ti2Se2S with out-of-plane altermagnetic ground state and giant valley splitting. The electronic properties of altermagnet Ti2Se2S are highly dependent on the onsite electron correlation. Through symmetry analysis, we find that the valleys of X and Y points are protected by the mirror Mxy symmetry rather than the time-reversal symmetry. Therefore, the multipiezo effect, including piezovalley and piezomagnetism, can be induced by the uniaxial strain. The total valley splitting of monolayer Ti2Se2S can be as high as ~500 meV. Most interestingly, the direction of valley polarization can be effectively tuned by the uniaxial strain, based on this, we have defined logical "0", "+1", and "-1" states for data transmission and storage. In addition, we have designed a schematic diagram for observing the anomalous Hall effect in experimentally. Our findings have enriched the candidate materials of two-dimensional altermagnet for the ultra-fast and low power consumption device applications.

cond-mat.mtrl-sci

Information-Maximized Soft Variable Discretization for Self-Supervised Image Representation Learning

Self-supervised learning (SSL) has emerged as a crucial technique in image processing, encoding, and understanding, especially for developing today's vision foundation models that utilize large-scale datasets without annotations to enhance various downstream tasks. This study introduces a novel SSL approach, Information-Maximized Soft Variable Discretization (IMSVD), for image representation learning. Specifically, IMSVD softly discretizes each variable in the latent space, enabling the estimation of their probability distributions over training batches and allowing the learning process to be directly guided by information measures. Motivated by the MultiView assumption, we propose an information-theoretic objective function to learn transform-invariant, non-travail, and redundancy-minimized representation features. We then derive a joint-cross entropy loss function for self-supervised image representation learning, which theoretically enjoys superiority over the existing methods in reducing feature redundancy. Notably, our non-contrastive IMSVD method statistically performs contrastive learning. Extensive experimental results demonstrate the effectiveness of IMSVD on various downstream tasks in terms of both accuracy and efficiency. Thanks to our variable discretization, the embedding features optimized by IMSVD offer unique explainability at the variable level. IMSVD has the potential to be adapted to other learning paradigms. Our code is publicly available at https://github.com/niuchuangnn/IMSVD.

cs.CV

Diffusion Prior Regularized Iterative Reconstruction for Low-dose CT

Computed tomography (CT) involves a patient's exposure to ionizing radiation. To reduce the radiation dose, we can either lower the X-ray photon count or down-sample projection views. However, either of the ways often compromises image quality. To address this challenge, here we introduce an iterative reconstruction algorithm regularized by a diffusion prior. Drawing on the exceptional imaging prowess of the denoising diffusion probabilistic model (DDPM), we merge it with a reconstruction procedure that prioritizes data fidelity. This fusion capitalizes on the merits of both techniques, delivering exceptional reconstruction results in an unsupervised framework. To further enhance the efficiency of the reconstruction process, we incorporate the Nesterov momentum acceleration technique. This enhancement facilitates superior diffusion sampling in fewer steps. As demonstrated in our experiments, our method offers a potential pathway to high-definition CT image reconstruction with minimized radiation.

eess.IV

Blind CT Image Quality Assessment Using DDPM-derived Content and Transformer-based Evaluator

Lowering radiation dose per view and utilizing sparse views per scan are two common CT scan modes, albeit often leading to distorted images characterized by noise and streak artifacts. Blind image quality assessment (BIQA) strives to evaluate perceptual quality in alignment with what radiologists perceive, which plays an important role in advancing low-dose CT reconstruction techniques. An intriguing direction involves developing BIQA methods that mimic the operational characteristic of the human visual system (HVS). The internal generative mechanism (IGM) theory reveals that the HVS actively deduces primary content to enhance comprehension. In this study, we introduce an innovative BIQA metric that emulates the active inference process of IGM. Initially, an active inference module, implemented as a denoising diffusion probabilistic model (DDPM), is constructed to anticipate the primary content. Then, the dissimilarity map is derived by assessing the interrelation between the distorted image and its primary content. Subsequently, the distorted image and dissimilarity map are combined into a multi-channel image, which is inputted into a transformer-based image quality evaluator. Remarkably, by exclusively utilizing this transformer-based quality evaluator, we won the second place in the MICCAI 2023 low-dose computed tomography perceptual image quality assessment grand challenge. Leveraging the DDPM-derived primary content, our approach further improves the performance on the challenge dataset.

eess.IV

Image Reconstruction Using a Mixture Score Function (MSF)

Computed tomography (CT) reconstructs volumetric images using X-ray projection data acquired from multiple angles around an object. For low-dose or sparse-view CT scans, the classic image reconstruction algorithms often produce severe noise and artifacts. To address this issue, we develop a novel iterative image reconstruction method based on maximum a posteriori (MAP) estimation. In the MAP framework, the score function, i.e., the gradient of the logarithmic probability density distribution, plays a crucial role as an image prior in the iterative image reconstruction process. By leveraging the Gaussian mixture model, we derive a novel score matching formula to establish an advanced score function (ADSF) through deep learning. Integrating the new ADSF into the image reconstruction process, a new ADSF iterative reconstruction method is developed to improve image reconstruction quality. The convergence of the ADSF iterative reconstruction algorithm is proven through mathematical analysis. The performance of the ADSF reconstruction method is also evaluated on both public medical image datasets and clinical raw CT datasets. Our results show that the ADSF reconstruction method can achieve better denoising and deblurring effects than the state-of-the-art reconstruction methods, showing excellent generalizability and stability.

physics.med-ph

Physics-/Model-Based and Data-Driven Methods for Low-Dose Computed Tomography: A survey

Since 2016, deep learning (DL) has advanced tomographic imaging with remarkable successes, especially in low-dose computed tomography (LDCT) imaging. Despite being driven by big data, the LDCT denoising and pure end-to-end reconstruction networks often suffer from the black box nature and major issues such as instabilities, which is a major barrier to apply deep learning methods in low-dose CT applications. An emerging trend is to integrate imaging physics and model into deep networks, enabling a hybridization of physics/model-based and data-driven elements. %This type of hybrid methods has become increasingly influential. In this paper, we systematically review the physics/model-based data-driven methods for LDCT, summarize the loss functions and training strategies, evaluate the performance of different methods, and discuss relevant issues and future directions.

eess.IV

Parallel Diffusion Model-based Sparse-view Cone-beam Breast CT

Breast cancer is the most prevalent cancer among women worldwide, and early detection is crucial for reducing its mortality rate and improving quality of life. Dedicated breast computed tomography (CT) scanners offer better image quality than mammography and tomosynthesis in general but at higher radiation dose. To enable breast CT for cancer screening, the challenge is to minimize the radiation dose without compromising image quality, according to the ALARA principle (as low as reasonably achievable). Over the past years, deep learning has shown remarkable successes in various tasks, including low-dose CT especially few-view CT. Currently, the diffusion model presents the state of the art for CT reconstruction. To develop the first diffusion model-based breast CT reconstruction method, here we report innovations to address the large memory requirement for breast cone-beam CT reconstruction and high computational cost of the diffusion model. Specifically, in this study we transform the cutting-edge Denoising Diffusion Probabilistic Model (DDPM) into a parallel framework for sub-volume-based sparse-view breast CT image reconstruction in projection and image domains. This novel approach involves the concurrent training of two distinct DDPM models dedicated to processing projection and image data synergistically in the dual domains. Our experimental findings reveal that this method delivers competitive reconstruction performance at half to one-third of the standard radiation doses. This advancement demonstrates an exciting potential of diffusion-type models for volumetric breast reconstruction at high-resolution with much-reduced radiation dose and as such hopefully redefines breast cancer screening and diagnosis.

eess.IV