SearcharxivSearch

arXiv subjects

Wenhao Dong

Publications and source records attributed to Wenhao Dong.

18 recordsLinked to original sources

LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection

The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN backbones built upon YOLO for feature extraction and focus on the design of feature fusion. In this paper, we introduce the Laplacian Decoupled Feature Enhancement block (LDFE) to fuse features from different stages of the dual-stream CNN backbone. By design, LDFE simultaneously considers the characteristics of modalities and structures for feature fusion by employing global-local decomposition, denoising, fusion, and reconstruction, sequentially. The LDFE first separates features into global and local components based on Laplacian Pyramid, and then performs denoising and fusion based on Global State Space Enhancement module (GS2E) and Local Convolutional Correlation Enhancement module (LC2E) separately. Specifically, the GS2E conducts a two-branch architecture for the main and auxiliary modalities. It dynamically suppresses noise in the main modality through cross-modal attention derived from the auxiliary modality, while employing a State Space Model to capture long-range dependencies within the global feature representations of the main modality. To obtain bidirectional interaction, the two modalities systematically alternate their main/auxiliary roles. Moreover, the LC2E suppresses noise in local features and leverages spatial and channel dimension along with triple convolution to extract fine-grained details for fusion. These innovative designs achieve a significant performance improvement, with mAP surpassing the SOTA methods 6.2%, 3.7%, 4.7%, 2.3%, 4.1% and 2.0% on M3FD, DroneVehicle, LLVIP, FLIR-Aligned, KAIST and VEDAI datasets,respectively.

cs.CV

Information Bottleneck-Guided Heterogeneous Graph Learning for Interpretable Neurodevelopmental Disorder Diagnosis

Developing interpretable models for neurodevelopmental disorders (NDDs) diagnosis presents significant challenges in effectively encoding, decoding, and integrating multimodal neuroimaging data. While many existing machine learning approaches have shown promise in brain network analysis, they typically suffer from limited interpretability, particularly in extracting meaningful biomarkers from functional magnetic resonance imaging (fMRI) data and establishing clear relationships between imaging features and demographic characteristics. Besides, current graph neural network methodologies face limitations in capturing both local and global functional connectivity patterns while simultaneously achieving theoretically principled multimodal data fusion. To address these challenges, we propose the Interpretable Information Bottleneck Heterogeneous Graph Neural Network (I2B-HGNN), a unified framework that applies information bottleneck principles to guide both brain connectivity modeling and cross-modal feature integration. This framework comprises two complementary components. The first is the Information Bottleneck Graph Transformer (IBGraphFormer), which combines transformer-based global attention mechanisms with graph neural networks through information bottleneck-guided pooling to identify sufficient biomarkers. The second is the Information Bottleneck Heterogeneous Graph Attention Network (IB-HGAN), which employs meta-path-based heterogeneous graph learning with structural consistency constraints to achieve interpretable fusion of neuroimaging and demographic data. The experimental results demonstrate that I2B-HGNN achieves superior performance in diagnosing NDDs, exhibiting both high classification accuracy and the ability to provide interpretable biomarker identification while effectively analyzing non-imaging data.

cs.CV

Measuring the crust-superfluid coupling time-scale for 105 UTMOST pulsars with a Kalman filter

Crust-superfluid coupling plays an important role in neutron star rotation, particularly with respect to timing noise and glitches. Here, we present new timing-noise-based estimates of the crust-superfluid coupling time-scale \(τ\) for 105 radio pulsars in the UTMOST dataset, by Kalman filtering the pulse times of arrival. The 105 objects are selected because they favor a two-component, crust-superfluid model over a one-component model with log Bayes factor \(\ln \mathfrak{B}_{\rm BF} \geq 5\). The median estimate of \(τ\) ranges from \(10^{4.6\pm0.4}\)\,s for PSR J2241$-$5236 to \(10^{7.7^{+0.7}_{-0.4}}\)\,s for PSR J1644$-$4559 among 28 out of 105 objects with sharply peaked \(τ\) posteriors. A hierarchical Bayesian analysis is performed on 101 out of 105 objects that are canonical (i.e.\ neither recycled nor magnetars) and reside in the populous core of the \(Ω_{\rm c}\)-\(\dotΩ_{\rm c}\) plane. It returns the population-level scaling \(τ\propto Ω_{\rm c}^{0.19^{+0.50}_{-0.52}} |\dotΩ_{\rm c}|^{0.18^{+0.18}_{-0.19}}\), where \(Ω_{\rm c}\) and \(\dotΩ_{\rm c}\) are the angular velocity and spin-down rate of the crust respectively. The variances of the stochastic crust and superfluid torques are also estimated hierarchically, with \(Q_{\rm c} \propto Ω_{\rm c}^{1.23^{+0.80}_{-0.75}} |\dotΩ_{\rm c}|^{0.49^{+0.27}_{-0.32}}\) and \(Q_{\rm s} \propto Ω_{\rm c}^{0.71^{+0.76}_{-0.78}} |\dotΩ_{\rm c}|^{1.27^{+0.30}_{-0.28}}\) respectively. Implications for the physical origin of crust-superfluid coupling, e.g.\ through mutual friction, are discussed briefly.

astro-ph.HE

STARFormer: A Novel Spatio-Temporal Aggregation Reorganization Transformer of FMRI for Brain Disorder Diagnosis

Many existing methods that use functional magnetic resonance imaging (fMRI) classify brain disorders, such as autism spectrum disorder (ASD) and attention deficit hyperactivity disorder (ADHD), often overlook the integration of spatial and temporal dependencies of the blood oxygen level-dependent (BOLD) signals, which may lead to inaccurate or imprecise classification results. To solve this problem, we propose a Spatio-Temporal Aggregation eorganization ransformer (STARFormer) that effectively captures both spatial and temporal features of BOLD signals by incorporating three key modules. The region of interest (ROI) spatial structure analysis module uses eigenvector centrality (EC) to reorganize brain regions based on effective connectivity, highlighting critical spatial relationships relevant to the brain disorder. The temporal feature reorganization module systematically segments the time series into equal-dimensional window tokens and captures multiscale features through variable window and cross-window attention. The spatio-temporal feature fusion module employs a parallel transformer architecture with dedicated temporal and spatial branches to extract integrated features. The proposed STARFormer has been rigorously evaluated on two publicly available datasets for the classification of ASD and ADHD. The experimental results confirm that the STARFormer achieves state-of-the-art performance across multiple evaluation metrics, providing a more accurate and reliable tool for the diagnosis of brain disorders and biomedical research. The codes are available at: https://github.com/NZWANG/STARFormer.

eess.IV

WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection

Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB and IR decomposed by Discrete Wavelet Transform (DWT). An improved detection head incorporating the Inverse Discrete Wavelet Transform (IDWT) is also proposed to reduce information loss and produce the final detection results. The core of our approach is the introduction of WaveMamba Fusion Block (WMFB), which facilitates comprehensive fusion across low-/high-frequency sub-bands. Within WMFB, the Low-frequency Mamba Fusion Block (LMFB), built upon the Mamba framework, first performs initial low-frequency feature fusion with channel swapping, followed by deep fusion with an advanced gated attention mechanism for enhanced integration. High-frequency features are enhanced using a strategy that applies an ``absolute maximum" fusion approach. These advancements lead to significant performance gains, with our method surpassing state-of-the-art approaches and achieving average mAP improvements of 4.5% on four benchmarks.

cs.CV

HRGS: Hierarchical Gaussian Splatting for Memory-Efficient High-Resolution 3D Reconstruction

3D Gaussian Splatting (3DGS) has made significant strides in real-time 3D scene reconstruction, but faces memory scalability issues in high-resolution scenarios. To address this, we propose Hierarchical Gaussian Splatting (HRGS), a memory-efficient framework with hierarchical block-level optimization. First, we generate a global, coarse Gaussian representation from low-resolution data. Then, we partition the scene into multiple blocks, refining each block with high-resolution data. The partitioning involves two steps: Gaussian partitioning, where irregular scenes are normalized into a bounded cubic space with a uniform grid for task distribution, and training data partitioning, where only relevant observations are retained for each block. By guiding block refinement with the coarse Gaussian prior, we ensure seamless Gaussian fusion across adjacent blocks. To reduce computational demands, we introduce Importance-Driven Gaussian Pruning (IDGP), which computes importance scores for each Gaussian and removes those with minimal contribution, speeding up convergence and reducing memory usage. Additionally, we incorporate normal priors from a pretrained model to enhance surface reconstruction quality. Our method enables high-quality, high-resolution 3D scene reconstruction even under memory constraints. Extensive experiments on three benchmarks show that HRGS achieves state-of-the-art performance in high-resolution novel view synthesis (NVS) and surface reconstruction tasks.

cs.CV

Neural-MCRL: Neural Multimodal Contrastive Representation Learning for EEG-based Visual Decoding

Decoding neural visual representations from electroencephalogram (EEG)-based brain activity is crucial for advancing brain-machine interfaces (BMI) and has transformative potential for neural sensory rehabilitation. While multimodal contrastive representation learning (MCRL) has shown promise in neural decoding, existing methods often overlook semantic consistency and completeness within modalities and lack effective semantic alignment across modalities. This limits their ability to capture the complex representations of visual neural responses. We propose Neural-MCRL, a novel framework that achieves multimodal alignment through semantic bridging and cross-attention mechanisms, while ensuring completeness within modalities and consistency across modalities. Our framework also features the Neural Encoder with Spectral-Temporal Adaptation (NESTA), a EEG encoder that adaptively captures spectral patterns and learns subject-specific transformations. Experimental results demonstrate significant improvements in visual decoding accuracy and model generalization compared to state-of-the-art methods, advancing the field of EEG-based neural visual representation decoding in BMI. Codes will be available at: https://github.com/NZWANG/Neural-MCRL.

cs.CV

A Tale of Single-channel Electroencephalogram: Devices, Datasets, Signal Processing, Applications, and Future Directions

Single-channel electroencephalogram (EEG) is a cost-effective, comfortable, and non-invasive method for monitoring brain activity, widely adopted by researchers, consumers, and clinicians. The increasing number and proportion of articles on single-channel EEG underscore its growing potential. This paper provides a comprehensive review of single-channel EEG, focusing on development trends, devices, datasets, signal processing methods, recent applications, and future directions. Definitions of bipolar and unipolar configurations in single-channel EEG are clarified to guide future advancements. Applications mainly span sleep staging, emotion recognition, educational research, and clinical diagnosis. Ongoing advancements of single-channel EEG in AI-based EEG generation techniques suggest potential parity or superiority over multichannel EEG performance.

cs.CV

State-space algorithm for detecting the nanohertz gravitational wave background

The stochastic gravitational wave background (SGWB) can be observed in the nanohertz band using a pulsar timing array (PTA). Here a computationally efficient state-space framework is developed for analysing SGWB data, in which the stochastic gravitational wave strain at Earth is tracked with a non-linear Kalman filter and separated simultaneously from intrinsic, achromatic pulsar spin wandering. The filter is combined with a nested sampler to estimate the parameters of the model, and to calculate a Bayes factor for selecting between models with and without a SGWB. The procedure extends previous state-space formulations of PTA data analysis applied to individually resolvable binary black hole sources. The performance of the new algorithm is tested on synthetic data from the first International PTA Mock Data Challenge. It is shown that the algorithm distinguishes a SGWB from pure noise for $A_{\rm gw} \geq 3 \times 10^{-14}$, where $A_{\rm gw}$ denotes the standard normalization factor for a power spectral density with power-law exponent $-13/3$. Additional, systematic validation tests are also performed with synthetic data generated independently by adjusting the injected parameters to cover astrophysically plausible ranges. Full posterior distributions are recovered and tested for accuracy. The state-space procedure is memory-light and evaluates the likelihood for a standard-sized PTA dataset in $\lesssim 10^{-1}$ s without optimization on a standard central processing unit.

astro-ph.IM

Gravitational waves from r-mode oscillations of stochastically accreting neutron stars

$r$-mode oscillations in rotating neutron stars are a source of continuous gravitational radiation. We investigate the excitation of $r$-modes by the mechanical impact on the neutron star surface of stochastically accreted clumps of matter, assuming that the Chandrasekhar-Friedman-Schutz instability is not triggered. The star is idealised as a slowly-rotating, unmagnetised, one-component fluid with a barotropic equation of state in Newtonian gravity. It is found that the $r$-mode amplitude depends weakly on the equation of state but sensitively on the rotation frequency $ν_{\rm s}$. The gravitational wave strain implicitly depends on the equation of state through the damping timescale. The root-mean-square strain is $h_{\rm rms} \approx 10^{-35} (ν_{\rm s}/ 10 {\rm Hz})^{2} (R_*/10 {\rm km})^2 (Δt_{\rm acc}/1 {\rm yr})^{1/2} (f_{\rm acc}/1 {\rm kHz})^{-1/2} (\dot{M}/10^{-8} \text{M}_{\odot} \text{yr}^{-1}) (v/0.4c) (d/1 {\rm kpc})^{-1}$, which is comparable to the strain from $g$-, $p$- and $f$-modes excited by stochastic accretion, where $R_*$ is the radius of the star, $Δt_{\rm acc}$ is the uninterrupted duration of an accretion episode, $f_{\rm acc}$ is the mean clump impact frequency, $\dot{M}$ is the accretion rate, $v$ is the impact speed, and $d$ is the distance of the star from the Earth. An observational test is proposed, based on the temporal autocorrelation function of the gravitational wave signal, to discern whether the Chandrasekhar-Friedman-Schutz instability switches on and coexists with impact-excited $r$-modes before or during a gravitational wave observation.

astro-ph.HE

MHNet: Multi-view High-order Network for Diagnosing Neurodevelopmental Disorders Using Resting-state fMRI

Background: Deep learning models have shown promise in diagnosing neurodevelopmental disorders (NDD) like ASD and ADHD. However, many models either use graph neural networks (GNN) to construct single-level brain functional networks (BFNs) or employ spatial convolution filtering for local information extraction from rs-fMRI data, often neglecting high-order features crucial for NDD classification. Methods: We introduce a Multi-view High-order Network (MHNet) to capture hierarchical and high-order features from multi-view BFNs derived from rs-fMRI data for NDD prediction. MHNet has two branches: the Euclidean Space Features Extraction (ESFE) module and the Non-Euclidean Space Features Extraction (Non-ESFE) module, followed by a Feature Fusion-based Classification (FFC) module for NDD identification. ESFE includes a Functional Connectivity Generation (FCG) module and a High-order Convolutional Neural Network (HCNN) module to extract local and high-order features from BFNs in Euclidean space. Non-ESFE comprises a Generic Internet-like Brain Hierarchical Network Generation (G-IBHN-G) module and a High-order Graph Neural Network (HGNN) module to capture topological and high-order features in non-Euclidean space. Results: Experiments on three public datasets show that MHNet outperforms state-of-the-art methods using both AAL1 and Brainnetome Atlas templates. Extensive ablation studies confirm the superiority of MHNet and the effectiveness of using multi-view fMRI information and high-order features. Our study also offers atlas options for constructing more sophisticated hierarchical networks and explains the association between key brain regions and NDD. Conclusion: MHNet leverages multi-view feature learning from both Euclidean and non-Euclidean spaces, incorporating high-order information from BFNs to enhance NDD classification performance.

cs.CV

State-space analysis of a continuous gravitational wave source with a pulsar timing array: inclusion of the pulsar terms

Pulsar timing arrays can detect continuous nanohertz gravitational waves emitted by individual supermassive black hole binaries. The data analysis procedure can be formulated within a time-domain, state-space framework, in which the radio timing observations are related to a temporal sequence of latent states, namely the intrinsic pulsar spin frequency. The achromatic wandering of the pulsar spin frequency is tracked using a Kalman filter concurrently with the pulse frequency modulation induced by a gravitational wave from a single source. The modulation is the sum of terms proportional to the gravitational wave strain at the Earth and at every pulsar in the array. Here we generalize previous state-space formulations of the pulsar timing array problem to include the pulsar terms; that is, we copy the pulsar terms from traditional, non-state-space analyses over to the state-space framework. The performance of the generalized Kalman filter is tested using astrophysically representative software injections in Gaussian measurement noise. It is shown that including the pulsar terms corrects for previously identified biases in the parameter estimates (especially the sky position of the source) which also arise in traditional matched-filter analyses that exclude the pulsar terms. Additionally, including the pulsar terms decreases the minimum detectable strain by $14\%$. Overall, the study verifies that the pulsar terms do not raise any special extra impediments for the state-space framework, beyond those studied in traditional analyses. The inspiral-driven evolution of the wave frequency at the Earth and at the retarded time at every pulsar in the array is also investigated.

astro-ph.HE

Kalman tracking and parameter estimation of continuous gravitational waves with a pulsar timing array

Continuous nanohertz gravitational waves from individual supermassive black hole binaries may be detectable with pulsar timing arrays. A novel search strategy is developed, wherein intrinsic achromatic spin wandering is tracked simultaneously with the modulation induced by a single gravitational wave source in the pulse times of arrival. A two-step inference procedure is applied within a state-space framework, such that the modulation is tracked with a Kalman filter, which then provides a likelihood for nested sampling. The procedure estimates the static parameters in the problem, such as the sky position of the source, without fitting for ensemble-averaged statistics such as the power spectral density of the timing noise, and therefore complements traditional parameter estimation methods. It also returns the Bayes factor relating a model with a single gravitational wave source to one without, complementing traditional detection methods. It is shown via astrophysically representative software injections in Gaussian measurement noise that the procedure distinguishes a gravitational wave from pure noise down to a characteristic wave strain of $h_0 \approx 2 \times 10^{-15}$. Full posterior distributions of model parameters are recovered and tested for accuracy. There is a bias of $\approx 0.3$ rad in the marginalised one-dimensional posterior for the orbital inclination $ι$, introduced by dropping the so-called `pulsar terms'. Smaller biases $\lesssim 10 \%$ are also observed in other static parameters.

astro-ph.HE

Gravitational waves from nonradial oscillations of stochastically accreting neutron stars

Oscillating neutron stars are sources of continuous gravitational waves. We study analytically the excitation of stellar oscillations by the mechanical impact on the stellar surface of ''clumps'' of stochastically accreted matter. We calculate the waveform and spectrum of the gravitational wave signal emitted by the accretion-driven pulsations. Results are generated for an idealised model of a nonrotating, unmagnetised, one-component star with uniform polytropic index $n_{\rm poly}$ assuming Newtonian gravity and the Cowling approximation. We find that the excited mode amplitudes grow with increasing $n_{\rm poly}$ and mode order $n$. The gravitational wave signal forms a sequence of amplitude-modulated packets for $n_{\rm poly}=1$, lasting $\sim 10^{-3}$s after each impact. The gravitational wave strain increases with increasing $n_{\rm poly}$, but decreases with increasing $n$ and increasing multipole order $l$ for $n_{\rm poly}=1$. In the observing band of current long-baseline interferometers, $g$-modes emit higher, narrower peaks in the amplitude spectral density than $f$- and $p$-modes, with the highest peaks reaching $\sim 10^{-26}$Hz$^{-1/2}$ for modes with damping time $τ_{nl} \sim 10^{8}$yr. The root-mean-square strain $h_{\text{rms}}$, calculated by summing over modes with $2\leq l\leq4$ and $τ_{nl} \leq 10^{8}$yr, spans the range $10^{-33} \leq h_{\text{rms}} \leq 10^{-32}$ for $1\leq n_{\text{poly}}\leq 2$.

astro-ph.HE

Fusion-Mamba for Cross-modality Object Detection

Cross-modality fusing complementary information from different modalities effectively improves object detection performance, making it more useful and robust for a wider range of applications. Existing fusion strategies combine different types of images or merge different backbone features through elaborated neural network modules. However, these methods neglect that modality disparities affect cross-modality fusion performance, as different modalities with different camera focal lengths, placements, and angles are hardly fused. In this paper, we investigate cross-modality fusion by associating cross-modal features in a hidden state space based on an improved Mamba with a gating mechanism. We design a Fusion-Mamba block (FMB) to map cross-modal features into a hidden state space for interaction, thereby reducing disparities between cross-modal features and enhancing the representation consistency of fused features. FMB contains two modules: the State Space Channel Swapping (SSCS) module facilitates shallow feature fusion, and the Dual State Space Fusion (DSSF) enables deep fusion in a hidden state space. Through extensive experiments on public datasets, our proposed approach outperforms the state-of-the-art methods on $m$AP with 5.9% on $M^3FD$ and 4.9% on FLIR-Aligned datasets, demonstrating superior object detection performance. To the best of our knowledge, this is the first work to explore the potential of Mamba for cross-modal fusion and establish a new baseline for cross-modality object detection.

cs.CV

Development of a resource-efficient FPGA-based neural network regression model for the ATLAS muon trigger upgrades

This paper reports on the development of a resource-efficient FPGA-based neural network regression model for potential applications in the future hardware muon trigger system of the ATLAS experiment at the Large Hadron Collider (LHC). Effective real-time selection of muon candidates is the cornerstone of the ATLAS physics programme. With the planned ATLAS upgrades for the High Luminosity LHC, an entirely new FPGA-based hardware muon trigger system will be installed that will process full muon detector data within a 10 $μs$ latency window. The large FPGA devices planned for this upgrade should have sufficient spare resources to allow deployment of machine learning methods for improving identification of muon candidates and searching for new exotic particles. Our neural network regression model promises to improve rejection of the dominant source of background trigger events in the central detector region, which are due to muon candidates with low transverse momenta. This model was implemented in FPGA using 157 digital signal processors and about 5,000 lookup tables. The simulated network latency and deadtime are 122 and 25 ns, respectively, when implemented in the FPGA device using a 320 MHz clock frequency. Two other FPGA implementations were also developed to study the impact of design choices on resource utilisation and latency. The performance parameters of our FPGA implementation are well within the requirements of the future muon trigger system, therefore opening a possibility for deploying machine learning methods for future data taking by the ATLAS experiment.

physics.ins-det

A low dead time, resource efficient encoding method for FPGA based high-resolution TDL TDCs

This paper presents a novel encoding method for fine time data of a tapped delay line (TDL) time-to-digital Converter (TDC). It is based on divide-and-conquer strategy, and has the advantage of significantly reducing logic resource utilization while retaining low dead-time performance. Furthermore, the problem of high bubble depth in advanced devices can be resolved with this method. Four examples are demonstrated, which were implemented in a Xilinx Artix-7 Field Programmable Gate Array (FPGA) device, and encoding method presented in this paper was employed to encode fine time data for normal TDL TDC, a half-length delay line TDC, and double-edge and four-edge wave union TDCs. Compared with TDCs from the latest published papers that adopt traditional encoders, the logic utilization of TDCs in this paper were reduced by a factor of 45% to 70% in different situations, while the encoding dead time can be restricted in one clock cycle. Acceptable resolutions of the demonstrated TDCs were also obtained, proving the functionality of the encoding method.

physics.space-ph

The DAQ and control system for JadePix3

The silicon pixel sensor is the core component of the vertex detector for the Circular Electron Positron Collider~(CEPC). The JadePix3 is a full-function large-size CMOS chip designed for the CEPC vertex detector. To test all the functions and the performance of this chip, we designed a test system based on the IPbus framework. The test system controls the parameters and monitors the status of the pixel chip. By integrating the jumbo frame feature into the IPbus suite, the block read/write speed is further extended in order to meet the specifications of the JadePix3. The robustness, scalability, and portability of this system have been verified by pulse test, cosmic test and laser test in the laboratory. This paper summarizes the DAQ and control system of the JadePix3 and presents the first results of the tests.

physics.ins-det