SearcharxivSearch

arXiv subjects

Peihong Zhang

Publications and source records attributed to Peihong Zhang.

At least 19 recordsLinked to original sources

Silicon-Germanium Heterostructures with Enhanced Valley Splitting for Spin Qubits

Achieving valley splittings well in excess of the thermal energy of electrons and avoiding valley excitations is essential for the consistent initialization, operation and readout of gate-defined Si spin qubits. In this work, we present a device-level optimization strategy for pushing valley splittings to between 1 and 5 meV, well beyond values reported in nearly all previous theoretical studies. Using device-scale simulations that incorporate atomistic alloy disorder through a 1D tight-binding theory, we demonstrate that our proposed approach yields large valley splittings with a tight distribution across disorder realizations, a key requirement for reproducible qubit performance at scale. The approach rests on an unorthodox Si/SiGe heterostructure design combining a narrow quantum well, a small Ge spike, and a pure-Ge cap. We corroborate these predictions with targeted atomistic density functional theory calculations. These results offer a clear path forward for scalable Si/SiGe spin qubit devices and, if realized experimentally, effectively eliminate valley splitting as an existential problem for large scale SiGe-based quantum processors.

cond-mat.mes-hall

Co-optimization of spin coherence and valley splitting in Si/SiGe heterostructures

Single electron spins can be used to encode and process information in semiconductor quantum devices. Progress has been hindered by materials challenges, such as the small energy splitting between low-lying valley states and hyperfine coupling to nuclear spins. Here we use density functional theory to optimize the valley splitting and spin dephasing time in realistic Si/SiGe heterostructures. Reductions in the Si quantum well width generally increase the valley splitting. However, in narrow quantum wells, a larger fraction of the electronic wavefunction resides in the SiGe buffer layers, which increases the hyperfine coupling with spinful $^{73}$Ge. Our work shows that Si/SiGe heterostructures with 3~--~4~nm wide quantum wells and $^{73}$Ge and $^{29}$Si concentrations of 50 ppm should support average valley splittings $E_{v}$~$>$~500~$μ$eV and spin dephasing times $T_2^*$ exceeding 15~$μ$s assuming an effective quantum dot area of 700 nm$^2$. In addition, sharper Si/SiGe interfaces in general result in larger valley splittings and longer spin dephasing times.

cond-mat.mtrl-sci

Membership Inference Attack Against Music Diffusion Models via Generative Manifold Perturbation

Membership inference attacks (MIAs) test whether a specific audio clip was used to train a model, making them a key tool for auditing generative music models for copyright compliance. However, loss-based signals (e.g., reconstruction error) are weakly aligned with human perception in practice, yielding poor separability at the low false-positive rates (FPRs) required for forensics. We propose the Latent Stability Adversarial Probe (LSA-Probe), a white-box method that measures a geometric property of the reverse diffusion: the minimal time-normalized perturbation budget needed to cross a fixed perceptual degradation threshold at an intermediate diffusion state. We show that training members, residing in more stable regions, exhibit a significantly higher degradation cost.

cs.SD

A New Workflow for Materials Discovery Bridging the Gap Between Experimental Databases and Graph Neural Networks

Incorporating Machine Learning (ML) into material property prediction has become a crucial step in accelerating materials discovery. A key challenge is the severe lack of training data, as many properties are too complicated to calculate with high-throughput first principles techniques. To address this, recent research has created experimental databases from information extracted from scientific literature. However, most existing experimental databases do not provide full atomic coordinate information, which prevents them from supporting advanced ML architectures such as Graph Neural Networks (GNNs). In this work, we propose to bridge this gap through an alignment process between experimental databases and Crystallographic Information Files (CIF) from the Inorganic Crystal Structure Database (ICSD). Our approach enables the creation of a database that can fully leverage state-of-the-art model architectures for material property prediction. It also opens the door to utilizing transfer learning to improve prediction accuracy. To validate our approach, we align NEMAD with the ICSD and compare models trained on the resulting database to those trained on NEMAD originally. We demonstrate significant improvements in both Mean Absolute Error (MAE) and Correct Classification Rate (CCR) in predicting the ordering temperatures and magnetic ground states of magnetic materials, respectively.

cond-mat.mtrl-sci

DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification under Domain Shift

Acoustic scene classification (ASC) suffers from device-induced domain shift, especially when labels are limited. Prior work focuses on curriculum-based training schedules that structure data presentation by ordering or reweighting training examples from easy-to-hard to facilitate learning; however, existing curricula are static, fixing the ordering or the weights before training and ignoring that example difficulty and marginal utility evolve with the learned representation. To overcome this limitation, we propose the Dynamic Dual-Signal Curriculum (DDSC), a training schedule that adapts the curriculum online by combining two signals computed each epoch: a domain-invariance signal and a learning-progress signal. A time-varying scheduler fuses these signals into per-example weights that prioritize domain-invariant examples in early epochs and progressively emphasize device-specific cases. DDSC is lightweight, architecture-agnostic, and introduces no additional inference overhead. Under the official DCASE 2024 Task~1 protocol, DDSC consistently improves cross-device performance across diverse ASC baselines and label budgets, with the largest gains on unseen-device splits.

cs.SD

TopSeg: A Multi-Scale Topological Framework for Data-Efficient Heart Sound Segmentation

Deep learning approaches for heart-sound (PCG) segmentation built on time-frequency features can be accurate but often rely on large expert-labeled datasets, limiting robustness and deployment. We present TopSeg, a topological representation-centric framework that encodes PCG dynamics with multi-scale topological features and decodes them using a lightweight temporal convolutional network (TCN) with an order- and duration-constrained inference step. To evaluate data efficiency and generalization, we train exclusively on PhysioNet 2016 dataset with subject-level subsampling and perform external validation on CirCor dataset. Under matched-capacity decoders, the topological features consistently outperform spectrogram and envelope inputs, with the largest margins at low data budgets; as a full system, TopSeg surpasses representative end-to-end baselines trained on their native inputs under the same budgets while remaining competitive at full data. Ablations at 10% training confirm that all scales contribute and that combining H_0 and H_1 yields more reliable S1/S2 localization and boundary stability. These results indicate that topology-aware representations provide a strong inductive bias for data-efficient, cross-dataset PCG segmentation, supporting practical use when labeled data are limited.

cs.SD

NMCSE: Noise-Robust Multi-Modal Coupling Signal Estimation Method via Optimal Transport for Cardiovascular Disease Detection

The coupling signal refers to a latent physiological signal that characterizes the transformation from cardiac electrical excitation, captured by the electrocardiogram (ECG), to mechanical contraction, recorded by the phonocardiogram (PCG). By encoding the temporal and functional interplay between electrophysiological and hemodynamic events, it serves as an intrinsic link between modalities and offers a unified representation of cardiac function, with strong potential to enhance multi-modal cardiovascular disease (CVD) detection. However, existing coupling signal estimation methods remain highly vulnerable to noise, particularly in real-world clinical and physiological settings, which undermines their robustness and limits practical value. In this study, we propose Noise-Robust Multi-Modal Coupling Signal Estimation (NMCSE), which reformulates coupling signal estimation as a distribution matching problem solved via optimal transport. By jointly aligning amplitude and timing, NMCSE avoids noise amplification and enables stable signal estimation. When integrated into a Temporal-Spatial Feature Extraction (TSFE) network, the estimated coupling signal effectively enhances multi-modal fusion for more accurate CVD detection. To evaluate robustness under real-world conditions, we design two complementary experiments targeting distinct sources of noise. The first uses the PhysioNet 2016 dataset with simulated hospital noise to assess the resilience of NMCSE to clinical interference. The second leverages the EPHNOGRAM dataset with motion-induced physiological noise to evaluate intra-state estimation stability across activity levels. Experimental results show that NMCSE consistently outperforms existing methods under both clinical and physiological noise, highlighting it as a noise-robust estimation approach that enables reliable multi-modal cardiac detection in real-world conditions.

eess.SP

Electron Localization in Non-Compact Covalent Bonds Captured by the r2SCAN+V Approach

In density functional theory, the SCAN (Strongly Constrained and Appropriately Normed) and r2SCAN functionals significantly improve over generalized gradient approximation functionals such as PBE (Perdew-Burke-Ernzerhof) in predicting electronic, magnetic, and structural properties across various materials, including transition-metal compounds. However, there remain puzzling cases where SCAN and r2SCAN underperform, such as in calculating the band structure of graphene, the magnetic moment of Fe, the potential energy curve of the Cr2 molecule, and the bond length of VO2. This research identifies a common characteristic among these challenging materials: non-compact covalent bonding through s-s, p-p, or d-d electron hybridization. While SCAN and r2SCAN excel at capturing electron localization at local atomic sites, they struggle to accurately describe electron localization in non-compact covalent bonds, resulting in a biased improvement. To address this issue, we propose the r2SCAN+V approach as a practical modification that improves accuracy across all the tested materials. The parameter V is 4 eV for metallic Fe, but substantially lower for the other cases. Our findings provide valuable insights for the future development of advanced functionals.

cond-mat.mtrl-sci

MAIA: An Inpainting-Based Approach for Music Adversarial Attacks

Music adversarial attacks have garnered significant interest in the field of Music Information Retrieval (MIR). In this paper, we present Music Adversarial Inpainting Attack (MAIA), a novel adversarial attack framework that supports both white-box and black-box attack scenarios. MAIA begins with an importance analysis to identify critical audio segments, which are then targeted for modification. Utilizing generative inpainting models, these segments are reconstructed with guidance from the output of the attacked model, ensuring subtle and effective adversarial perturbations. We evaluate MAIA on multiple MIR tasks, demonstrating high attack success rates in both white-box and black-box settings while maintaining minimal perceptual distortion. Additionally, subjective listening tests confirm the high audio fidelity of the adversarial samples. Our findings highlight vulnerabilities in current MIR systems and emphasize the need for more robust and secure models.

cs.SD

Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack

Music Information Retrieval (MIR) systems are highly vulnerable to adversarial attacks that are often imperceptible to humans, primarily due to a misalignment between model feature spaces and human auditory perception. Existing defenses and perceptual metrics frequently fail to adequately capture these auditory nuances, a limitation supported by our initial listening tests showing low correlation between common metrics and human judgments. To bridge this gap, we introduce Perceptually-Aligned MERT Transformer (PAMT), a novel framework for learning robust, perceptually-aligned music representations. Our core innovation lies in the psychoacoustically-conditioned sequential contrastive transformer, a lightweight projection head built atop a frozen MERT encoder. PAMT achieves a Spearman correlation coefficient of 0.65 with subjective scores, outperforming existing perceptual metrics. Our approach also achieves an average of 9.15\% improvement in robust accuracy on challenging MIR tasks, including Cover Song Identification and Music Genre Classification, under diverse perceptual adversarial attacks. This work pioneers architecturally-integrated psychoacoustic conditioning, yielding representations significantly more aligned with human perception and robust against music adversarial attacks.

cs.SD

Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds

Self-supervised speech models have demonstrated impressive performance in speech processing, but their effectiveness on non-speech data remains underexplored. We study the transfer learning capabilities of such models on bioacoustic detection and classification tasks. We show that models such as HuBERT, WavLM, and XEUS can generate rich latent representations of animal sounds across taxa. We analyze the models properties with linear probing on time-averaged representations. We then extend the approach to account for the effect of time-wise information with other downstream architectures. Finally, we study the implication of frequency range and noise on performance. Notably, our results are competitive with fine-tuned bioacoustic pre-trained models and show the impact of noise-robust pre-training setups. These findings highlight the potential of speech-based self-supervised learning as an efficient framework for advancing bioacoustic research.

cs.LG

Out-of-plane displacement of quantum color centers in monolayer h-BN

Color centers exhibiting deep-level states within the wide bandgap h-BN monolayer possess substantial potential for quantum applications. Uncovering precise geometric characteristics at the atomic scale is crucial for understanding defect performance. In this study, first-principles calculations were performed on the most extensively investigated CBVN and NBVN color centers in h-BN, focusing on the out-of-plane displacement and their specific impacts on electronic, vibrational, and emission properties. We demonstrate the competition between the σ*-like antibonding state and the π-like bonding state, which determines the out-of-plane displacement. The overall effect of vibronic coupling on geometry is elucidated using a pseudo Jahn-Teller model. Local vibrational analysis reveals a series of distinct quasi-local phonon modes that could serve as fingerprints for experimental identification of specific point defects. The critical effects of out-of-plane displacement during the quantum emission process are carefully elucidated to answer the distinct observations in experiments, and these revelations are universal in quantum point defects in other layered materials.

cond-mat.mes-hall

Efficiently charting the space of mixed vacancy-ordered perovskites by machine-learning encoded atomic-site information

Vacancy-ordered double perovskites (VODPs) are promising alternatives to three-dimensional lead halide perovskites for optoelectronic and photovoltaic applications. Mixing these materials creates a vast compositional space, allowing for highly tunable electronic and optical properties. However, the extensive chemical landscape poses significant challenges in efficiently screening candidates with target properties. In this study, we illustrate the diversity of electronic and optical characteristics as well as the nonlinear mixing effects on electronic structures within mixed VODPs. For mixed systems with limited local environment options, the information regarding atomic-site occupation in-principle determines both structural configurations and all essential properties. Building upon this concept, we have developed a model that integrates a data-augmentation scheme with a transformer-inspired graph neural network (GNN), which encodes atomic-site information from mixed systems. This approach enables us to accurately predict band gaps and formation energies for test samples, achieving Root Mean Square Errors (RMSE) of 21 meV and 3.9 meV/atom, respectively. Trained with datasets that include (up to) ternary mixed systems and supercells with less than 72 atoms, our model can be generalized to medium- and high-entropy mixed VODPs (with 4 to 6 principal mixing elements) and large supercells containing more than 200 atoms. Furthermore, our model successfully reproduces experimentally observed bandgap bowing in Sn-based mixed VODPs and reveals an unconventional mixing effect that can result in smaller band gaps compared to those found in pristine systems.

cond-mat.mtrl-sci

TF-SepNet: An Efficient 1D Kernel Design in CNNs for Low-Complexity Acoustic Scene Classification

Recent studies focus on developing efficient systems for acoustic scene classification (ASC) using convolutional neural networks (CNNs), which typically consist of consecutive kernels. This paper highlights the benefits of using separate kernels as a more powerful and efficient design approach in ASC tasks. Inspired by the time-frequency nature of audio signals, we propose TF-SepNet, a CNN architecture that separates the feature processing along the time and frequency dimensions. Features resulted from the separate paths are then merged by channels and directly forwarded to the classifier. Instead of the conventional two dimensional (2D) kernel, TF-SepNet incorporates one dimensional (1D) kernels to reduce the computational costs. Experiments have been conducted using the TAU Urban Acoustic Scene 2022 Mobile development dataset. The results show that TF-SepNet outperforms similar state-of-the-arts that use consecutive kernels. A further investigation reveals that the separate kernels lead to a larger effective receptive field (ERF), which enables TF-SepNet to capture more time-frequency features.

cs.SD

Quasiparticle and Excitonic Structures of Few-layer and Bulk GaSe: Interlayer Coupling, Self-energy, and Electron-hole Interaction

Metal monochalcogenide GaSe is a classic layered semiconductor that has received increasing research interest due to its highly tunable electronic and optical properties for ultrathin electronics applications. Despite intense research efforts, a systematic understanding of the layer-dependent electronic and optical properties of GaSe remains to be established, and there appear significant discrepancies between different experiments. We have performed GW plus Bethe-Salpeter equation (BSE) calculations for few-layer and bulk GaSe, aiming at understanding the effects of interlayer coupling and dielectric screening on excited state properties of GaSe, and how the electronic and optical properties evolve from strongly two-dimensional (2D) like to intermediate thick layers, and to three-dimensional (3D) bulk character. Using a new definition of the exciton binding energy, we are able to calculate the binding energies of all excitonic states. Our results reveal an interesting correlation between the binding energy of an exciton and the spread of its wave function in the real and momentum spaces. We find that the existence of (nearly) parallel valence and conduction bands facilitates the formation of excitonic states that spread out in the momentum space. Thus, these excitons tend to be more localized in real space and have large exciton binding energies. The interlayer coupling substantially suppresses the Mexican-hat-like dispersion of the top valence band seen in monolayer system, explaining the greatly enhanced photoluminescence (PL) as layer thickness increases. Our results also help resolve apparent discrepancies between different experiments. After including the quasiparticle and excitonic effects as well the optical activities of excitons, our results compare well with available experimental results.

cond-mat.mtrl-sci

Stacking up electron-rich and electron-deficient monolayers to achieve extraordinary mid- to far-infrared excitonic absorption: Interlayer excitons in the C3B/C3N bilayer

Our ability to efficiently detect and generate far-infrared (i.e., terahertz) radiation is vital in areas spanning from biomedical imaging to interstellar spectroscopy. Despite decades of intense research, bridging the terahertz gap between electronics and optics remains a major challenge due to the lack of robust materials that can efficiently operate in this frequency range, and two-dimensional (2D) type-II heterostructures may be ideal candidates to fill this gap. Herein, using highly accurate many-body perturbation theory within the GW plus Bethe-Salpeter equation approach, we predict that a type-II heterostructure consisting of an electron rich C3N and an electron deficient C3B monolayers can give rise to extraordinary optical activities in the mid- to far-infrared range. C3N and C3B are two graphene-derived 2D materials that have attracted increasing research attention. Although both C3N and C3B monolayers are moderate gap 2D materials, and they only couple through the rather weak van der Waals interactions, the bilayer heterostructure surprisingly supports extremely bright, low-energy interlayer excitons with large binding energies of 0.2 ~ 0.4 eV, offering an ideal material with interlayer excitonic states for mid-to far-infrared applications at room temperature. We also investigate in detail the properties and formation mechanism of the inter- and intra-layer excitons.

cond-mat.mtrl-sci

Giant excitonic effects in bulk vacancy-ordered double perovskites

Using first-principles GW plus Bethe-Salpeter equation calculations, we identify anomalously strong excitonic effects in several vacancy-ordered double perovskites Cs2MX6 (M = Ti, Zr; X = I, Br). Giant exciton binding energies about 1 eV are found in these moderate-gap, inorganic bulk semiconductors, pushing the limit of our understanding of electron-hole (e-h) interaction and exciton formation in solids. Not only are the exciton binding energies extremely large compared with any other moderate-gap bulk semiconductors, but they are also larger than typical 2D semiconductors with comparable quasiparticle gaps. Our calculated lowest bright exciton energy agrees well with the experimental optical band gap. The low-energy excitons closely resemble the Frenkel excitons in molecular crystals, as they are highly localized in a single [MX6]2- octahedron and extended in the reciprocal space. The weak dielectric screening effects and the nearly flat frontier electronic bands, which are derived from the weakly bonded [MX6]2- units, together explain the significant excitonic effects. Spin-orbit coupling effects play a crucial role in red-shifting the lowest bright exciton by mixing up spin-singlet and spin-triplet excitons, while exciton-phonon coupling effects have minor impacts on the strong exciton binding energies.

physics.comp-ph

Prediction of protected band edge states and dielectric tunable quasiparticle and excitonic properties of monolayer MoSi$_2$N$_4$

The electronic structure of two-dimensional (2D) materials are inherently prone to environmental perturbations, which may pose significant challenges to their applications in electronic or optoelectronic devices. A 2D material couples with its environment through two mechanisms: local chemical coupling and nonlocal dielectric screening effects. The local chemical coupling is often difficult to predict or control experimentally. Nonlocal dielectric screening, on the other hand, can be tuned by choosing the substrates or layer thickness in a controllable manner. Therefore, a compelling 2D electronic material should offer band edge states that are robust against local chemical coupling effects. Here it is demonstrated that the recently synthesized MoSi$_2$N$_4$ is an ideal 2D semiconductor with robust band edge states protected from capricious environmental chemical coupling effects. Detailed many-body perturbation theory calculations are carried out to illustrate how the band edge states of MoSi$_2$N$_4$ are shielded from the direct chemical coupling effects, but its quasiparticle and excitonic properties can be modulated through the nonlocal dielectric screening effects. This unique property, together with the moderate band gap and the thermodynamic and mechanical stability of this material, paves the way for a range of applications of MoSi$_2$N$_4$ in areas including energy, 2D electronics, and optoelectronics.

cond-mat.mtrl-sci