SearcharxivSearch

arXiv subjects

Panpan Zhang

Publications and source records attributed to Panpan Zhang.

At least 19 recordsLinked to original sources

Bayesian Variable Selection for High-Dimensional Predictors with Missing Psychometric Outcomes

High-dimensional, multimodal predictors and partially observed multivariate outcomes are common in psychometric research. However, existing regularization methods often do not accommodate hierarchical predictor structures and are primarily designed for univariate outcomes. We propose SHIM, a Bayesian framework for structured variable selection that combines hierarchical horseshoe shrinkage with a Bayesian treatment of missing outcomes. The framework jointly accommodates predictor hierarchies, dependence among outcomes, and incomplete multivariate responses. We establish theoretical properties of the proposed prior specification and evaluate SHIM through simulation studies. The results demonstrate that SHIM balances sensitivity with false-positive control while yielding accurate coefficient estimates and well-calibrated uncertainty quantification. We further apply SHIM to data from an Alzheimer's disease cohort to characterize associations between multimodal neuroimaging measures and multivariate neuropsychological outcomes and to generate posterior-based multiple imputations for downstream analyses of the relationships between fluid biomarkers and cognition. An R package, shim, is publicly available to facilitate implementation.

stat.ME

Gaussian Graphical Models for Functional Connectivity Analysis: A Statistical Review with Applications to Alzheimer's Disease

Functional connectivity analysis is an important tool for characterizing interactions among brain regions, particularly in studies of neurodegenerative disorders such as Alzheimer's disease (AD). Gaussian graphical models (GGMs) provide a promising statistical framework for estimating functional connectivity by capturing conditional dependence relationships among brain regions. Although a variety of regularized precision matrix estimators have been proposed to estimate sparse conditional dependency structures for GGMs, their comparative performance and practical implications for neuroimaging studies are not well understood. In this work, we present a comprehensive statistical review and empirical evaluation of widely used GGM estimation methods, including the graphical lasso (glasso), ridge-based glasso, graphical elastic net, adaptive glasso, smoothly clipped absolute deviation (SCAD), minimax concave penalty (MCP), constrained $\ell_1$ minimization for inverse matrix estimation (CLIME), and tuning-insensitive graph estimation and regression (TIGER). Their performance is evaluated through extensive data-driven simulations designed to reflect realistic neuroimaging settings, along with an application to an AD cohort study to illustrate methodological differences and their impact on downstream network analysis. In addition, a user-friendly R package, spice, is provided to facilitate implementation and enhance the reproducibility of empirical studies.

stat.ME

Quantum Dynamics of Enantiomers in Chiral Optical Cavities

Chirality, the absence of mirror symmetry, is a fundamental molecular property with far-reaching consequences from chemistry to biology. Yet enantiosensitive optical responses are very weak. Here, we introduce a theoretical framework in which a chiral optical cavity under strong coupling directly lifts the degeneracy of opposite enantiomers at the electronic-dipole level. The cavity's parity-breaking field inside the cavity induces distinct site-energy shifts for left- versus right-handed molecules, producing robust enantioselective polariton states that overcome the weakness of traditional chiroptical effects. Using cavity quantum electrodynamics simulations, we show that strong light-matter coupling reshapes the polaritonic energy landscape and leads to enantiomer-specific coherence lifetimes and relaxation pathways. To reveal these dynamics, we propose ultrafast two-dimensional electronic spectroscopy (2DES) as a probe, capable of resolving polaritonic splittings on femtosecond timescales. Simulated 2DES spectra exhibit unambiguous enantioselective signatures of the cavity-induced asymmetry. These findings establish that chiral cavities provide a powerful platform for detecting and controlling molecular handedness beyond the limits of conventional optical methods.

physics.optics

When Does the Silhouette Score Work? A Comprehensive Study in Network Clustering

Selecting the number of communities is a fundamental challenge in network clustering. The silhouette score offers an intuitive, model-free criterion that balances within-cluster cohesion and between-cluster separation. Albeit its widespread use in clustering analysis, its performance in network-based community detection remains insufficiently characterized. In this study, we comprehensively evaluate the performance of the silhouette score across unweighted, weighted, and fully connected networks, examining how network size, separation strength, and community size imbalance influence its performance. Simulation studies show that the silhouette score accurately identifies the true number of communities when clusters are well separated and balanced, but it tends to underestimate under strong imbalance or weak separation and to overestimate in sparse networks. Extending the evaluation to a real airline reachability network, we demonstrate that the silhouette-based clustering can recover geographically interpretable and market-oriented clusters. These findings provide empirical guidance for applying the silhouette score in network clustering and clarify the conditions under which its use is most reliable.

cs.SI

Controlling Nonadiabatic Transitions Through Engineered Ultrafast Laser Fields at Conical Intersections

In this paper, we investigate coherent control of nonadiabatic dynamics at a conical intersection (CI) using engineered ultrafast laser pulses. Within a model vibronic system, we tailor pulse chirp and temporal profile and compute the resulting wave-packet population and coherence dynamics using projections along the reaction coordinate. This approach allows us to resolve the detailed evolution of wave-packets as they traverse the degeneracy region with strong nonadiabatic coupling. By systematically varying pulse parameters, we demonstrate that both chirp and pulse duration modulate vibrational coherence and alter branching between competing pathways, leading to controlled changes in quantum yield. Our results elucidate the dynamical mechanisms underlying pulse-shaped control near conical intersections and establish a general framework for manipulating ultrafast nonadiabatic processes.

quant-ph

Covariate Connectivity Combined Clustering for Weighted Networks

Community detection is a central task in network analysis, with applications in social, biological, and technological systems. Traditional algorithms rely primarily on network topology, which can fail when community signals are partly encoded in node-specific attributes. Existing covariate-assisted methods often assume the number of clusters is known, involve computationally intensive inference, or are not designed for weighted networks. We propose $\text{C}^4$: Covariate Connectivity Combined Clustering, an adaptive spectral clustering algorithm that integrates network connectivity and node-level covariates into a unified similarity representation. $\text{C}^4$ balances the two sources of information through a data-driven tuning parameter, estimates the number of communities via an eigengap heuristic, and avoids reliance on costly sampling-based procedures. Simulation studies show that $\text{C}^4$ achieves higher accuracy and robustness than competing approaches across diverse scenarios. Application to an airport reachability network demonstrates the method's scalability, interpretability, and practical utility for real-world weighted networks.

stat.ME

ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows

As large language models (LLMs) advance, the ultimate vision for their role in science is emerging: we could build an AI collaborator to effectively assist human beings throughout the entire scientific research process. We refer to this envisioned system as ResearchGPT. Given that scientific research progresses through multiple interdependent phases, achieving this vision requires rigorous benchmarks that evaluate the end-to-end workflow rather than isolated sub-tasks. To this end, we contribute CS-54k, a high-quality corpus of scientific Q&A pairs in computer science, built from 14k CC-licensed papers. It is constructed through a scalable, paper-grounded pipeline that combines retrieval-augmented generation (RAG) with multi-stage quality control to ensure factual grounding. From this unified corpus, we derive two complementary subsets: CS-4k, a carefully curated benchmark for evaluating AI's ability to assist scientific research, and CS-50k, a large-scale training dataset. Extensive experiments demonstrate that CS-4k stratifies state-of-the-art LLMs into distinct capability tiers. Open models trained on CS-50k with supervised training and reinforcement learning demonstrate substantial improvements. Even 7B-scale models, when properly trained, outperform many larger proprietary systems, such as GPT-4.1, GPT-4o, and Gemini 2.5 Pro. This indicates that making AI models better research assistants relies more on domain-aligned training with high-quality data than on pretraining scale or general benchmark performance. We release CS-4k and CS-50k in the hope of fostering AI systems as reliable collaborators in CS research.

cs.LG

Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge

State-space models (SSMs) have emerged as efficient alternatives to Transformers for sequence modeling, offering superior scalability through recurrent structures. However, their training remains costly and the ecosystem around them is far less mature than that of Transformers. Moreover, the structural heterogeneity between SSMs and Transformers makes it challenging to efficiently distill knowledge from pretrained attention models. In this work, we propose Cross-architecture distillation via Attention Bridge (CAB), a novel data-efficient distillation framework that efficiently transfers attention knowledge from Transformer teachers to state-space student models. Unlike conventional knowledge distillation that transfers knowledge only at the output level, CAB enables token-level supervision via a lightweight bridge and flexible layer-wise alignment, improving both efficiency and transferability. We further introduce flexible layer-wise alignment strategies to accommodate architectural discrepancies between teacher and student. Extensive experiments across vision and language domains demonstrate that our method consistently improves the performance of state-space models, even under limited training data, outperforming both standard and cross-architecture distillation methods. Our findings suggest that attention-based knowledge can be efficiently transferred to recurrent models, enabling rapid utilization of Transformer expertise for building a stronger SSM community.

cs.LG

CALF-SBM: A Covariate-Assisted Latent Factor Stochastic Block Model

We propose a novel network generative model extended from the standard stochastic block model by concurrently utilizing observed node-level information and accounting for network-enabled nodal heterogeneity. The proposed model is so so-called covariate-assisted latent factor stochastic block model (CALF-SBM). The inference for the proposed model is done in a fully Bayesian framework. The primary application of CALF-SBM in the present research is focused on community detection, where a model-selection-based approach is employed to estimate the number of communities which is practically assumed unknown. To assess the performance of CALF-SBM, an extensive simulation study is carried out, including comparisons with multiple classical and modern network clustering algorithms. Lastly, the paper presents two real data applications, respectively based on an extremely new network data demonstrating collaborative relationships of otolaryngologists in the United States and a traditional aviation network data containing information about direct flights between airports in the United States and Canada.

stat.ME

Effects of alternating interactions and boundary conditions on quantum entanglement of three-leg Heisenberg ladder

The spin-12 three-leg antiferromagnetic Heisenberg spin ladder is studied under open boundary condition (OBC) and cylinder boundary condition (CBC), using the density matrix renormalization group and matrix product state methods, respectively. Specifically, we calculate the energy density, entanglement entropy, and concurrence while discussing the effects of interleg interaction J2 and the alternating coupling parameter gamma on these quantities. It is found that the introduction of gamma can completely reverse the concurrence distribution between odd and even bonds. Under CBC, the generation of the interleg concurrence is inhibited when gamma=0, and the introduction of gamma can cause interleg concurrence between chains 1 and 3, in which the behavior is more complicated due to the competition between CBC and gamma. Additionally, we find that gamma induces two types of long-distance entanglement (LDE) in the system under OBC: intraleg LDE and inter-leg one. When the system size is sufficiently large, both types of LDE reach similar strength and stabilize at a constant value. The study indicates that the three-leg ladder makes it easier to generate LDE compared with the two-leg system. However, the generation of LDE is inhibited under CBC which the spin frustration exists. In addition, the calculated results of energy, entanglement entropy and concurrence all show that there are essential relations between these quantities and phase transitions of the system. Further, we predict a phase transition point near gamma=0.54 under OBC. The present study provides valuable insights into understanding the phase diagram of this class of systems.

cond-mat.str-el

DisenGCD: A Meta Multigraph-assisted Disentangled Graph Learning Framework for Cognitive Diagnosis

Existing graph learning-based cognitive diagnosis (CD) methods have made relatively good results, but their student, exercise, and concept representations are learned and exchanged in an implicit unified graph, which makes the interaction-agnostic exercise and concept representations be learned poorly, failing to provide high robustness against noise in students' interactions. Besides, lower-order exercise latent representations obtained in shallow layers are not well explored when learning the student representation. To tackle the issues, this paper suggests a meta multigraph-assisted disentangled graph learning framework for CD (DisenGCD), which learns three types of representations on three disentangled graphs: student-exercise-concept interaction, exercise-concept relation, and concept dependency graphs, respectively. Specifically, the latter two graphs are first disentangled from the interaction graph. Then, the student representation is learned from the interaction graph by a devised meta multigraph learning module; multiple learnable propagation paths in this module enable current student latent representation to access lower-order exercise latent representations, which can lead to more effective nad robust student representations learned; the exercise and concept representations are learned on the relation and dependency graphs by graph attention modules. Finally, a novel diagnostic function is devised to handle three disentangled representations for prediction. Experiments show better performance and robustness of DisenGCD than state-of-the-art CD methods and demonstrate the effectiveness of the disentangled learning framework and meta multigraph module. The source code is available at \textcolor{red}{\url{https://github.com/BIMK/Intelligent-Education/tree/main/DisenGCD}}.

cs.LG

Diverse Transient Chiral Dynamics in Evolutionary distinct Photosynthetic Reaction Centers

The evolution of photosynthetic reaction centers (RCs) from anoxygenic bacteria to oxygenic cyanobacteria and plants reflects their structural and functional adaptation to environmental conditions. Chirality plays a significant role in influencing the arrangement and function of key molecules in these RCs. This study investigates chirality-related energy transfer in two distinct RCs: Thermochromatium tepidum (BRC) and Thermosynechococcus vulcanus (PSII RC) using two-dimensional electronic spectroscopy (2DES). Circularly polarized laser pulses reveal transient chiral dynamics, with 2DCD spectroscopy highlighting chiral contributions. BRC displays more complex chiral behavior, while PSII RC shows faster coherence decay, possibly as an adaptation to oxidative stress. Comparing the chiral dynamics of BRC and PSII RC provides insights into photosynthetic protein evolution and function.

physics.chem-ph

Competitive optimal portfolio selection in a non-Markovian financial market: A backward stochastic differential equation study

This paper studies a competitive optimal portfolio selection problem in a model where the interest rate, the appreciation rate and volatility rate of the risky asset are all stochastic processes, thus forming a non-Markovian financial market. In our model, all investors (or agents) aim to obtain an above-average wealth at the end of the common investment horizon. This competitive optimal portfolio problem is indeed a non-zero stochastic differential game problem. The quadratic BSDE theory is applied to tackle the problem and Nash equilibria in suitable spaces are found. We discuss both the CARA and CRRA utility cases. For the CARA utility case, there are three possible scenarios depending on market and competition parameters: a unique Nash equilibrium, no Nash equilibrium, and infinite Nash equilibria. The Nash equilibrium is given by the solutions of a quadratic BSDE and a linear BSDE with unbounded coefficient when it is unique. Different from the wealth-independent Nash equilibria in the existing literature, the equilibrium in our paper is of feedback form of wealth. For the CRRA utility case, the issue is a bit more complicated than the CARA utility case. We prove the solvability of a new kind of quadratic BSDEs with unbounded coefficients. A decoupling technology is used to relate the Nash equilibrium to a series of 1-dimensional quadratic BSDEs. With the help of this decoupling technology, we can even give the limiting strategies for both cases when the number of agent tends to be infinite.

math.OC

Macroscopic electro-optical modulation of solution-processed molybdenum disulfide

Molybdenum disulfide (MoS2) has drawn great interest for tunable photonics and optoelectronics advancement. Its solution processing, though scalable, results in randomly networked ensembles of discrete nanosheets with compromised properties for tunable device fabrication. Here, we show via density-functional theory calculations that the electronic structure of the individual solution-processed nanosheets can be modulated by external electric fields collectively. Particularly, the nanosheets can form Stark ladders, leading to variations in the underlying optical transition processes and thus, tunable macroscopic optical properties of the ensembles. We experimentally confirm the macroscopic electro-optical modulation employing solution-processed thin-films of MoS2 and ferroelectric P(VDF-TrFE), and prove that the localized polarization fields of P(VDF-TrFE) can modulate the optical properties of MoS2, specifically, the optical absorption and photoluminescence on a macroscopic scale. Given the scalability of solution processing, our results underpin the potential of electro-optical modulation of solution-processed MoS2 for scalable tunable photonics and optoelectronics. As an illustrative example, we successfully demonstrate solution-processed electro-absorption modulators.

cond-mat.mtrl-sci

Generating General Preferential Attachment Networks with R Package wdnet

Preferential attachment (PA) network models have a wide range of applications in various scientific disciplines. Efficient generation of large-scale PA networks helps uncover their structural properties and facilitate the development of associated analytical methodologies. Existing software packages only provide limited functions for this purpose with restricted configurations and efficiency. We present a generic, user-friendly implementation of weighted, directed PA network generation with R package wdnet. The core algorithm is based on an efficient binary tree approach. The package further allows adding multiple edges at a time, heterogeneous reciprocal edges, and user-specified preference functions. The engine under the hood is implemented in C++. Usages of the package are illustrated with detailed explanation. A benchmark study shows that wdnet is efficient for generating general PA networks not available in other packages. In restricted settings that can be handled by existing packages, wdnet provides comparable efficiency.

stat.CO

Multidimensional indefinite stochastic Riccati equations and zero-sum stochastic linear-quadratic differential games with non-Markovian regime switching

This paper is concerned with zero-sum stochastic linear-quadratic differential games in a regime switching model. The coefficients of the games depend on the underlying noises, so it is a non-Markovian regime switching model. Based on the solutions of a new kind of multidimensional indefinite stochastic Riccati equation (SRE) and a multidimensional linear backward stochastic differential equation (BSDE) with unbounded coefficients, we provide closed-loop optimal feedback control-strategy pairs for the two players. The main contribution of this paper, which is of great importance in its own right from the BSDE theory point of view, is to prove the existence and uniqueness of the solution to the new kind of SRE. Notably, the first component of the solution (as a process) is capable of taking positive and negative values simultaneously. For homogeneous systems, we obtain the optimal feedback control-strategy pairs under general closed convex cone control constraints. Finally, these results are applied to portfolio selection games with full or partial no-shorting constraint in a regime switching market with random coefficients.

math.OC

DiffAug: Enhance Unsupervised Contrastive Learning with Domain-Knowledge-Free Diffusion-based Data Augmentation

Unsupervised Contrastive learning has gained prominence in fields such as vision, and biology, leveraging predefined positive/negative samples for representation learning. Data augmentation, categorized into hand-designed and model-based methods, has been identified as a crucial component for enhancing contrastive learning. However, hand-designed methods require human expertise in domain-specific data while sometimes distorting the meaning of the data. In contrast, generative model-based approaches usually require supervised or large-scale external data, which has become a bottleneck constraining model training in many domains. To address the problems presented above, this paper proposes DiffAug, a novel unsupervised contrastive learning technique with diffusion mode-based positive data generation. DiffAug consists of a semantic encoder and a conditional diffusion model; the conditional diffusion model generates new positive samples conditioned on the semantic encoding to serve the training of unsupervised contrast learning. With the help of iterative training of the semantic encoder and diffusion model, DiffAug improves the representation ability in an uninterrupted and unsupervised manner. Experimental evaluations show that DiffAug outperforms hand-designed and SOTA model-based augmentation methods on DNA sequence, visual, and bio-feature datasets. The code for review is released at \url{https://github.com/zangzelin/code_diffaug}.

cs.LG

A Mixed-Membership Model for Social Network Clustering

We propose a simple mixed membership model for social network clustering in this paper. A flexible function is adopted to measure affinities among a set of entities in a social network. The model not only allows each entity in the network to possess more than one membership, but also provides accurate statistical inference about network structure. We estimate the membership parameters using an MCMC algorithm. We evaluate the performance of the proposed algorithm by applying our model to two empirical social network data, the Zachary club data and the bottlenose dolphin network data. We also conduct some numerical studies based on synthetic networks for further assessing the effectiveness of our algorithm. In the end, some concluding remarks and future work are addressed briefly.

stat.AP