SearcharxivSearch

arXiv subjects

Shengbin Ye

Publications and source records attributed to Shengbin Ye.

4 recordsLinked to original sources

Theoretical and experimental studies of energy modulation to demodulation in seeded free-electron lasers

Laser manipulation plays a critical role in precisely tailoring relativistic electron beams through energy modulation, enabling the generation of coherent, intense, and ultrashort radiation in accelerator-based light sources such as synchrotron radiation facilities and free-electron lasers (FELs). However, laser-induced energy modulation inevitably degrades electron beam quality by increasing the energy spread, thereby limiting high-repetition-rate operation. Here, we investigate energy modulation and demodulation in a seeded FEL using two modulators separated by a tunable phase shifter. Analytical analysis and three-dimensional simulations show that a $\pi$ phase delay can nearly reverse the laser-beam interaction and substantially suppress the residual modulation. Diagnostics based on coherent undulator radiation and time-resolved measurements are established to characterize weak residual modulation, and a dedicated demodulation undulator is designed for controlled studies. Preliminary experiments performed at the Shanghai soft X-ray FEL facility using the existing seeding beamline demonstrate laser-induced energy-modulation suppression. Together with the analytical and numerical studies, these results establish a practical framework for investigating the transition from energy modulation to demodulation in seeded FELs, with potential applications in high-repetition-rate, fully coherent X-ray sources with improved preservation of electron beam quality.

physics.acc-ph

Posterior Summarization for Variable Selection in Bayesian Tree Ensembles

Variable selection remains a fundamental challenge in statistics, especially in nonparametric settings where model complexity can obscure interpretability. Bayesian tree ensembles, particularly the popular Bayesian additive regression trees (BART) and their rich variants, offer strong predictive performance with interpretable variable importance measures. We modularize variable selection with Bayesian tree ensembles into two components, the tree prior and the posterior summary, and show that, although typically framed as a modeling task, it often hinges on posterior summarization, which remains underexplored. To this end, we introduce the VC-measure (Variable Count and its rank variant) with a clustering-based threshold. This posterior summary is a simple, tuning-free plug-in that requires no sampling beyond the standard model fits used by existing methods, integrates with any BART variant, and avoids the instability of the median probability model and the computational cost of permutations. In a large-scale benchmark of 3,600 settings built on 100 nonlinear physics equations, it yields uniform $F_1$ gains for both general-purpose and sparsity-inducing priors; when paired with the Dirichlet Additive Regression Tree (DART), it overcomes pitfalls of the original summary and attains the best overall balance of recall, precision, and efficiency. Practical guidance on aligning summaries and downstream goals is discussed.

stat.ME

Ab Initio Nonparametric Variable Selection for Scalable Symbolic Regression with Large $p$

Symbolic regression (SR) is a powerful technique for discovering symbolic expressions that characterize nonlinear relationships in data, gaining increasing attention for its interpretability, compactness, and robustness. However, existing SR methods do not scale to datasets with a large number of input variables (referred to as extreme-scale SR), which is common in modern scientific applications. This ``large $p$'' setting, often accompanied by measurement error, leads to slow performance of SR methods and overly complex expressions that are difficult to interpret. To address this scalability challenge, we propose a method called PAN+SR, which combines a key idea of ab initio nonparametric variable selection with SR to efficiently pre-screen large input spaces and reduce search complexity while maintaining accuracy. The use of nonparametric methods eliminates model misspecification, supporting a strategy called parametric-assisted nonparametric (PAN). We also extend SRBench, an open-source benchmarking platform, by incorporating high-dimensional regression problems with various signal-to-noise ratios. Our results demonstrate that PAN+SR consistently enhances the performance of 19 contemporary SR methods, enabling several to achieve state-of-the-art performance on these challenging datasets.

stat.ML

Operator-induced structural variable selection for identifying materials genes

In the emerging field of materials informatics, a fundamental task is to identify physicochemically meaningful descriptors, or materials genes, which are engineered from primary features and a set of elementary algebraic operators through compositions. Standard practice directly analyzes the high-dimensional candidate predictor space in a linear model; statistical analyses are then substantially hampered by the daunting challenge posed by the astronomically large number of correlated predictors with limited sample size. We formulate this problem as variable selection with operator-induced structure (OIS) and propose a new method to achieve unconventional dimension reduction by utilizing the geometry embedded in OIS. Although the model remains linear, we iterate nonparametric variable selection for effective dimension reduction. This enables variable selection based on ab initio primary features, leading to a method that is orders of magnitude faster than existing methods, with improved accuracy. To select the nonparametric module, we discuss a desired performance criterion that is uniquely induced by variable selection with OIS; in particular, we propose to employ a Bayesian Additive Regression Trees (BART)-based variable selection method. Numerical studies show superiority of the proposed method, which continues to exhibit robust performance when the input dimension is out of reach of existing methods. Our analysis of single-atom catalysis identifies physical descriptors that explain the binding energy of metal-support pairs with high explanatory power, leading to interpretable insights to guide the prevention of a notorious problem called sintering and aid catalysis design.

stat.ME